Somewhat Scalable Slide Scanning
Photography has been a meaningful part of my life and a reasonable chunk of the posts here in large part because of my stepfather, who was a professional photographer and educator for over 50 years. Sadly, he passed away this past summer, though fortunately we were able to capture some of his basic material on film photography in a series of videos.

Much of his work is captured on film, in a variety of sizes (35mm, 120, panorama, large format), in both negative film and positive slide formats. And there’s a lot of it; we don’t have a count, but it’s likely on the order of 50,000 – 100,000 photos, filling multiple floor-to-ceiling cabinets. This leaves us with the daunting task of trying to archive the collection – via high quality digital scans, and retention of the film for what feel like the best shots. Commercial scanning at that volume, especially for mounted slides, is prohibitively expensive. A digital camera with a macro lens can generate a fairly high quality transfer, but using say Nikon’s ES-2 film scanning adapter (image theirs) for a 60mm macro lens still requires manually inserting one slide/negative strip at a time:

Since most of the collection was in 35mm slide format, we decided to start with that. I wanted to automate as much of this as possible given the volume – but didn’t anticipate it would wind up being a fairly complex systems project that explored several flavors of AI…
Hardware Attempt #1
One thing I did in during six of the years where I didn’t post anything was coach a team my kids were part of, competing in FIRST Robotics at the middle school level. That league uses Lego Mindstorms (EV3, and later Spike Prime) as their platform for creating a set of challenges on a 4′ x 8′ board. It’s a great program! It also meant I had a lot of Lego to build with.

This first attempt feeds slides out of a stack that can hold up to ~50 into an arm that raises the slide in front of a light source, the idea being that software running on a PC connected to the camera would automate capturing the image and trigger advancing to the next slide. This did ultimately sort of work!
Sadly, there were several significant issues. The size and corresponding amount of force was pretty high for Lego, and that meant the rig was relatively slow, and not always reliable. It was hard to ensure that the angle was very close to a perfect 90 degrees to avoid any skew in the image. Also, I’d planned to use a color sensor stuck to a laptop screen as a trigger, but that only works with the older EV3 color sensors; the newer Spike Prime used here is immune to any incident light that’s not reflected from its own hardware. That’s a real problem because there are no supported methods for communicating with Lego software from a PC (unless you root your brick and install your own firmware, which I didn’t want to do).
Hardware Attempt #2
The second attempt simplified things significantly having the camera mounted vertically, shooting downward to eliminate the need to have a big arm to lift the slide into position.

A key part of the simplification was having a small LED video light – the kind that can go in a flash hotshoe – mounted under the main part of the rig, instead of using a large softbox.

To solve the communication issue, I stuck with using the color sensor, but trigger it by waving a physical object in front of it (in this case, a magenta Lego brick) using a stepper motor controlled by an Arduino that responds to serial communication over USB from the host PC:

The slide feeder itself is fairly straightforward, a door alllows a stack of slides to be placed inside with slides being pushed out the bottom one at a time using Lego’s motor and gear system.

It took some tweaking and there are still occasional jams in the feeder, but all in all this approach was sufficiently reliable to automate a lot of the scanning for the first batch of ~1,000 slides.
A big thanks to digicamcontrol, a free, open source utility for tethered capture that supports a wide range of cameras (including the D850 used in this project). It includes a CLI for scriping despite being primarily a GUI-based tool, which was very handy for this effort!
Cropping Complexity
35mm are slides are roughly 36mm x 24mm, but can be either landscape or portrait orientation depending on how the camera was oriented when the picture was taken. I thus had the choice of either (1) placing all slides in one arbitrary orientation and rotating digitally after the fact, or (2) placing slides in their original orientation and cropping, at the expense of some resolution. I chose the latter, since (1) would have required a semi-manual pass to verify the orientation of every slide after scanning.
Still, even without inferring orientation based on the photo, cropping was tricker than I expected. When you look at a slide above, the white plastic slide case seems easy to distinguish from the photo, but during scanning the film is illuminated more brightly – though dark areas in the photo will still be darker than the plastic case. The image on the film itself also has a border, Lego isn’t 100% precise so the position on each scan will be slightly different, and the camera isn’t going to be pixel perfect parallel to the slide.
All this still seemed pretty easy to solve, so I asked AI (using Google’s Gemini) to do so. It was great at building the scaffolding, and correctly cited a number of canonical algorithms for edge finding, but even after much prompting and attempted refinement, what AI produced completely failed on the actual images. It seemed really credible and knew much more about computer vision than I did, but in the end I had to code a heuristic by hand that iterates on finding each edge, ranking candidates based on a combination of contrast and variation on one side of the line.

Exposure Bracketing
One of the most amazing things about the film days is that photographers generally wouldn’t know for days or weeks if they actually got a shot, since film needed to be developed and printed (or mounted) first. One technique to mitigate the risk of having something wrong is exposure bracketing – figure out what you think the right exposure is, then take a series of 3 shots at +1, 0, -1 Ev so that even if the target exposure was a little off, you still have a photo that works. However, what that meant for me is that there were often 3, 4, or 5 of essentially the same photo at different exposure levels. Besides the exposure (brightness) being different between shots, things will often have shifted slightly (flowing water, trees blowing, people or animals moving, camera motion when not using a tripod). So no simple comparison is possible.
After the AI fail with cropping, I started by trying to code something by hand to compare histograms, expecting that the histogram for any region would be basically the same across exposure stacked photos, modulo a factor for the exposure difference. This completely did not work; I’m not sure if it would even have worked on digital photos but film behaves very differently with significant color shifts at different exposures – I was stunned that even to my eye the histograms were totally different.
Fortunately, I stumbled upon an AI-based approach that worked better. Not AI as in “please write me the code”, but using a generative AI image model in reverse (so to speak) to derive feature vectors corresponding to an image, and then looking at the distance between the feature vectors for two images as a measure of their similarity. This Medium post describes the process; it was straightforward to adapt that approach to clustering similar images in a group of ~40 since the feature vectors are small and faster to compare once computed. The AI model it uses, CLIP-ViT-B-16-plus-240, is free, small (<1GB), and can easily be run locally on modest hardware. It doesn’t get everything right, but it’s close enough to save a good amount of time over stacking everything manually.
Organization and GUI
Once you have even 1,000 photos and expect ~99,000 more, a system for organizing them is important. I’d assumed some folder convention would suffice, but needing to review and override exposure compensation stacking, apply date codes and location/group names, and distinguish which images should be ignored, kept, or highlighted meant I needed a better way to organize things.
Here, I went with a Python backend using Flask to serve a frontend. I wrote the backend entirely by hand, but Gemini did a fantastic job of creating a functional HTML UI given a description of the backend API.

While there’s rightfully some concern about AI replacing the craft that humans have dedicated a long time to refining – coding, if anything, was mine – I’m truly amazed that the AI tools we have make it practical to create a bespoke software system like this, even though it is for a single user (me), for a use case with personal value but zero economic value.
Uploading and Antigravity
A minor add-on piece was to take all photos that I’d tagged as 2 stars or higher, and that were selected as the primary if they were in an exposure compensation stack, and upload them to SmugMug for sharing with family/colleagues – creating galleries as needed. I did this in the last month, so was interested in trying out Google’s newly released Antigravity IDE (vs. copying and pasting code around).
Though this was a straighforward task, I was incredibly impressed by Antigravity. It read through all the existing code to figure out how everything was organized on disk and with JSON sidecar files, and I just gave it a pointer to the online documentation for the SmugMug API, and it more or less did the rest. Sure, it required a number of debugging steps and some simple corrections where it was using the wrong attributes when invoking an API, but overall I was hugely impressed and look forward to playing with Antigravity on a more complex project sometime.
Now all I need is for these AI tools to be capable of going through trays of slides and removing them from the plastic sheets they are in :).
If you somehow made it all the way to the bottom of this long post, I hope something in it was interesting to you!
One Comment
Lim bee hong
Absolutely amazing. Thought designing a movable Lego structure was already a feat. Didn’t realise the further complexity perfecting the images of scanned slides. Also great you created software to select the best 2. Glad this provided the opportunity to test what AI, Gemini’s usefulness.
Thank you for all your hard work Mark….and on behalf of Sugawara san. He could not have believed it. He wanted to test Nikon 850’s scanner but you took it many steps forward with this automatic feeding, capturing and (near to) perfect cropping! with your home grien Lego contraption!