A while ago I remembered a website I made for a shipyard in 2001. In the days following, I couldn't shake the idea that there might still be remnants of that site around, so on a Friday night I started searching...
I will keep the actual company unnamed, but I'm sure some of you can do your own research ;-)
I quickly found out that the domain name I remembered had been redirected to the website of a municipality since the late '90s. The shipyard's current web address also wasn't the one I remembered. The first dead end.
Then I turned to AI to see if it knew if there had been other owners of the domain I remembered. It said there had been a legal debate around the domain name ('valued at 350 euros') in 2007, but in the legal documents it presented, there was no mention of that particular domain name. Another dead end.
But what if I had remembered the domain name incorrectly? Could AI find out if the current company had had any previous domains? Bingo! And trying it with a different domain extension, it also returned hits in the cache of the Wayback Machine.
It was a joy to see my old work back: the Flash-intro with sound (still working flawlessly thanks to Ruffle), the bordeaux-red matching my taste at the time, and the little language switcher. The site was much smaller than I remembered, but then again, I created it on my 800×600-pixel blue iMac.
Clicking further I immediately noticed that — just like any old wreck — some parts were missing. At first there was the joy of being face to face with that memory, and then facing the reality. To stay in shipyard terms: the bow was gone. Other parts were still in immaculate condition. Unweathered by 25 years of web turmoil.
Of course I wanted to save what was left from further digital degradation and link-rot, so I decided the whole thing had to come back into my possession. It had to find its last resting place on my laptop's SSD. To be taken care of (and backed up) endlessly.
I tried a simple 'download this webpage', but that version contained a lot of archive.org-specific code, links weren't localized, and Flash didn't run. I was eager, but I wasn't willing to download all of the 120 pages by hand, stripping out the archive.org code and renaming links and assets. I tried some of the existing tools, but they would start pulling in the whole of archive.org or following external links to long-dead websites.
I started writing my own tool that could batch-download a single domain from the Wayback Machine, strip out particular code and localize links. And also serve the site on a local server, so Ruffle could also do its thing. Then I let it run for a while.
Needless to say, the website is now neatly tucked away in one of my folders, and backed up to my private cloud. But something more significant happened than just this one website being saved. What else was out there, in crummy corners of the cache? What else could I resurrect and put out there for myself (and others) to enjoy? That's how I became the Web Archeologist.