One sort of bad thing came out of it: our trusty server box really didn't like being shut down. When we set it up at the new place, it didn't POST the first time we turned it on. But did we panic? We did not. Hit the reset button and try again. Ah there, it's fine. We shifted it a week later because its location was less than optimal, pretty much the same story. Again, we're not worry warts, we just hit reset until it booted.
That was stupid move #1.
Stupid move #2. The unit was giving off some sort of high-pitched whine - I have to take everybody's word for it, because I lost that range a long time ago. Anyhow, to combat the noise, we put the box in a cabinet. The cabinet isn't hermetically sealed or anything, but there was only about 5 or 6 inches of space at most around the unit, and I expect there wasn't a great deal of circulation happening in there.
Stupid move #3. SVN started acting funky. People would check stuff in, and the next time somebody updated, they'd get various error messages and the update invariably failed. Huh. So my partner in crime/sysadmin and I went to work on the SVN revisions with a pair of pliers and a blow torch. We restored what we could from the backups, and found some scripts that were known to fix symptoms of this kind for the newer revisions.
After 11 hours of this, we had a kernel oops. For those who have not yet experienced an oops, think of it as a not-quite-panic. The system definitely freaks out, but the kernel doesn't lock it up; it carries on and hopes for the best. The best was not our lot that evening, and a lot of zombie processes starting showing up. Fine.
shutdown -r now. Except that didn't work either - shutdown was zombifying (if I may coin that term). Sigh. Hard boot.Now. After 3 weeks of being cooped up in a hot cabinet, the system decided that I should serve as the example to others. Wouldn't POST. I thought perhaps if I could bring up the BIOS setup, maybe I could figure out what was up. The setup screen froze after about 10 seconds. 3 times in a row. Suck. So I waited 15 minutes and tried again. No dice. 30 minutes. Fuck you, Jack. Gave it an hour, the thing finally POSTs. I have 15 minutes to spare before I'd miss my last bus. Not happy, but at least everything's working and I can go home.
I vowed to get the boss to buy us new hardware the very next day. He wasn't planning on replacing it until the end of the year, but he's a technical person too, so he understands this needs to be done now. He commits to getting the hardware next week (this is now Friday) and starts speccing new gear. Great.
However, the cosmos had other plans for me. Over the long weekend, unbeknownst to me, the SVN is still acting up and the system gets 2 or 3 more hard reboots. I come in on Tuesday to find the RAID on which the repositories live is dead. D-E-A-D. As a fucking doornail. Neat.
I have a Linux box under my desk that I was using as a test bed for MySQL cluster. That testing's all done now, and except for the database, everything else is basically a stock install. I installed what I needed to get SVN working more or less as it did on the old box, and copied the (week-old) backups over. Took a little while because our repos are huge - well over half my day is eaten up by copying, exporting, reloading, verifying.
But now it's all done, and with a little care we managed to save all the history minus the past week, and no data was lost because of course everybody has the files they worked on still sitting on their own desktops.
So what's the moral of the story? If it looks like a duck, walks like a duck, and quacks like a duck - it's probably a duck. Pretending that a dying system isn't really dying is probably not the best strategy for data integrity.
That's good times right there.
ReplyDelete- Arch