Warning: This post is long. While working through this massive server upgrade/migration process, tears were shed, many cuss words were said, along with a general feeling of frustration, which ultimately culminated into extreme happiness once the migration was completed. The scale and complexity of the implementation factor into the length of this post, and I’ll share my thought process on how this was executed, so here goes.
n a perfect world, one would upgrade databases one version at a time and not let them get too old. But our databases are where the “crown jewels” are. They must stay up 24×7. When performance is acceptable, it’s acceptable, and sometimes old versions stay around too long. We don’t live in a perfect world. This idea applies to so many things. There’s almost never a perfect data model. There is always some type of resource constraint be it storage, memory, CPU, IOPS, or just plain dollars.
I will bring this concept of not living in a perfect world into a discussion about upgrades.
Ideally there would be…
- …time to do two upgrades. One upgrade to 5.6, the other to 5.7. This is the way sane, normal people upgrade.
- …a lot of extra hardware. It sure would be nice to maybe combine a maintenance like this with a hardware refresh so that we could just have an entirely new MySQL 5.7 database shard, slap a MySQL 5.6 box in the middle as a relay, and once things are all caught up just do a VIP cutover to new machines.
The extra hardware idea I have actually worked with in the past. A client of a company I worked for just had a standalone 5.5 machine with no replicants. They were able to allocate two more servers, one to be a 5.6 relay and one new 5.7 machine that would serve as a final destination. Added bonus… At the end the 5.6 server would be upgraded and left in place for a redundant machine.
In the past I setup some new Pacemaker clustered nodes with a fresh Debian Stretch installation. I followed our standard installation guide, created also shared replicated DRBD storage, but whenever I tried to mount the ext4 storage DRBD detached the disks on both node sides with I/O errors. After recreating it, using other storage volumes and testing my ProLiant hardware (whop I thought it had got a defect..) it still occurs, but somewhere in the middle of testing, a quicker setup without LVM it worked fine, hum..