March 16th: Planned Bazooka Hadoop cluster upgrade, reorganisation of backup servers
Bazooka Hadoop cluster
The increasing popularity of the Big Data statistics courses hosted on the Stats section's Bazooka Hadoop cluster has been the catalyst for some further upgrading of this cluster, which was last updated in late 2016. An additional node optimised for teaching use, with 64 processor cores, 528 GB of memory and 8 TB of local disk storage is being added to the cluster; this node, called athena, will be a second 'head node' in parallel with the existing bazooka head node although it will normally be available to research users, during courses it will be dedicated exclusively to teaching use.
At the same time a major upgrade of the Mapr Hadoop ecosystem is planned for around Easter time, taking it from the current version 5.2 to 6.1 along with a Ubuntu operating system upgrade. The new athena node has been installed and user data stored in the existing HDFS distributed filesystem is currently being backed up to a non-HDFS server but there is a lot it and this will take a few days to complete. After the new cluster set-up has been trialled on the Churchill test cluster, it is planned to upgrade the Mortar test cluster as well to provide an alternative facility to the Bazooka cluster while it is being upgraded, which is likely to take several days during which time it will be unavailable for use.
Reorganisation of backup servers
Recent storage upgrades to various group and sectional compute servers - especially the Stats' modal and medial systems - has had the knock-on effect of requiring more back-up storage capacity so that we can continue to hold full backups of users' data. Reorganisation of the contents of the three on-site backup servers is now under way, with backups being moved about between these servers to make better use of the available space. Unfortunately, even with the dedicated internal inter-server networks within Huxley 616, the limiting factor currently is the gigabit (1000 Mbits/second) network speeds and this operation will take time to complete since there is a huge amount of data being moved about.