While the thoughts of almost everyone else turn to annual holidays, staycations, BBQs, painting the outside of the house or simply lazing in the garden the summer holiday season is actually the busiest time of year for research IT, when we get ready for the next academic year beginning in October. Apart from the summer MSc project season, usage of IT facilities is typically low at this time so it's a good time to undertake updates and upgrades; here are some of the things to look forward to in the next few months:
NextGen cluster upgrade: two major upgrades are planned for this cluster:
replacement of macomps09-16: following the upgrade of the eight macomp01-08 nodes with Dell R630 servers last autumn, another eight macomp nodes macomp09-16 have been replaced with identical R630 servers. Fitted with two 10 core CPUs and 192GB of memory each, these have added a significant boost to the cluster's job handling capacity and throughput. At present there are no scratch disks in the new nodes owing to the supply chain problems affecting the IT industry as a whole, but these will be added later when we can get hold of them. The old R410 nodes from this cluster have been moved to Huxley and will join the Stats section's Hadoop cluster which is also being rebuilt & modernised this summer (see below).
replacement of mablads01-16: by today's standards the 16 blade nodes mablad01-mablad16 have a very modest specification; dating from 2008, they have dual quad-core Xeon CPUs and 16 GB of memory each (a few blades that were subsequently fitted as replacments for failed blades have 32GB of installed memory). The CPUs do not support the more recent instruction sets including AVX, AVX2, etc which sometimes means compiling different versions of specialist applications and libraries specifically to run on these older processors. All 16 blades and the chassis into which they plug into will soon be replaced by 10 Dell R630 servers the same as in the macomp node upgrade (above).
Bazooka Hadoop cluster upgrade/rebuild: this 15 node specialist cluster began life in late 2013 running the then brand-new MapR Hadoop distribution based on Apache's software suite of the same name. Using the free community edition of MapR, this was used both for research and Big Data/AI teaching courses. Unfortunately, MapR started getting into financial difficulties in 2016 and we have had to keep both the Linux and the MapR versions frozen as of summer 2016 since no further community edition updates were released and the cluster was required every spring for a Big Data course involving many students. (Eventually MapR went bankrupt, was bought out by HP Enterprise and relaunched as an expensive commercial suite, running on HPE servers and requiring costly support contracts for continued use). In the meantime, although still suitable for teaching existing MapReduce and AI courses, the usability of the cluster for research has declined over the years since it is python2-based and does not support modern python3 applications and libraries.
With this year's Big Data course now completed, the decision has been taken to completely rebuild the cluster using the same hardware but with the very latest Linux and Apache Hadoop components, following a successful pilot earlier this summer using a single node psuedo-cluster server (aphrodite.ma) for teaching AI courses. Using only open source code compiled locally on the cluster ensures we are not trapped into using commercial Hadoop distributions that subsequently run into commercial problems, acquisitions, etc. At the same time the opportunity will be taken to expand the cluster by adding the eight ex-NextGen R410 servers to it (see above).
Stats HPC cluster upgrade: since its introduction in January this year the Stats HPC has been little-used so only the submission node fallas, with 24 processors, has remained in operation throughout with the 8 compute nodes stats01-08 being powered off. In early August it was fully powered up for an upgrade & update before the compute nodes were again shut down as a precaution during the heatwave. However, if the demand arises it is easy to power one or more nodes back on to povide the full 88 processor capability.
rizzuto storage upgrade: this compute server, a sister to gehrig, now has the same fast 5 disk XFS-based pool as its sibling to increase the available local storage to nearly 8 TB.
Huxley server room: it is now becoming clear that simple expansion of the Huxley 616 server room will not be a good solution long-term and we are now looking at relocating the entire facility to another much larger room.