Current news...


September 2024 update:

Failure of both the clustor2 storage server and its backup mirror, clustor2-backup, through overheating in the ICT data centre on August 22nd has underlined the fact that our research IT accommodation is now sub-standard and more needs to be done to make systems more resilient to unexpected incidents. Unexpectedly, it has also shown SSDs (solid state disks) are possibly less reliable than the older spinning mechanical hard disks when exposed to excess temperatures - one of the SSDs in clustor2 actually went short-circuit across its power supply feed, which suggests the silicon within melted. It also means adding more complexity (and therefore cost) to storage systems where some disks are used for filesystem housekeeping roles rather than the actual data storage; this is the case with Maths large storage servers which use SSDs to accelerate read caching and mapping functions in the ZFS filesystems they use.

To this end another server (clustor5) is now being used as a test bed for using mirrored pairs of nvme SSDs instead of single SATA SSD devices in ZFS pools; SSDs are used for these caching and logging functions since they have a read/write speed many times faster than mechanical hard disks, which is important in servers used for real-time computing in HPC clusters and GPU servers where data throughput can be very high. On clustor2 for example, using separate SSDs to handle caching and logging instead of distributing these functions randomly around the spinning data disks allowed us to achieve data write speeds of over 18 gigabytes a second across the whole disk pool.

However, a failure of the logging disk wipes out the entire dataset on the disk pool whereas a failure of any other single disk in the system will not cause data loss. The clustor2-backup server was intended to provide insurance against such an event happening but unfortunately, being located in the same room, the overheating incident affected this server too, leading to the loss of all data after November 22nd 2022 which was on a different ZFS filesystem. For added resilience we did originally plan to keep the backup server in Huxley using the standard College gigabit network to link the two but the latter proved far too slow once some users started creating around half a terabyte of fresh data during the course of a single night - the backups never ever finished! So the server was moved to the data centre and the two linked together with a short 10 Gbit/s fibre link.

Mirroring both the logging and the caching SSDs with a second identical SSD (essentially a RAID 1 mirror within a RAIDz1 pool!) means one logging SSD can fail without affecting the filesystem in any way. Of course, it is true that an overheating data centre might affect all of the SSDs in a system so we are also adding thermal monitoring to the NextGen cluster to shut it down completely if the ambient temperature becomes excessive. A fairly recent development in the SSD world has been nvme SSDs, which connect directly to a system's PCIe bus without an intermediate traditional disk interface such as SATA or SAS. These are extremely fast and mounted onto an adapter card fitting into a PCIe slot, they don't need to use traditional hard disk bays, which on Maths servers will free up 2 disk bays currently used for SATA SSDs that can be used for adding additional spinning disks.

Once development work and testing of the new SSD set-up has been completed, this upgrade will be rolled out to all storage servers that use ZFS filesystems with separate logging & caching disks. Unfortunately, supply chain problems are delaying this work with only half of the new SSDs required having been delivered to date (September 19th).

Older news items:

August 15th: August 2024 update, introducing clustor3
March 20th: March 2024 update, substantial upgrade of the NextGen HPC cluster started
November 20th: November 2023 update
September 25th: September 2023 update
August 18th: August 2023 update
June 30th: June 2023 update
March 15th: March 2023 update
August 13th: Mid-August 2022 update
July 5th: Early July update, RStudio upgrade for apollo Stats MSc server
June 15th: June update, Huxley server room expansion, Huxley 616 network management improvements
May 4th: April update, tape backups, development of a new Hadoop/Spark platform completed
March 16th: March update, job management for the forrest GPU server, development of a new Hadoop/Spark platform
December 23rd: a mini-HPC for Stats
August 14th: two more compute servers added to the Stats MSc compute pool
July 19th: July update: extensive updates to the Stats MSc compute servers, CUDA and cuDNN updates for nvidia4
June 19th: June update: new archival server, storage upgrade for Keaveny cluster, Firedrake build servers & Big Blue Button server
May 21st: Degond Cluster becomes a full HPC
March 24th: Degond Cluster fully upgraded
February 25th: a GPU server for the StatML CDT
February 1st: Bazooka Hadoop cluster control network switch replaced
September 24th: all Stats' general purpose compute systems now running Ubuntu 18.04
August 9th: new GPU server nvidia4 introduced
April 7th: Magma software upgraded to version 2.25-4
March 17th: a new 8 card GPU server installed
March 2nd: another backup server added
February 28th: internal network expansion, ma-backup4 and GPU servers coming soon
February 6th: more large compute servers for Stats
November 1st: NextGen is shutting down on November 1st in readiness for relocation
September 29th: matlab2018 queue has now been discontinued on the NextGen cluster
August 24th: matlab2018 queue to be discontinued on the NextGen cluster
August 14th: more servers for Stats, remote monitoring upgrades, better system status reporting and more power supplies
June 23rd: more servers in the server room, expanded Bazooka Hadoop cluster now available for use
May 29th: R Shiny server memory replaced & remote management added
April 12th: NextGen cluster Maple 2019 upgrade completed
March 16th: Planned Bazooka Hadoop cluster upgrade, reorganisation of backup servers
February 19th: ma-offsite2 now online
January 19th: Matlab 2018b upgrade ongoing
December 14th: Matlab upgrade to version R2018b started, Stats section compute & storage enhancements completed, silos3 and 4 introduced
September 18th: more local storage for Stats modal server and new PostgreSQL database server launched
August 29th: new 'du' command options, cluster R upgrade and ma-backup3
July 2nd: nvidia3 now has two GPU cards
May 15th: Early summer update
March 29th: Easter update
March 24th: spring update
March 10th: late winter update
December 15th: pre-Christmas update
November 22nd: late November update
October 8th: start of 2017/2018 academic year update
2017: Midsummer's Day update
June 16th, 2017: mid-June update
June 2nd, 2017: Early summer update
April 20th, 2017: Spring update 2
March 22nd, 2017: Early spring update
March 10th, 2017: Winter update 2
February 22nd, 2017: Winter update
November 2nd, 2016: Autumn update
October 21st, 2016: Late summer update 2
October 14th, 2016: Late summer update
February 19th, 2016: Winter update
December 11th, 2015: Autumn update
September 14th, 2015: Late summer update 2
May 2nd, 2015: Spring update 2
April 26th, 2015: Spring update
November 11th, 2014: Autumn update
September 17th, 2014: Summer update 2
July 17th, 2014: Summer update
March 15th, 2014: Spring update
November 2nd, 2013: Summer update
May 24th, 2013: Spring update
January 23rd, 2013: Happy New Year!
November 22nd, 2012: No news is good news...
November 17th, 2011: A revamp for the Maths SSH gateways
September 7th, 2011: Failed systems under repair
August 14th, 2011: Introducing calculus, a new NFS home directory server for research users
July 19th, 2011: a new staging server for the compute cluster
July 19th, 2011: A new Matlab queue and improved queue documentation
June 30th, 2011: Updated laptop backup scripts
June 18th, 2011: More storage for the silo...
June 16th, 2011: Yet more storage for the SCAN...
June 10th, 2011: 3 new nodes added to the Maths compute cluster
May 21st, 2011: Announcing SCAN large storage and subversion (SVN) servers
May 26th, 2011: Reporting missing scratch disk on macomp01
May 21st, 2011: Announcing completion of silo upgrades
May 16th, 2011: Announcing upgrades for silo
April 14th, 2011: Goodbye SCAN 3, hello SCAN 4
March 26th, 2011: quickstart guide to using the Torque/Maui cluster job queueing system
March 9th, 2011: automatic laptop backup/sync service, new collaboration systems launched
May 20th, 2010: Scratch disks are now available on all macomp and mablad compute cluster systems
March 11th, 2010: Introduing job queueing on the Fünf Gruppe compute cluster
October 16th, 2008: Introduing the Fünf Gruppe compute cluster
June 18th, 2008: German City compute farm now expanded to 22 machines
February 7th, 2008: new applications on the Linux apps server, unclutter your desktop
November 13th, 2007: aragon and cathedral now general access computers, networked Linux Matlab installation upgraded to R2007a
September 14th, 2007: Problems with sending outgoing mail for UNIX & Linux users
July 23rd, 2007: SCAN available full-time over the summer vacation, closure of Imperial's Usenet news server
May 15th, 2007: Temporary SCAN suspension, closure of the Maths Physics computer room, new research computing facilities
January 14th, 2005: Exchange mail server upgrade, spam filtering with pine and various other enhancements


Andy Thomas

Research Computing Manager,
Department of Mathematics

last updated: 19.09.2024