Stats section computing facilities


Six different types of compute facilities are available to users in the Stats section:

Because many of the various research projects undertaken in this section overlap to some extent and/or use overlapping/complementary computing technologies, these compute facilities are interlinked to varying degrees so that data stored in one research area can be made available to another. Computing demands in Stats can be very reactive with a need to set up new facilities at short notice, so a flexible environment that supports this is required.

All of the compute systems, Hadoop cluster and compute servers run the server edition of Ubuntu Linux and have a wide range of mathematical and statistical software packages installed in addition to the standard Linux applications. The four storage servers run FreeBSD UNIX, using ZFS as the disk storage pool technology.

General purpose HPC cluster

Replacing the original five high performance general purpose standalone compute systems fallas, festival, fiesta, fira and hustler, the Stats HPC went live at the beginning of 2022, with the first three servers listed above moving to a new rack where they joined 8 ex-NextGen HPC compute nodes which were replaced in October of 2021, forming a new 11 node HPC cluster with 128 processors.

fallas is now the submission node for this cluster but to retain some backward compatibility with the 'traditional' interactive nature of the original standalone Stats compute servers, users in the Stats research section can continue to log into fallas remotely using ssh and from there, log into festival and fiesta as before and run jobs such as R, Matlab, Maple and many other packages. Using the cluster in this way bypasses the job scheduling and cluster management facilities so you are free to run whatever you wish whenever you wish.

Home directories on this cluster are by default on the department's own in-house clustor fileserver and in addition, other Maths servers including silos 1-4, calculus and clustor2 are mounted on these compute systems as well as the College's central ICNFS service; the Stats sections' secure servers can also be mounted on this cluster.

MSc project compute systems

Four systems specifically intended for MSc project use are available:

apollo          2 CPUs, 12 cores, 192 GB memory
artemis       2 CPUs, 12 cores, 192 GB memory
hydra          4 CPUs, 64 cores, 512 GB memory
zeus            4 CPUs, 64 cores, 512 GB memory

and in addition to all the software provided by the general purpose systems above, these can be customised with additional software to support a particular project. For example, artemis has Hadoop hdfs utilities installed to enable working with hdfs archives while apollo is directly accessible via ssh from outside the College (no need to use VPN, ssh gateways, etc) which makes it possible to use Jupyter Notebook over an ordinary ssh connection as described here (for Windows users) and also, the Spyder Python IDE and R Studio Server are recent additions to the apollo server. For more information on the Stats MSc systems please see the dedicated MSc compute server documentation.

These systems also have access to the same specialist data repositories that the large M-series Machines compute servers have, with access permissions configurable on a 'need to have' basis.

Bazooka Hadoop cluster

A 16 node Hadoop cluster is available to Stats users - this is often called the Bazooka cluster since the head node which research users log into to use the cluster is known as bazooka.ma. Completely rebuilt in January 2023 with vanilla reference-standard Hadoop components from the Apache Foundation now replacing the MaPR distribution used previously, this cluster provides 512 processor cores, 1739 GB of memory and 167 terabytes of storage. Another node called athena with 64 processors and half a terabyte of memory and intended primarily for teaching courses, was added to this cluster in June 2019.

If you want to use the Bazooka Hadoop cluster, just ask for an account; this will include both a local conventional home directory on bazooka as well as a Hadoop home directory whose storage is distributed throughout the entire cluster using the HDFS filesystem.

Mortar and Churchill Hadoop test clusters

There are two other small 4 node Hadoop clusters, with 8 processor cores, 32 GB of memory and 27 TB of storage which are normally used for test and development purposes but these can be made available for general use on request.

aphrodite, a single-node Hadoop compute server

With 16 processor cores, 96 GB of memory and 130 GB of storage aphrodite was introduced in March 2022 specifically to meet the needs of a new machine learning/AI course module. Running the latest available Apache Hadoop and Spark software along with Jupyter Notebook and directly accessible by users off-campus, this single-node implementation offers the same modern Hadoop environment as the Bazooka Hadoop cluster.

The M-series Machines - large compute servers for special projects

Known as modal, model, medial, madul and midal, five compute servers are available each having 64 CPU cores (four 16-core AMD Opteron CPUs) and 512 GB of memory; in addition modal has 8 TB of resilient storage shared across all of the M-series machines and medial also has its own 10 TB of local disk storage. These servers are used mainly for cyber security-related projects with accounts being set up on request - all run Ubuntu 22.04 and have R version 4.4.1 installed.

Storage servers

Four dedicated fileservers are installed in the Stats section - two have over 10 TB capacity and are named fusion and enkidu; fusion can be used by any Stats user and accounts are set up on request while enkidu is reserved for security-related projects. The other two servers, flowdata3 and an identical mirror server flowdata3-backup each with a capacity of 60 TB, contain the Netflow data archive.

Data on fusion is available on all of the other Stats compute systems under /home/fusion while flowdata3 is connected in a similar way and contains several distinct Netflow archives named netflow_2013, netflow_2016, etc and these can be found under /home/netflow_2013, /home/netflow_2016 and so on on connected compute systems. (The Netflow archive is a continous feed of network analytic data gathered from all of the routers and network switches in the Imperial College networks, updated every 5 minutes and stored in Cisco NetFlow 9 format as a rolling archive of 2 year's worth of data and is used for developing statistical methods for use in cyber-security research).

Note that not all datasources are accessible to all users - this is for security reasons with access being configured at both individual user and group levels on a "need to have" basis.

Miscellaneous servers

Statistics has for a long time hosted one of the two UK CRAN mirrors (Comprehensive R Archive Network) and now also hosts the LANL cyber security events database as well as a number of other lesser-known data sets and a dedicated server for a joint Imperial College/Microsoft cyber-security project.

Internal networks

Because of the size of the datasets now being worked with in the Stats section, the network bandwidth - that is, the speed of the interconnections between the Stats systems - is an important issue. With the exception of hustler, all of the Stats compute systems are accommodated in the Maths server room where three dedicated internal networks - one each for M-series machine storage, Hadoop cluster data and Netflow data - have now been installed to enhance both bandwidth and security for Stats data. In addition to a normal College gigabit (1000 Mbits/second) network connection, each system now has additional connections to each of these three separate gigabit networks to ensure quick data transfer between systems. (hustler is in a staff office on another floor where it is not possible to connect systems to server room networks).



Andy Thomas

Research Computing Manager,
Department of Mathematics

last updated: 10.12.2024