Access

From iVEC User Documenation
Main Page > Access
Jump to: navigation, search

Contents


Introduction

iVEC has three supercomputers, named Epic, Fornax, and Magnus, which provide complementary resources for a wide range of computational research. User access to each supercomputer is controlled by a scheduler (specifically, PBS Pro), which organizes the computational jobs from different users to ensure an efficient, effective, and fair scheme for sharing the resources. Each system runs a Linux operating system, with a wide range of applications and libraries. The software environment is set up to be as consistent as possible across the three systems, to make it easier for users to move from using one system to another. Furthermore, module files provide pre-configured environments for commonly used software, to ensure a simple and consistent setup process.

Magnus

Magnus is a Cray XC30 supercomputer. That is, it is a massively parallel architecture of many individual nodes that are connected by a high speed network. Each node is based on the x86-64 architecture, which is popular in desktop PCs and enterprise servers.

Magnus is a homogeneous cluster which means there is just one type of processor across all compute nodes; the Intel Xeon E5-2690 v3 (Haswell). Each node has two E5-2690 v3 processors and each E5-2690 v3 has 12 cores. These can be referred to as “dual 12-core” nodes. In the simplest of terms Magnus can be considered to have twenty four cores per node providing a total of 35,712 cores across the 1488 nodes. Each node has 64GB of memory.

Connecting all the nodes together is the HPC optimized Aries interconnect, that utilizes the Dragonfly network topology to enable parallel codes using MPI – Message Passing Interface, to communicate between the compute nodes. Each compute blade has a single Aries ASIC, which provides about 72Gbits/sec for each of the 4 nodes on the blade. The Aries/Dragonfly interconnect has been designed for low latency and efficient small message transfer.

Storage on Magnus is provided by the Lustre parallel file system. Lustre can use multiple servers to store metadata and data, which can improve throughput significantly for large data files. The Lustre filesystem operates over the Infiniband network.

Magnus has two login nodes. The generic hostname magnus.ivec.org distributes users onto either login node via round-robin DNS. To run jobs on the compute nodes, submit jobs from the login nodes. Compute nodes cannot be accessed from the internet.

Magnus is the second stage of the petascale system that is housed in the Pawsey Centre.


Epic

Epic is a HP commodity Linux cluster. That is, it has many individual nodes that are connected by a high speed network. Each node is based on the x86-64 architecture, which is popular in desktop PCs and enterprise servers.

Epic is a homogeneous cluster which means there is just one type of processor across all compute nodes; the Intel Xeon X5660 (Westmere). Each node has two X5660 processors and each X5660 has 6 cores. These can be referred to as “dual hex-core” nodes. In the simplest of terms EPIC can be considered to have twelve cores per node providing a total of 9600 cores across the 800 nodes. Each node has 24GB of memory.

Connecting all the nodes together is an Infiniband switch; this is the high speed interconnect switch. This switch is what enables parallel codes using MPI- Message Passing Interface to communicate between the compute nodes. Each node contains a quad data rate (QDR) Infiniband card which has a peak speed of 40Gbits/sec. Infiniband has a low latency so that small messages can be sent and received very fast.

A HP ProLiant BL2x220c G6 blade, with two CPUs and their memory banks

Storage on Epic is provided by the Lustre parallel file system. Lustre can use multiple servers to store metadata and data, which can improve throughput significantly for large data files. The Lustre filesystem operates over the Infiniband network.

Epic has two login nodes. The generic hostname epic.ivec.org distributes users onto either login node via round-robin DNS. To run jobs on the compute nodes, submit jobs from the login nodes. Compute nodes cannot be accessed from the internet. Epic has two datamover nodes, each with multiple 10Gbps Ethernet cards facing the internet, providing high speed and resilience. These are the preferred route for moving data to and from Epic’s Lustre filesystems. The datamover nodes are accessed via the name epicdata.ivec.org.

Epic is made up of 25 HP BladeSystem c7000 enclosures, each of which contains 16 double-dense HP ProLiant BL2x220c G6 blades. It is housed in a Performance Optimised Datacentre (POD), which is similar in size to a shipping container.


Fornax

Fornax is a SGI Linux cluster. That is, it has many individual nodes that are connected by a high speed network. Each node is based on the x86-64 architecture, which is popular in desktop PCs and enterprise servers. Each node has a NVidia Tesla GPU, which has some similarities with NVidia gaming graphics cards but is optimized for double precision scientific computing. The GPUs provide the majority of the compute power of Fornax, and are better suited to repeated operations on streamed data than CPUs.

Fornax is a heterogeneous cluster which means there is more than one type of processor. Fornax has Intel Xeon X5650 (Westmere) CPUs and Nvidia Tesla C2075 GPUs. Each node has two X5660 processors and each X5660 has 6 cores. These can be referred to as “dual hex-core” nodes. Each node has one C2075 GPU. Fornax has 96 nodes and therefore 1152 CPU cores and 96 GPUs. Each node has 72GB of memory.

Connecting all the nodes together is an Infiniband switch; this is the high speed interconnect switch. This switch is what enables parallel codes using MPI- Message Passing Interface to communicate between the compute nodes. Each node contains a quad data rate (QDR) Infiniband card which has a peak speed of 40Gbits/sec. Infiniband has a low latency so that small messages can be sent and received very fast.

Storage on Fornax is provided by the Lustre parallel file system. Lustre can use multiple servers to store metadata and data, which can improve throughput significantly for large data files. The Lustre filesystem operates over the Infiniband network. On Fornax the Lustre network traffic is isolated from the MPI traffic to maintain high MPI performance during heavy I/O. Fornax also has 7TB per node of local storage, used by some application areas.

Fornax has two login nodes. The generic hostname fornax.ivec.org distributes users onto either login node via round-robin DNS. To run jobs on the compute nodes, submit jobs from the login nodes. Compute nodes cannot be accessed from the internet.

Fornax has two datamover nodes, each with multiple 10Gbps Ethernet cards facing the internet, providing high speed and resilience. These are the preferred route for moving data to and from Fornax’s Lustre filesystems. The datamover nodes are accessed via the name fornaxdata.ivec.org.


Galaxy

Galaxy is a Cray XC30 supercomputer. That is, it is a massively parallel architecture of many individual nodes that are connected by a high-speed network. Each node is based on the x86-64 architecture, which is popular in desktop PCs and enterprise servers: a small number of nodes also have a GPU accelerator to support, amongst others, CUDA-enabled applications.

Galaxy is only available for radio-astronomy-focused research. In particular, it is used to support ASKAP, which is one of the Square Kilometre Array precursor projects currently under way in the north-west of Western Australia. For ASKAP, it acts as a real-time computer, allowing direct processing of data delivered to the Pawsey Centre from the Murchison Radio Observatory.

Galaxy consists of three cabinets that contain 134 compute blades, each of which has four nodes. The majority of nodes (484, in 118 blades) support two 10-core Intel Xeon E5-2690V2 ‘Ivy Bridge’ processors operating at 3.00 GHz with 64GB of memory, for a total of 9,440 cores delivering around 200 TFLOPS of compute power. There are also 64 nodes that each support an NVIDIA K20X ‘Kepler’ GPU, alongside an Intel Xeon E5-2690 ‘Sandy Bridge’ host processor at 2.6 GHz, with 32GB of (host) memory.

Connecting all the nodes together is the HPC-optimized Cray Aries interconnect, which utilizes the Dragonfly network topology to enable parallel codes using MPI – Message Passing Interface, to communicate between the compute nodes. Each compute blade has a single Aries ASIC, which provides about 72Gbits/sec for each of the 4 nodes on the blade. The Aries interconnect and Dragonfly topology has been designed for low latency and efficient small message transfer.

Galaxy has a 1.3 Petabyte Lustre scratch file system, provided by a Cray Sonexion 1600 appliance storage system and connected via FDR Infiniband.

Galaxy has two login nodes. The generic hostname galaxy.ivec.org distributes users onto either login node via round-robin DNS. To run jobs on the compute nodes, submit jobs from the login nodes. Compute nodes cannot be accessed from the internet.

Zeus

Zeus is an SGI Linux cluster. Zeus is primarily used for pre- and post-processing of data, large shared memory computations and remote visualization work. Zeus is a heterogeneous cluster with 30 nodes in various configurations. Zeus shares the same /home, /scratch and /group file systems with Magnus. Users ssh to zeus.ivec.org to access all the 30 nodes in Zeus.

In Zeus, there is one large-memory node called Zythos, which is an SGI UV2000 system. Zythos consists of 24 UV blades, and they appear as a single system with huge shared memory of 6TB. 20 of the blades each contain two 6-core Intel Xeon E5-4610 CPUs, and the other 4 blades each contain one 6-core Intel Xeon E5-4610 CPU and one Nvidia Tesla K20 GPU. Multiple users are permitted to run on Zythos at the same time.

Apart from Zythos, Zeus’s other 29 nodes all have two 8-core Intel Xeon E5-2670 CPUs. 27 of these nodes also feature GPUs, with 20 of them each containing one Nvidia Quadro K5000 GPU and the other 7 nodes each containing one Nvidia Tesla K20 GPU. On these 27 GPU-featuring nodes, only one user is permitted to run at a time on each node.

Time on Zythos is allocated through the iVEC Director’s Share. For Zeus’s other nodes, access is automatically granted to anyone who is a user in an active project on either Magnus, Galaxy or Zythos.

Zeus is part of the portfolio of Pawsey Project resources: it resides in the Pawsey Centre and is closely integrated with other Pawsey Centre infrastructure including the Magnus supercomputer, the RDSI Collection Development node, and the Hierarchical Storage Management system, allowing a diverse range of workflows to be undertaken.

Zythos

Zythos is part of the Zeus cluster. Zythos has 6 TB of memory, 264 Intel Xeon processor cores, and 4 NVIDIA Kepler GPUs. Users ssh to zeus.ivec.org and submit jobs to the zythos partition to access the node Zythos. Zythos uses cpusets to keep user programs separate from each other.

Zythos is primarily for pre- and post-processing of data associated with Magnus or Galaxy. Due to its large amount of shared memory, Zythos is also well suited to scientific problems that cannot easily make use of distributed memory. Zythos provides researchers with an opportunity to undertake novel and previously impractical computational, data-analysis, and visualisation workflows.

AccessSubscribeUser Portal
Main PageSupercomputersData StorageVisualisation