UTHPC - UT HPC Cluster Rocket and Slurm upgrade – Hooldustöö üksikasjad

Kõik teenused on töökorras!

UT HPC Cluster Rocket and Slurm upgrade

Tehtud
Aeg 21. juuli 2026 kell 6:00:00 –  28. juuli 2026 kell 16:05:48UTC

Mõjutab

rocket.hpc.ut.ee

Hooldustöö alates 6:00 AM kuni 4:05 PM

Services

Hooldustöö alates 6:00 AM kuni 4:05 PM

Galaxy

Hooldustöö alates 6:00 AM kuni 4:05 PM

RStudio

Hooldustöö alates 6:00 AM kuni 4:05 PM

Open OnDemand

Hooldustöö alates 6:00 AM kuni 4:05 PM

Värskendused
  • Tehtud
    28. juuli 2026 kell 16:05:48UTC
    Tehtud
    28. juuli 2026 kell 16:05:48UTC

    Rocket cluster maintenace and SLURM update has been completed. Due to unfortunate circumstances, a subset of running jobs were impacted.

  • Uuendus
    22. juuli 2026 kell 13:40:21UTC
    Uuendus
    22. juuli 2026 kell 13:40:21UTC

    Update: Both login node updates have been completed successfully ahead of schedule. Login1 and Login2 are now fully available, and all temporary SSH connection restrictions have been lifted.

    Next, the Slurm upgrade on the 28th of July will follow. Until then, no interruptions in Rocket cluster computing are expected.

  • Töös
    21. juuli 2026 kell 6:00:01UTC
    Töös
    21. juuli 2026 kell 6:00:01UTC
    Hooldustöö on käimas
  • Lisatud
    17. juuli 2026 kell 12:08:06UTC
    Lisatud
    17. juuli 2026 kell 12:08:06UTC

    Here is the summer maintenance schedule for the HPC Rocket Cluster.

    HPC cluster Rocket updates are scheduled for July 2026 to improve the cluster's performance and capabilities.

    1. Login Node Updates

    We'll be performing system updates on both login nodes on the following dates:

    * Login1: July 21st

    * Login2: July 28th

    To minimize disruption, we'll close new SSH connections to the corresponding login node one week before each update, allowing existing connections to naturally expire. One of the login nodes will remain available at all times, so you won't experience any service downtime.

    2. Slurm Update

    On July 28th, starting at 15:00, we'll be upgrading the Slurm version. Your running jobs won't be affected. However, during the update, submitting new jobs, SLURM commands like sacct, sacctmgr, and related tools will be unavailable. The process should take about two hours, but may run longer.

    The compute nodes will also receive OS, BIOS, and firmware upgrades.

    After July 28th, the compute nodes will be updated in a rolling fashion. This means some nodes will be temporarily drained until all updates are complete, which may result in longer queue times depending on cluster usage.

    The Open OnDemand service, as the UTHPC cluster portal, is also affected.