Cluster Usage Guidelines

Review these guidelines before running research workloads on Pegasus or Triton. They summarize the main access, login-node, storage, software, and job submission expectations for the shared clusters.

Accounts, Projects, and Policies

Access to IDSC cluster resources is managed by project. Before using Pegasus or Triton, review IDSC ACS Policies and confirm that you are a member of an approved project.

  • Triton: Join a Triton project with an appropriate resource type, such as triton_faculty, triton_student, or triton_education.

  • Pegasus: Join a Pegasus project with the pegasus_project resource type.

Cluster resources, including compute time and scratch storage, are allocated to projects. Contact the project’s Principal Investigator to request membership. See Projects & Resources for details.

Network Access

Connect from the University of Miami network. When working remotely, connect through the University VPN before accessing Pegasus or Triton.

Use Login Nodes Responsibly

Login nodes are shared access points for lightweight tasks, including:

  • editing scripts and configuration files;

  • transferring and organizing files;

  • loading and inspecting software modules; and

  • submitting and monitoring LSF jobs.

Warning

Do not run production, long-running, or resource-intensive computations directly on a login node. Misuse of login nodes can affect other users and may result in account restrictions.

Submit computational work as an LSF batch job or request an interactive compute session. See Serial/Parallel Job Scripts, Pegasus and Triton Queues, and Interactive Jobs.

Use Scratch Storage for Active Jobs

Use the project scratch filesystem for active job input, temporary files, and high-throughput output:

/scratch/<project>/$USER

Home directories are appropriate for scripts, configuration files, source code, and user-installed software. Avoid high-throughput or data-intensive job I/O under /home or /nethome.

See the storage documentation for retention, backup, quota, and data-management requirements.

Request Appropriate Job Resources

Pegasus and Triton use the LSF scheduler to manage shared compute resources. Each submitted job should request resources appropriate for the workload, including:

  • project allocation;

  • CPU cores;

  • memory;

  • runtime; and

  • any required accelerator or special resource.

Requests that are too small may cause jobs to fail. Excessive requests may increase queue time and leave resources unused. Benchmark smaller runs first when resource requirements are unknown.

See LSF Overview for job submission and resource-request guidance.

Manage Software with Modules and Environments

Use the Environment Modules system to access centrally installed software:

module avail
module load <module-name>
module list

Use project-specific environments when additional packages or isolated software stacks are required. See g-modules and the relevant software guides.

Use the Data Transfer Node for I/O-Intensive Work

On Pegasus, do not run long-running or I/O-intensive data-management work on a login node. This includes rsync, scp, SFTP or rclone workflows involving many files; recursive find or du scans; checksum verification; and archive creation or extraction.

Submit this work through the transfers LSF queue, which runs on the dedicated Data Transfer Node (DTN). For example:

bsub -Is -q transfers -P <projectID> -n 1 -W 01:00 bash

Do not SSH directly to the DTN. See Use the Pegasus Data Transfer Node for large workflows for interactive and batch transfer-job examples, monitoring commands, cancellation, and safe parallelism guidance. For transferring data between your local computer and the cluster, use the supported external transfer endpoint described in Transferring Files.

Getting Help

When requesting support, include the cluster name, job ID, relevant commands, and complete error messages. Do not include passwords, private keys, or other credentials.