Pegasus and Triton QuickStart

Pegasus and Triton are shared high-performance computing (HPC) clusters available to authorized University of Miami researchers. Both systems use the LSF scheduler to run computational work on compute hosts.

Cluster

Compute hosts

Architecture

Typical use

QuickStart queue

Pegasus

282

x86_64

General-purpose HPC

general

Triton

72

IBM POWER9 (ppc64le)

GPU-accelerated and data-intensive workloads

short

This QuickStart demonstrates a complete workflow from a local computer to the cluster and back:

dataset + Python script + LSF job script
                   |
                   |  SCP or SFTP
                   v
           cluster scratch storage
                   |
                   |  LSF batch job
                   v
              result file
                   |
                   |  SCP or SFTP
                   v
              local computer

Before You Begin

You need:

  • an active IDSC account with access to Pegasus or Triton;

  • the name of an approved LSF project;

  • access to that project’s scratch directory; and

  • a connection method provided for your account, such as the University of Miami network, VPN, or a configured SSH jump host.

In the commands below:

  • replace <username> with the username associated with your IDSC account;

  • replace <project> with your LSF project name; and

  • use general as <queue> on Pegasus or short on Triton.

Note

The complete QuickStart workflow has been tested on both Pegasus and Triton. The same dataset and Python program work on both clusters. The login host, queue, and reported system architecture differ by cluster.

1. Connect to a Cluster

Open a terminal application. On Windows, use PowerShell or Windows Terminal. On macOS or Linux, use Terminal.

Pegasus

ssh <username>@pegasus2.idsc.miami.edu

Triton

ssh <username>@t2.idsc.miami.edu

Enter the password or complete the authentication method associated with your IDSC account when prompted. After login, you will be on a cluster login node.

Note

If your account uses a configured SSH alias or jump host, use the same SSH destination supplied for your account instead of the host shown above.

Important

Login nodes are intended for tasks such as editing files, transferring data, loading software, and submitting jobs. Run computational work through the LSF scheduler rather than directly on a login node.

For additional SSH options, see Connecting to IDSC Systems from offsite. For graphical applications, see Forwarding the display with x11.

2. Create a Scratch Working Directory

On the cluster, create a directory for the QuickStart files:

mkdir -p /scratch/<project>/$USER/quickstart
cd /scratch/<project>/$USER/quickstart

Confirm your location:

pwd

The output should resemble:

/scratch/<project>/<username>/quickstart

The environment variable $PROJECT is not assumed to be defined. Enter the project name directly wherever <project> appears.

3. Prepare the Example Files Locally

Open a second terminal on your local computer and create a directory for the example:

mkdir -p ~/hpc-quickstart
cd ~/hpc-quickstart

The example uses three files:

data.csv      Input dataset
analyze.py    Python analysis program
analyze.job   LSF batch-job script

Create the Dataset

Create data.csv with the following content:

sample,value
A,10
B,15
C,12
D,18

Create the Python Script

Create analyze.py:

#!/usr/bin/env python3

import csv
import sys
from pathlib import Path


def main():
    if len(sys.argv) != 3:
        raise SystemExit(
            f"Usage: {sys.argv[0]} INPUT_CSV OUTPUT_CSV"
        )

    input_file = Path(sys.argv[1])
    output_file = Path(sys.argv[2])

    values = []

    with input_file.open(newline="", encoding="utf-8") as csv_file:
        reader = csv.DictReader(csv_file)

        for row in reader:
            values.append(float(row["value"]))

    if not values:
        raise ValueError(f"No data values were found in {input_file}")

    results = {
        "count": len(values),
        "minimum": min(values),
        "maximum": max(values),
        "mean": sum(values) / len(values),
    }

    with output_file.open("w", newline="", encoding="utf-8") as csv_file:
        writer = csv.writer(csv_file)
        writer.writerow(["metric", "value"])

        for metric, value in results.items():
            writer.writerow([metric, value])

    print(f"Read {len(values)} values from {input_file}")
    print(f"Wrote the analysis results to {output_file}")


if __name__ == "__main__":
    main()

This program reads the values in data.csv and writes summary statistics to summary.csv. It uses only the Python standard library.

Create the LSF Job Script

Create analyze.job:

#!/bin/bash
#BSUB -J quickstart-python
#BSUB -P <project>
#BSUB -q <queue>
#BSUB -n 1
#BSUB -R "rusage[mem=512M]"
#BSUB -W 00:05
#BSUB -o quickstart.%J.out
#BSUB -e quickstart.%J.err

set -euo pipefail

WORKDIR="${LS_SUBCWD:-$PWD}"
cd "$WORKDIR"

echo "Job ID: $LSB_JOBID"
echo "Compute host: $(hostname)"
echo "Architecture: $(uname -m)"
echo "Working directory: $WORKDIR"
echo

python3 analyze.py data.csv summary.csv

Replace <project> with your LSF project. Replace <queue> according to the cluster:

Cluster

Queue

Pegasus

general

Triton

short

The main LSF directives in this example specify:

  • -J – the job name;

  • -P – the LSF project;

  • -q – the queue;

  • -n – the number of CPU cores;

  • -R – the requested memory;

  • -W – the runtime limit;

  • -o – the standard-output file; and

  • -e – the standard-error file.

The token %J is replaced with the LSF job ID.

Confirm that the local directory contains all three files:

ls -l data.csv analyze.py analyze.job

4. Transfer the Files to the Cluster

Transfer with SCP

Run scp from the local computer, not from the cluster login node.

Pegasus

scp data.csv analyze.py analyze.job \
    <username>@pegasus2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/

Triton

scp data.csv analyze.py analyze.job \
    <username>@t2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/

If your account uses an SSH alias or jump host, replace the destination host with the same destination you use successfully for ssh.

Transfer with an SFTP Application

You may instead use an SFTP application such as FileZilla, Cyberduck, or WinSCP.

Use these connection settings:

Setting

Value

Protocol

SFTP

Port

22

Username

Your IDSC account username

Pegasus host

pegasus2.idsc.miami.edu

Triton host

t2.idsc.miami.edu

Remote directory

/scratch/<project>/<username>/quickstart

Connect using the same authentication or jump-host configuration required for SSH. Upload data.csv, analyze.py, and analyze.job to the remote directory.

5. Verify the Uploaded Files

Return to the cluster terminal:

cd /scratch/<project>/$USER/quickstart
ls -l

The directory should contain:

analyze.job
analyze.py
data.csv

You can inspect the files before submission:

cat data.csv
cat analyze.job

6. Submit the LSF Job

Submit the job from the scratch working directory:

bsub < analyze.job

LSF returns a job ID similar to:

Job <123456> is submitted to queue <general>.

On Triton, the queue name in the message will be short.

7. Monitor the Job

Check active jobs:

bjobs

Common states include:

  • PEND – waiting for resources;

  • RUN – currently running;

  • DONE – completed successfully; and

  • EXIT – ended with an error.

This example runs quickly. It may finish before bjobs displays it. If LSF reports No unfinished job found, list recent jobs with:

bjobs -a

For detailed information about a specific job, use:

bjobs -a -l <jobid>

8. Inspect the Output

After the job finishes, list the directory:

ls -l

The job creates files similar to:

quickstart.123456.out
quickstart.123456.err
summary.csv

On a shared filesystem, the LSF output and error files may take a few seconds to appear after a very short job finishes.

View the standard output:

cat quickstart.*.out

Near the end of the file, you should see output similar to:

Job ID: 123456
Compute host: <compute-host>
Architecture: <architecture>
Working directory: /scratch/<project>/<username>/quickstart

Read 4 values from data.csv
Wrote the analysis results to summary.csv

Check the standard-error file:

cat quickstart.*.err

An empty error file indicates that the program did not write an error message.

View the generated result:

cat summary.csv

Expected result:

metric,value
count,4
minimum,10.0
maximum,18.0
mean,13.75

This confirms that LSF ran the Python program on a compute host, passed the dataset to the program, and wrote the result to scratch storage.

9. Download the Result

Run the download command from the local computer.

Pegasus

scp \
    <username>@pegasus2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/summary.csv \
    .

Triton

scp \
    <username>@t2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/summary.csv \
    .

If your account uses an SSH alias or jump host, use the same destination that worked for the upload.

Confirm the downloaded file:

cat summary.csv

You have now completed the full workflow:

prepare files -> upload -> submit -> monitor -> inspect -> download

10. Disconnect

To disconnect from the cluster, run:

exit

Next Steps

Before running production workloads, review the Cluster Usage Guidelines.

Continue with the detailed guides as needed: