Pegasus and Triton QuickStart
Pegasus and Triton are shared high-performance computing (HPC) clusters available to authorized University of Miami researchers. Both systems use the LSF scheduler to run computational work on compute hosts.
Cluster |
Compute hosts |
Architecture |
Typical use |
QuickStart queue |
|---|---|---|---|---|
Pegasus |
282 |
x86_64 |
General-purpose HPC |
|
Triton |
72 |
IBM POWER9 (ppc64le) |
GPU-accelerated and data-intensive workloads |
|
This QuickStart demonstrates a complete workflow from a local computer to the cluster and back:
dataset + Python script + LSF job script
|
| SCP or SFTP
v
cluster scratch storage
|
| LSF batch job
v
result file
|
| SCP or SFTP
v
local computer
Before You Begin
You need:
an active IDSC account with access to Pegasus or Triton;
the name of an approved LSF project;
access to that project’s scratch directory; and
a connection method provided for your account, such as the University of Miami network, VPN, or a configured SSH jump host.
In the commands below:
replace
<username>with the username associated with your IDSC account;replace
<project>with your LSF project name; anduse
generalas<queue>on Pegasus orshorton Triton.
Note
The complete QuickStart workflow has been tested on both Pegasus and Triton. The same dataset and Python program work on both clusters. The login host, queue, and reported system architecture differ by cluster.
1. Connect to a Cluster
Open a terminal application. On Windows, use PowerShell or Windows Terminal. On macOS or Linux, use Terminal.
Pegasus
ssh <username>@pegasus2.idsc.miami.edu
Triton
ssh <username>@t2.idsc.miami.edu
Enter the password or complete the authentication method associated with your IDSC account when prompted. After login, you will be on a cluster login node.
Note
If your account uses a configured SSH alias or jump host, use the same SSH destination supplied for your account instead of the host shown above.
Important
Login nodes are intended for tasks such as editing files, transferring data, loading software, and submitting jobs. Run computational work through the LSF scheduler rather than directly on a login node.
For additional SSH options, see Connecting to IDSC Systems from offsite. For graphical applications, see Forwarding the display with x11.
2. Create a Scratch Working Directory
On the cluster, create a directory for the QuickStart files:
mkdir -p /scratch/<project>/$USER/quickstart
cd /scratch/<project>/$USER/quickstart
Confirm your location:
pwd
The output should resemble:
/scratch/<project>/<username>/quickstart
The environment variable $PROJECT is not assumed to be defined. Enter the
project name directly wherever <project> appears.
3. Prepare the Example Files Locally
Open a second terminal on your local computer and create a directory for the example:
mkdir -p ~/hpc-quickstart
cd ~/hpc-quickstart
The example uses three files:
data.csv Input dataset
analyze.py Python analysis program
analyze.job LSF batch-job script
Create the Dataset
Create data.csv with the following content:
sample,value
A,10
B,15
C,12
D,18
Create the Python Script
Create analyze.py:
#!/usr/bin/env python3
import csv
import sys
from pathlib import Path
def main():
if len(sys.argv) != 3:
raise SystemExit(
f"Usage: {sys.argv[0]} INPUT_CSV OUTPUT_CSV"
)
input_file = Path(sys.argv[1])
output_file = Path(sys.argv[2])
values = []
with input_file.open(newline="", encoding="utf-8") as csv_file:
reader = csv.DictReader(csv_file)
for row in reader:
values.append(float(row["value"]))
if not values:
raise ValueError(f"No data values were found in {input_file}")
results = {
"count": len(values),
"minimum": min(values),
"maximum": max(values),
"mean": sum(values) / len(values),
}
with output_file.open("w", newline="", encoding="utf-8") as csv_file:
writer = csv.writer(csv_file)
writer.writerow(["metric", "value"])
for metric, value in results.items():
writer.writerow([metric, value])
print(f"Read {len(values)} values from {input_file}")
print(f"Wrote the analysis results to {output_file}")
if __name__ == "__main__":
main()
This program reads the values in data.csv and writes summary statistics to
summary.csv. It uses only the Python standard library.
Create the LSF Job Script
Create analyze.job:
#!/bin/bash
#BSUB -J quickstart-python
#BSUB -P <project>
#BSUB -q <queue>
#BSUB -n 1
#BSUB -R "rusage[mem=512M]"
#BSUB -W 00:05
#BSUB -o quickstart.%J.out
#BSUB -e quickstart.%J.err
set -euo pipefail
WORKDIR="${LS_SUBCWD:-$PWD}"
cd "$WORKDIR"
echo "Job ID: $LSB_JOBID"
echo "Compute host: $(hostname)"
echo "Architecture: $(uname -m)"
echo "Working directory: $WORKDIR"
echo
python3 analyze.py data.csv summary.csv
Replace <project> with your LSF project. Replace <queue> according to
the cluster:
Cluster |
Queue |
|---|---|
Pegasus |
|
Triton |
|
The main LSF directives in this example specify:
-J– the job name;-P– the LSF project;-q– the queue;-n– the number of CPU cores;-R– the requested memory;-W– the runtime limit;-o– the standard-output file; and-e– the standard-error file.
The token %J is replaced with the LSF job ID.
Confirm that the local directory contains all three files:
ls -l data.csv analyze.py analyze.job
4. Transfer the Files to the Cluster
Transfer with SCP
Run scp from the local computer, not from the cluster login node.
Pegasus
scp data.csv analyze.py analyze.job \
<username>@pegasus2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/
Triton
scp data.csv analyze.py analyze.job \
<username>@t2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/
If your account uses an SSH alias or jump host, replace the destination host
with the same destination you use successfully for ssh.
Transfer with an SFTP Application
You may instead use an SFTP application such as FileZilla, Cyberduck, or WinSCP.
Use these connection settings:
Setting |
Value |
|---|---|
Protocol |
SFTP |
Port |
|
Username |
Your IDSC account username |
Pegasus host |
|
Triton host |
|
Remote directory |
|
Connect using the same authentication or jump-host configuration required for
SSH. Upload data.csv, analyze.py, and analyze.job to the remote
directory.
5. Verify the Uploaded Files
Return to the cluster terminal:
cd /scratch/<project>/$USER/quickstart
ls -l
The directory should contain:
analyze.job
analyze.py
data.csv
You can inspect the files before submission:
cat data.csv
cat analyze.job
6. Submit the LSF Job
Submit the job from the scratch working directory:
bsub < analyze.job
LSF returns a job ID similar to:
Job <123456> is submitted to queue <general>.
On Triton, the queue name in the message will be short.
7. Monitor the Job
Check active jobs:
bjobs
Common states include:
PEND– waiting for resources;RUN– currently running;DONE– completed successfully; andEXIT– ended with an error.
This example runs quickly. It may finish before bjobs displays it. If LSF
reports No unfinished job found, list recent jobs with:
bjobs -a
For detailed information about a specific job, use:
bjobs -a -l <jobid>
8. Inspect the Output
After the job finishes, list the directory:
ls -l
The job creates files similar to:
quickstart.123456.out
quickstart.123456.err
summary.csv
On a shared filesystem, the LSF output and error files may take a few seconds to appear after a very short job finishes.
View the standard output:
cat quickstart.*.out
Near the end of the file, you should see output similar to:
Job ID: 123456
Compute host: <compute-host>
Architecture: <architecture>
Working directory: /scratch/<project>/<username>/quickstart
Read 4 values from data.csv
Wrote the analysis results to summary.csv
Check the standard-error file:
cat quickstart.*.err
An empty error file indicates that the program did not write an error message.
View the generated result:
cat summary.csv
Expected result:
metric,value
count,4
minimum,10.0
maximum,18.0
mean,13.75
This confirms that LSF ran the Python program on a compute host, passed the dataset to the program, and wrote the result to scratch storage.
9. Download the Result
Run the download command from the local computer.
Pegasus
scp \
<username>@pegasus2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/summary.csv \
.
Triton
scp \
<username>@t2.idsc.miami.edu:/scratch/<project>/<username>/quickstart/summary.csv \
.
If your account uses an SSH alias or jump host, use the same destination that worked for the upload.
Confirm the downloaded file:
cat summary.csv
You have now completed the full workflow:
prepare files -> upload -> submit -> monitor -> inspect -> download
10. Disconnect
To disconnect from the cluster, run:
exit
Next Steps
Before running production workloads, review the Cluster Usage Guidelines.
Continue with the detailed guides as needed:
Projects & Resources for project access and allocation information;
Transferring Files for additional transfer methods;
g-modules for software modules;
LSF Overview for batch-job submission;
Interactive Jobs for interactive jobs; and
IDSC ACS Policies for cluster policies.