My TLDR for Tsubame 4 @ISCT
Also see Altair Grid Engine Docs and Tsubame’s Docs.
Compared to Slurm
| Goal | Slurm | TSUBAME (Grid Engine) |
|---|---|---|
| Submit a batch | sbatch job.sh |
qsub -g [group] job.sh |
| Interactive | srun --pty bash |
qrsh -g [group] -l node_q=1 -l h_rt=1:00:00 |
| See my work | squeue --me |
qstat |
| Show a job | scontrol show job ID |
qstat -j ID |
| Cancel | scancel ID |
qdel ID |
| See Nodes/Queues | sinfo |
qstat -g c ※ / qhost ※ |
| History | sacct -j ID |
qacct -j ID ※ / portal ジョブ一覧 |
| Make one depend on another | --dependency=afterok:ID |
-hold_jid ID |
| Mail me when ending | --mail-type=END |
-m e |
| In-script command | #SBATCH |
#$ |
| Specify resource | --gres=gpu:A100:4 |
-l node_f=1 (node_f is a type) |
| Specify time | --time=24:00:00 |
-l h_rt=24:00:00 |
※: Maybe some differences
Resources (Nodes) List
Also see 58.html for reference.
node_f4*H100 192cores 1.00 (nominated pts per hour, not really deducted)node_h2*H100 96cores 0.50node_q1*H100 48cores 0.25gpu_11*H100 8cores 0.20
MPI plus multiple gpu_1’s are often a good choice.