BLAST (Basic Local Alignment Search Tool) compares nucleotide or protein sequences against a database and reports statistically significant matches. This guide covers running BLAST on the UCR High-Performance Computing Center (HPCC) cluster: logging in, choosing a partition, running interactive and batch searches, and tuning resource requests.
The HPCC documentation at hpcc.ucr.edu is the authority on partitions, limits and software. Where this guide and the HPCC site differ, follow the HPCC site.
Logging in to the HPCC
You need an HPCC account first. See Getting an HPCC account.
Connect with SSH. The address cluster.hpcc.ucr.edu sends you to one of the head nodes (bluejay or skylark):
ssh username@cluster.hpcc.ucr.edu
Replace username with your HPCC username. UCR users log in with password plus Duo; external users use SSH keys. See the HPCC login instructions. A browser option, Open OnDemand, is also available.
Head nodes are for submitting jobs, editing code and very small tests. Do not run BLAST searches on a head node. Submit them to compute nodes through Slurm, as shown below.
Choosing a partition
The HPCC uses the Slurm scheduler. Jobs go to partitions (queues), which are groups of compute nodes. The general-purpose CPU partitions suit most BLAST work:
epyc: AMD EPYC nodes (2021). A good default for BLAST.intel: Intel nodes (2016).batch: AMD nodes (2012).highmem: for very large memory needs. Jobs must request at least 100 GB.short: a mixed set of nodes from the batch, intel and lab partitions, with a 2-hour maximum. Useful for quick tests.
Default memory is 1 GB per job and default walltime is 7 days on epyc, intel and batch, so always request what you need. Per-user and per-lab limits are on the HPCC Queue Policies page, and node details are on the HPCC Managing Jobs and Hardware Details pages. Run slurm_limits on the cluster to see your current limits.
Set the partition with -p (or --partition) on every job:
sbatch -p epyc blast_job.sh
srun -p epyc --pty bash -l
Feature constraints on short. Because short mixes node types, you can ask for a CPU type with --constraint:
# Any Intel node
srun -p short -t 2:00:00 -c 8 --mem 8GB --constraint intel --pty bash -l
# AMD Rome or Milan node
srun -p short -t 2:00:00 -c 8 --mem 8GB --constraint "amd&(rome|milan)" --pty bash -l
Loading BLAST and the NCBI databases
BLAST+ is installed as the ncbi-blast module. Several versions are available; list them first:
module avail ncbi-blast
module load ncbi-blast # default version
# or a specific version, for example:
module load ncbi-blast/2.14.1+
The HPCC keeps copies of common NCBI databases (nr, nt, core_nt and others) through the db-ncbi module (db-swissprot for SwissProt). They are refreshed about every 6 months and kept for about 3 years. If your project needs a fixed database version for longer, copy it to your own storage. See the HPCC Databases page.
module load db-ncbi
Interactive BLAST searches
Interactive sessions are useful for testing commands and running small searches.
- Log in to the cluster.
-
Start a session on a compute node:
srun -p epyc --cpus-per-task=4 --mem=4G --time=00:30:00 --pty bash -l-p epyc: the partition (useshortfor quick tests).--cpus-per-task=4: 4 CPU cores. BLAST can use several threads.--mem=4G: 4 GB of memory. Large databases such as nt need more.--time=00:30:00: 30 minutes of walltime.--pty bash -l: starts a login shell on the compute node.
-
Load the modules:
module load ncbi-blast db-ncbi -
Run the search:
blastn -query input.fasta -db nt -out output.blast -num_threads 4blastn: nucleotide against nucleotide. Useblastp,blastx,tblastnortblastxas your data requires.-query input.fasta: your query sequences in FASTA format.-db nt: the NCBI nt database. Usenrfor proteins, or a custom database.-out output.blast: the results file.-num_threads 4: matches the 4 cores requested withsrun.
- Type
exitto end the session and return to the head node.
Batch BLAST searches (sbatch scripts)
For larger or repeated searches, submit a batch script so the job runs without an open session.
-
Create a script, for example
blast_job.sh:#!/bin/bash -l #SBATCH --job-name=blastn_search #SBATCH --partition=epyc #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --cpus-per-task=16 #SBATCH --mem=32G #SBATCH --time=24:00:00 #SBATCH --output=blastn_job_%j.out # %j is replaced with the job ID # Load BLAST and the NCBI databases module load ncbi-blast db-ncbi # Run the search, one thread per requested core blastn -query input.fasta -db nt -out output.blast -num_threads ${SLURM_CPUS_PER_TASK} # Record which node the job ran on and when it finished hostname dateWhat the directives do:
#!/bin/bash -l: runs a login shell, so themodulecommand is available.--job-name: a name shown bysqueue.--partition=epyc: the partition. Change it as needed.--nodes=1and--ntasks=1: BLAST runs as one multithreaded process on one node.--cpus-per-task=16: cores for BLAST's threads.${SLURM_CPUS_PER_TASK}passes the same number to-num_threads.--mem=32G: memory for the job. Running out of memory ends the job, so allow some headroom.--time=24:00:00: walltime. Estimate from a smaller test run.--output: the file for the job's standard output and error.
You can add
--mail-userand--mail-typeif you want email when the job starts or ends. -
Copy your input files to the cluster if they are not there yet, with
scp,sftpor another method from the HPCC Sharing Data page. Your home directory has a 50 GB quota (see HPCC Recharging Rates); large data belongs in your lab's/bigdataspace. See the HPCC Data Storage page. -
Submit the job from the directory that holds the script and input:
sbatch blast_job.sh -
Monitor the job:
squeue -u $USER --startThis lists your queued and running jobs, with an estimated start time when one is available.
-
Check the results. When the job ends, read
blastn_job_<JOBID>.outfor errors and the node name. The BLAST results (output.blast) are in the same directory.
Tuning BLAST searches
- Threads: set
-num_threadsto the number of cores you request (--cpus-per-task). - Database choice: large general databases such as nt and nr take longer to search than smaller, targeted ones. Use an organism-specific or custom database when it fits the question.
- E-value threshold: a stricter threshold (for example
-evalue 1e-6) returns fewer, more significant hits; a looser one (for example10) returns more hits, including weak ones. -
Check efficiency with
seff: after a job finishes, run:seff JOBID- Low CPU efficiency: confirm
-num_threadsmatches your core request. If it does, try fewer cores next time. - Low memory efficiency: lower
--memnext time, keeping about 20% above the memoryseffreports.
- Low CPU efficiency: confirm
- Test small first: run a subset of your queries to estimate runtime and memory before a large search.
Getting help
- HPCC accounts, the cluster, software and modules: support@hpcc.ucr.edu
- Other Research Computing questions: research-computing@ucr.edu
- UCR Research Computing Slack: https://ucr-research-compute.slack.com/