CephRDS supports the Amazon S3 API, so you can work with it from Python using the standard AWS SDK, boto3. For requesting storage and keys, see KB013: Connecting to CephRDS.
Prerequisites
- Python 3
- The
boto3library (pip install boto3) - Your CephRDS Access Key ID and Secret Access Key
- The name of your bucket
Two settings that matter
CephRDS runs on campus, not in AWS, so two settings differ from a standard AWS script:
- Endpoint: set
endpoint_urltohttps://rds.ucr.edu. - Path-style addressing: set
addressing_styletopath. You do not needregion_name.
Keep your keys out of your code
Do not type your keys into a script. Set them as environment variables in your shell (or load them from a file that is never committed to version control):
export CEPHRDS_ACCESS_KEY="your-access-key-id"
export CEPHRDS_SECRET_KEY="your-secret-access-key"
Code example
This script lists the objects in a bucket, uploads a file and downloads it again. Replace my-lab-bucket with your bucket name.
import os
import boto3
from botocore.client import Config
from botocore.exceptions import ClientError
ENDPOINT_URL = "https://rds.ucr.edu"
BUCKET_NAME = "my-lab-bucket"
s3 = boto3.client(
"s3",
endpoint_url=ENDPOINT_URL,
aws_access_key_id=os.environ["CEPHRDS_ACCESS_KEY"],
aws_secret_access_key=os.environ["CEPHRDS_SECRET_KEY"],
config=Config(s3={"addressing_style": "path"}),
)
# List objects (the paginator handles buckets with more than 1,000 objects)
paginator = s3.get_paginator("list_objects_v2")
for page in paginator.paginate(Bucket=BUCKET_NAME, Prefix="data/"):
for obj in page.get("Contents", []):
print(obj["Key"], obj["Size"])
# Upload a file
try:
s3.upload_file("local_data.csv", BUCKET_NAME, "data/local_data.csv")
print("Upload complete.")
except ClientError as err:
print("Upload failed:", err)
# Download it again
s3.download_file(BUCKET_NAME, "data/local_data.csv", "downloaded_data.csv")
print("Download complete.")
upload_file and download_file split large files into parts and transfer them in parallel automatically.
Troubleshooting
AccessDeniedorInvalidAccessKeyId: check that the environment variables hold the right keys and that your key has access to that bucket.NoSuchBucket: check the bucket name spelling. Bucket names are case-sensitive.- Connection timeouts: CephRDS is reachable only from the campus network. Off campus, connect to the UCR campus VPN (Cisco Secure Client) first. Contact research-computing@ucr.edu if the problem continues.
Security note
Never commit keys to GitHub or another repository. If a key may have been exposed, contact research-computing@ucr.edu so it can be replaced.