Most of our work has resulted in scholarly publications. On this page you can review our publications to get an idea about our work.
Realizou-se um levantamento das Plantas Alimentícias Não Convencionais (PANCs) nos municípios de Pedro do Rosário e Zé Doca, Maranhão, Brasil. Com objetivo de… Realizou-se um levantamento das Plantas Alimentícias Não Convencionais (PANCs) nos municípios de Pedro do Rosário e Zé Doca, Maranhão, Brasil. Com objetivo de catalogar a diversidade e potencial nutricional de PANCs, visando contribuir para a valorização destes recursos como alternativa sustentável de alimentação para famílias de baixa renda e para a conservação do conhecimento sobre a agrobiodiversidade local. Foram catalogadas 49 espécies distribuídas em 34 famílias botânicas e 46 gêneros, com destaque para Asteraceae, Amaranthaceae, Lamiaceae e Malvaceae. Espécies como vinagreira (Hibiscus sabdariffa), ora-pro-nóbis (Pereskia bleo), taioba (Xanthosoma taioba) e caruru (Amaranthus spp.) destacaram-se pelo elevado valor nutricional. Os resultados evidenciam o potencial das PANCs como alternativa alimentar sustentável e nutritiva para populações de baixa renda, reforçando a necessidade de políticas públicas e ações educativas para sua valorização. This repository contains the supporting dataset, analysis code, supporting tables, and final figure files for the study entitled “Pressure behavior in a roller pump-driven low-flow VACC mock cir… This repository contains the supporting dataset, analysis code, supporting tables, and final figure files for the study entitled “Pressure behavior in a roller pump-driven low-flow VACC mock circuit with inactive de-airing pump: a descriptive bench characterization under reduced patient-side effective filling pressure conditions”.
The repository includes the replicate-level dataset, a Google Colab-compatible Jupyter notebook, S1 and S2 Tables in DOCX format, and final TIFF figure files. The notebook reproduces the statistical analyses, S1 and S2 Tables, Figs 2–4, and S1 and S2 Figs from the replicate-level dataset. This repository contains the processed data, derived variables, source data for figures and tables, model outputs, and reproducibility materials supporting the study:
"Solar and Wind Expansion and Une… This repository contains the processed data, derived variables, source data for figures and tables, model outputs, and reproducibility materials supporting the study:
"Solar and Wind Expansion and Uneven Fossil-Generation Displacement Across 29 Emerging and Developing Economies”
The study examines whether growth in solar and wind electricity generation was accompanied by absolute reductions in fossil-fuel electricity generation across a balanced panel of 29 emerging and developing economies from 2010 to 2024. The dataset supports descriptive electricity-balance accounting, country and regional comparisons, two-way fixed-effects estimation, small-cluster inference, leave-one-country-out classification, calibration analysis, and robustness checks.
Dataset scope
The harmonized level panel contains:
29 countries;
annual observations from 2010 to 2024;
435 country-year observations.
The annual-change panel contains:
annual first differences from 2011 to 2024;
406 country-year observations;
317 observations with positive solar-wind generation growth;
129 observations with an annual decline in fossil generation.
The dataset includes countries from South Asia, Southeast Asia, the Middle East and North Africa, Sub-Saharan Africa, Latin America, and Central Asia/Eurasia.
Repository contents
The workbook contains:
the harmonized 2010–2024 country-year panel;
the 2011–2024 annual-change panel;
electricity-generation, demand, emissions, trade, and socioeconomic variables;
gross fossil-displacement ratio and clean-growth-coverage indicators;
annual transition classifications;
country-level, regional, and equal-weighted summary results;
fixed-effects and robustness outputs;
wild-cluster-bootstrap and multicollinearity diagnostics;
leave-one-country-out regression estimates;
out-of-fold classification probabilities and predictions;
fold-specific logistic-regression coefficients;
majority-class and sensitivity-model results;
source-consistency checks against Our World in Data;
electricity-balance residual checks;
source data underlying the manuscript figures and tables;
a complete variable dictionary;
a workbook README describing individual worksheets.
A separate Markdown README provides additional information on dataset structure, variables, units, sources, missing-value conventions, and the reproduction workflow.
Main indicators
The dataset includes two principal accounting indicators.
Gross fossil-displacement ratio
GDR is calculated when annual solar-wind generation growth is positive. Positive values indicate contemporaneous fossil-generation decline, while negative values indicate that fossil and solar-wind generation increased together. GDR is a descriptive accounting ratio and should not be interpreted as a causal efficiency estimate.
Clean-growth coverage
CGC is calculated when annual electricity-demand growth is positive. Values of at least one indicate that solar-wind growth equalled or exceeded the annual increase in electricity demand.
The dataset also classifies positive solar-wind-growth observations into four mutually exclusive outcomes:
fossil decline exceeding solar-wind growth;
fossil decline below solar-wind growth;
demand-dominant co-expansion;
fossil co-expansion despite demand coverage.
Data sources
The processed dataset was constructed from three public sources:
Ember Yearly Electricity Data for electricity generation by source, electricity demand, net imports, power-sector emissions, and carbon intensity;
World Bank World Development Indicators for GDP, population, urbanization, industrial structure, manufacturing, and related socioeconomic variables;
Our World in Data Energy Dataset for source-consistency checks.
The repository contains harmonized and derived research data rather than replacing the original public source datasets. Users should consult and cite the original data providers when reusing the underlying source variables.
Units
The main units are:
electricity generation, electricity demand, and net imports: TWh;
power-sector emissions: Mt CO₂;
carbon intensity: g CO₂ kWh⁻¹;
generation shares and socioeconomic shares: percent;
GDP growth and population growth: percent;
GDP per capita: constant US dollars or log-transformed constant US dollars;
GDR and CGC: dimensionless ratios;
fossil-decline outcome: binary indicator.
Analytical outputs
The workbook supports reproduction of:
aggregate electricity trends;
country and regional displacement accounting;
fixed-effects regression estimates;
country-clustered, Driscoll–Kraay, and wild-cluster-bootstrap inference;
influence and leave-one-country-out checks;
prospective structural logistic classification;
ex-post diagnostic logistic classification;
full electricity-balance sensitivity analysis;
ROC, precision-recall, calibration, and Brier-score results;
manuscript figures, tables, and supplementary tables.
The prospective classifier uses lagged electricity-system and socioeconomic variables. The ex-post diagnostic classifier additionally uses contemporaneous normalized changes in electricity demand, solar-wind generation, and hydropower. These diagnostic models identify annual fossil-decline outcomes after electricity-balance changes are observed and should not be interpreted as causal models or true ex-ante forecasts.
Reproducibility
The associated software repository contains:
the Google Colab notebook;
a Python script;
README.md;
requirements.txt;
environment.yml;
exact package versions;
fixed random seed and bootstrap settings;
instructions for obtaining the public input datasets;
procedures for regenerating the processed panels, model outputs, figures, and tables.
The default analysis uses a random seed of 42, 999 wild-cluster-bootstrap replications, and 2,000 country-cluster bootstrap replications for the principal classification analyses.
Related resource
The complete analysis code and computing environment are deposited in a separate linked Zenodo software record. Short paper (programme paper P2) of the CAOS_LDA_HSI series. Latent Dirichlet Allocation (LDA) and its neural variants are increasingly used as interpretable spectral mixture models on hyperspectral i… Short paper (programme paper P2) of the CAOS_LDA_HSI series. Latent Dirichlet Allocation (LDA) and its neural variants are increasingly used as interpretable spectral mixture models on hyperspectral imagery. Their seed and capacity stability has been studied at length; their robustness to the choice of spectral window has not.We introduce a band-mask robustness diagnostic that refits a canonical LDA model under four band-restriction policies (VNIR-only, SWIR-only, atmospheric-water-band removal, top-50 Fisher discriminant) and compares the masked dominant-topic maps against the canonical fit via the adjusted Rand index after a Hungarian topic-id alignment. Applied to six standard hyperspectral scenes (Indian Pines, Salinas, Salinas-A, Pavia University, Kennedy Space Center, Botswana) and five HIDSAG mineral subsets, the diagnostic yields a striking result: topic identities on Salinas-A under SWIR-only restriction recover the canonical assignment with paired ARI = 0.77, while KSC and Botswana paired ARI is ~0.01 under every mask. The diagnostic therefore separates scenes on which any interpretable-spectral-mixture claim is band-robust from scenes on which it is not.We release the full sweep (44 LDA refit attempts, 43 successful) as a public, deterministic artefact set under a permissive licence to support follow-on work.Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper .Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
[in development] Imputation (TODO add link/info)
Architecture
WIP architecture doc
Linked LucidChart
Development
This codebase is in initial development.
Requirements
This service is written in Java 17, and uses Postgres 13.
To run locally, you'll also need:
jq - install with brew install jq
vault - see DSP's setup instructions here
Note that for Step 7, "Create a GitHub Personal Access Token", you'll want to choose
the "Tokens (classic)" option, not the fine-grained access token option.
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 13 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately
Postgres.app https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@13
Tech stack
Java 17 temurin
Postgres 13.1
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access Vault.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/updatePipelineWorkspaceId/{pipelineName}/{workspaceId} to set your pipeline's workspace id. Workspace id can be found through the terra ui workspace dashboard or through the Rawls GET workspace endpoint.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/tsps-sa.json in the Run/Debug configuration Environment Variables.
Running Tests/Linter Locally
Testing
Run ./gradlew service:test to run tests
Linting
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from vault
before running the gradle task.
export SONAR_TOKEN=$(vault read -field=sonar_token secret/secops/ci/sonarcloud/tsps)
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./firecloud/psql-connect.sh dev tsps
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace. (This is not yet fully set up here - to be done in TSPS-107).
See this DSP blog post for more info.
Running the end-to-end tests
The end-to-end test is specified in .github/workflows/run-e2e-tests.yaml. It calls the test script defined
in the dsp-reusable-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli.PLANTAS ALIMENTÍCIAS NÃO CONVENCIONAIS (PANCS) NO MARANHÃO: POTENCIAL NUTRICIONAL E SUSTENTÁVEL
Pressure-characterization VACC mock-circuit study: data and analysis code
Data supporting "Solar and Wind Expansion and Uneven Fossil-Generation Displacement Across 29 Emerging and Developing Economies"
A Band-Mask Robustness Diagnostic for Latent Dirichlet Allocation on Hyperspectral Imagery
github.com/DataBiosphere/terra-scientific-pipelines-service/GatkConcordanceValidation
github.com/DataBiosphere/terra-scientific-pipelines-service/BeagleImputationValidation
github.com/DataBiosphere/terra-scientific-pipelines-service/UpdateVcfDictionaryHeader
github.com/DataBiosphere/terra-scientific-pipelines-service/SubsetVcfByBedFile
github.com/DataBiosphere/terra-scientific-pipelines-service/ReshapeReferencePanel
github.com/DataBiosphere/terra-scientific-pipelines-service/CreateImputationRefPanelBeagle
On Losses, Pauses, Jumps and the Wideband E-Model – IEEE Xplore Document
There is an increasing interest in upgrading the EModel, a parametric tool for speech quality estimation, to the wideband and super-wideband contexts. The
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles – IEEE Xplore Document
Contemporary models of Unmanned Aerial Vehicles (UAVs) are largely developed using simulators. In a typical scheme, a flight simulator is dovetailed with a
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles
Simulators as Drivers of Cutting Edge Research – IEEE Xplore Document
Undertaking engineering research can be compounding for beginning graduate students and thwarting even for seasoned researchers. With a wealth of academic
Simulators as Drivers of Cutting Edge Research
Evolutionary speech quality estimation in VoIP
A Methodology for Deriving VoIP Equipment Impairment Factors for a Mixed NB/WB Context
Real-Time, Non-intrusive Speech Quality Estimation: A Signal-Based Mod
Real-Time, Non-intrusive Evaluation of VoIP
VoIP speech quality estimation in a mixed context with genetic programming
An Evolutionary Approach to Speech Quality Estimation
Real-Time Non-Intrusive VoIP Evaluation Using Second Generation Network Processor
Non-intrusive quality evaluation of VoIP using genetic programming
