---
language: "en"
---
# ARK Portal Docs

## Pinned articles

*

  ### [Getting Started](https://help.arkportal.org/help/getting-started.md)

  Welcome to the ARK Portal docs site! This site exists to help orient you to the portal and use it with ease to make the most of your experience. What's here fo...
*

  ### [About the Portal](https://help.arkportal.org/help/about.md)

  The ARK Portal is a public data repository that stores and shares data and research knowledge generated by a network of research teams focused on autoimmune an...

## Documentation

*

  ### [Getting Started](https://help.arkportal.org/help/getting-started.md)

### [About the Portal](https://help.arkportal.org/help/about.md)

*

  ### [Navigating the Portal](https://help.arkportal.org/help/navigating-the-portal.md)

*

  ### [Access \& Attribution](https://help.arkportal.org/help/data-use-certificate.md)

*

  ### [About the Data](https://help.arkportal.org/help/accessing-data.md)

*

  ### [ARK Portal Data Standards](https://help.arkportal.org/help/ark-portal-data-standards.md)

*

  ### [Sending Data to Cloud Analytical Platforms](https://help.arkportal.org/help/sending-data-to-cloud-analytical-platforms.md)

---
language: "en"
---
# About the Portal

The ARK Portal is a public data repository that stores and shares data and research knowledge generated by a network of research teams focused on autoimmune and immune-related diseases.

The ARK Portal is funded by the [National Institute of Arthritis and Musculoskeletal and Skin Diseases (NIAMS)](https://www.niams.nih.gov/)National Institute of Arthritis and Musculoskeletal and Skin Diseases (NIAMS), and the [National Institute of Allergy and Infectious Diseases (NIAID)](https://www.niaid.nih.gov/)National Institute of Allergy and Infectious Diseases (NIAID). It is developed and maintained by Sage Bionetworks. If you have questions, suggestions, or feedback about the ARK Portal, please contact us [here](https://sagebionetworks.jira.com/servicedesk/customer/portal/11).

## History

The ARK Portal was originally established as part of the [Accelerating Medicines Partnership® Rheumatoid Arthritis and Systemic Lupus Erythematosus (AMP® RA/SLE) Program](https://www.niams.nih.gov/grants-funding/funded-research/accelerating-medicines/RA-SLE) which was initiated in 2014 between the National Institutes of Health (NIH) and several nonprofit organizations and pharmaceutical companies. Together, this public‐private partnership has funded an open, pre‐competitive discovery effort with the goal of working collaboratively to deepen the understanding of autoimmune diseases through a focus on rheumatoid arthritis (RA) and systemic lupus erythematosus (SLE).

## ARK Portal ↔︎ Synapse

At this point, you know what the ARK Portal is, and you may have come across the term Synapse - but how do they fit together? Let's break this down:  

|----------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Sage Bionetworks** | First, there's Sage Bionetworks - a name you may or may not have come across. While Sage is not a tool you'll be using, you should know who we are: a non-profit organization based out of Seattle, Washington. Sage is dedicated to promoting and advancing open science, as well as engaging patients in the research process. Sage acts as the Data Coordinating Center (DCC) for several different portals, including the ARK Portal. The scientists, developers, and designers that built the tools you're using are all employed by Sage. You can learn more about Sage Bionetworks and its initiatives [here](https://sagebionetworks.org/). |
| **Synapse**          | [Synapse](https://www.synapse.org/Home:x) is a collaborative research platform that helps you and your team share, organize, and discuss your scientific research. It allows users to upload, store, curate, and share data privately and publicly.                                                                                                                                                                                                                                                                                                                                                                                                 |
| **ARK Portal**       | ARK Portal is a custom website that disseminates data generated by a network of research teams working collaboratively to deepen the understanding of Arthritis and Autoimmune and Related Diseases. It was established by the National Institute of Arthritis and Musculoskeletal and Skin Diseases (NIAMS) and includes data from the Accelerating Medicines Partnership® (AMP®).                                                                                                                                                                                                                                                                 |

---
language: "en"
---
# About Data Sharing

## Why share data?

Data sharing is central to open science. When you share your data, you show your support for the future of science---for its openness, its reproducibility, and its longevity. Responsible and open data sharing allows you to demonstrate the rigor and reliability of your work to others, and by doing so, you invite them to review, reproduce, and reuse your materials, potentially advancing new discoveries. For biomedical data, data sharing can lead to new treatments, therapies, and even cures, improving patient outcomes. Data sharing also contributes to the public good---it can build trust in science and increase access to scientific knowledge.

Beyond philosophical reasons, data sharing is also required by many funding organizations, including the [U.S. National Institutes of Health (NIH)](https://grants.nih.gov/policy/sharing.htm), and many other governing bodies, foundations, and journal publications.

## What is data sharing?

In general, data sharing means making data available to others in a responsible way. This can include the following additional steps:

* providing information about the data such as abstracts, code, and protocols

* ensuring access controls are in place for sensitive data

* embargoing data for a specified period

* de-identifying data

* adding information about the data (called metadata or annotations) to enable data discovery and querying

These days, many researchers share data via online repositories equipped with features such as file storage, data annotation tools, security protocols, and search functionality. Some well-known scientific data repositories include GEO, cBioPortal, figshare, etc. At Sage, data are stored and curated in a platform called [Synapse](https://sagebionetworks.org/platform/synapse), and made explorable through data portals.

## How do I share my data?

While the exact process has some variations across our data portals and repositories, data sharing at Sage consists of the following steps:

1. **Get involved in a community.** First, make contact with the appropriate community for your data. In some cases, this might involve contacting the portal maintainers, or it might involve joining a consortium, as well as securing funding and setting up a data sharing plan with a specific organization.

2. **Prepare data.** Before generating data, review your chosen community's onboarding materials, which could include documentation, submission forms, webinars, one-on-one meetings, and other resources. Gather all supplemental information, and ensure you've met the data sharing requirements.

3. **Deposit data** and add information about the data. This step includes uploading and annotating data, as well as providing any supplementary information needed to understand and curate the data. This step may also include data quality checks and metadata validation. Consortia and/or funders may have additional requirements, such as milestone reports.

4. **Determine data access controls.** At Sage, we use the term [data governance](https://docs.synapse.org/synapse-docs/synapse-governance) to refer to the practice of determining how data should be shared. This stage encompasses data licensing, as well as deciding how the data should be accessed, and by whom. Many datasets on our portals are unrestricted and open to the public, but some data require access controls, such as data use agreements and institutional review board approval.

5. **Share data.** After the above steps are complete, you're finally ready to share your data with others, whether fully open or with restrictions. Often, this step occurs some time after the above processes, once a publication embargo period lifts.

6. **Access data.** Once your data is shared, it's now accessible to others, including yourself and your colleagues. Because of the steps above, that data will be [FAIR](https://www.go-fair.org/fair-principles/) --- a standard representing findable, accessible, interoperable, and reusable data. FAIR data are discoverable to users through precise metadata, understandable in terms of how the data can be used, machine-readable to enable computational analysis, and ultimately, fit for reuse.

---
language: "en"
---
# About the Data

This is likely the reason you're visiting the ARK Portal in the first place---to access data. The portal offers helpful filtering tools to help you find data of interest. All of the data and resources uploaded into the portal are associated with metadata, so they can be easily used to help query the list of resources in each page.

## Get started exploring data

### [Dataset Collections](https://arkportal.synapse.org/Explore/Collections)

Data in the portal is grouped into Datasets Collections of which there are two flavors: Experimental Data and Publications. Experimental Data are datasets related to a specific project, assay, or disease. Publications are datasets used in specific manuscripts. A Publication dataset may have a subset of files, or files that span multiple Experimental Data datasets. Use the dataset annotations to query for specific assays, tissue types, and which program and project the data is associated with. See also the short description in the dataset table.

### Individual Datasets

Each dataset has three components:

* **Description**: This is a summary describing the dataset

* **Acknowledgement**: A condition of use for any data obtained through the portal is the attribution of data contributors and funders. Use the acknowledgement statement as listed on the dataset page.

* **Files:** Find and download the experimental data and metadata files. The following metadata is provided

  * *Phenotypes*: This is a list of all individuals in a specific Program. The file is linked to all datasets relevant to that program, and may list more individuals than what is relevant to a specific dataset

  * *Biospecimen*: This is information on the specimens used in the assays. The specimen ID is linked to the individual ID in the Phenotype file

  * *Quality Assessments/Protocols*: Where available, additional information about QC and protocols are provided

### [**All Data**](https://arkportal.synapse.org/Explore/All%20Data)

This is a table of all the files in the portal. It can be queried in the same way as files within specific Datasets. Note that datafiles are annotated with the Experimental Data datasets they are part of.

## Downloading Data

See the Download Options in the upper right of the file tables found in Datasets and the All Data view.  
![Screen Shot 2022-12-11 at 1.12.19 PM.png](https://help.arkportal.org/__attachments/a_2076f8ddae4683e2ae41b5ecb1309bc41a304815dba69a444576c3a1e740a3f3/Screen%20Shot%202022-12-11%20at%201.12.19%20PM.png?cb=f433d8e1eefe92e2076d52e2d9c487e8)

You have 3 options:

* **Export Table**: Download a csv or tsv file with filenames and associated annotations

* **Add To Download Cart**: Add files to a cart to modify or download later. Note, this is a great way to determine the size of the data you have selected.

* **Programmatic Options**: See Command Line, R, or Python code for downloading files directly.

Learn the basic commands for[installing the Synapse API Clients](https://help.synapse.org/docs/Installing-Synapse-API-Clients.1985249668.html).

---
language: "en"
---
# ARK Portal Data Standards

## What are data standards?

Data standards are a set of rules that define how data is recorded, described, and shared. They help ensure that data is consistent, accurate, and comply with FAIR data practices.

This document, and any accompanying pages, describe standards for common assay data types found in the ARK Portal and outlines expectations for data contributions.

## When to apply standards?

Please be familiar with ARK Portal data standards requirements and plan to have your data files and metadata tables conform **before uploading to Synapse**. This means that files should be named and organized according to the conventions outlined here.

## Olink data standards

| **Folder Name**  |                             **Data Description**                             | **Expectation** |                                **File Formats**                                | **Data Level\*** |
|------------------|------------------------------------------------------------------------------|-----------------|--------------------------------------------------------------------------------|------------------|
| `raw_data`       | raw protein abundance measurements                                           | required        | `parquet` or `CSV`, one per plate                                              | 1                |
| `processed_data` | processed aggregated data                                                    | required        | single, aggregated `parquet` or `CSV` consisting of all finalized data points. | 2+               |
| `metadata`       | target panel^a^, standardized metadata\*\*, user-defined metadata (optional) | required        | a tabular file formats (e.g., `csv`, `xlsx`)                                   | N/A              |

At a minimum, all Olink data contributions to the ARK Portal should include all raw data in `parquet` or `CSV` format, one file per plate profiled, and a final aggregated data object in `parquet` or `CSV` format, that has been used to derive research findings reported in publications. This finalized data object should include any normalization, integration, transforms, etc. that have been applied to the raw data.

^a^For more details about target panels see [Supplemental Standards: Target Panel](https://help.arkportal.org/help/ark-portal-data-standards.md#Target-Panel).

\*<https://ark-portal.github.io/data_model/docs/attributes/dataLevel.html>

\*\*Data contributors are also required to provide standardized metadata conforming to the ARK Portal Data model. Templates will be provided to guide the collection of these critical metadata.

## scRNA-seq/sn-RNA-seq data standards

|             **Folder Name**              |                    **Data Description**                     | **Expectation** |                                                               **File Formats** ^**a**^                                                               | **Data Level\*** |
|------------------------------------------|-------------------------------------------------------------|-----------------|------------------------------------------------------------------------------------------------------------------------------------------------------|------------------|
| `fastq_files`                            | raw fastq files                                             | required        | gzipped fastq files                                                                                                                                  | 1                |
| `bam_files`                              | read alignment files                                        | optional        | bam or cram files                                                                                                                                    | 2                |
| `CellRanger_counts` or `raw_gene_counts` | raw gene counts                                             | preferred       | compressed tar archive (e.g., `.tgz`) or `.h5`file - these are readily available after running Cell Ranger counts on 10x Genomics sc/sn-RNA-seq data | 3                |
| `processed_data`                         | processed aggregated gene counts                            | preferred       | AnnData (as an `h5ad`file) or SeuratObj (as an `Rds` or similarly binary compressed R-compatible file)                                               | 4+               |
| metadata                                 | standardized metadata\*\*, user-defined metadata (optional) |                 |                                                                                                                                                      |                  |

^a^File name conventions are described at [Supplemental Standards: File Names](https://help.arkportal.org/help/ark-portal-data-standards.md#File-Names)

\*<https://ark-portal.github.io/data_model/docs/attributes/dataLevel.html>

^b^For data processed by 10x Genomics Cell Ranger software, if contributors wish to upload the MEX output they should first convert either the `raw_feature_bc_matrix/` or `filtered_feature_bc_matrix/` folder to a gzip compressed tar archive.

At a minimum, all sc/snRNA-seq data contributions to the ARK Portal should include the raw fastq files. We additionally request that contributors provide raw gene counts, either aggregated or split by library/sample, and a final aggregated AnnData (as an `h5ad`file) or SeuratObj (as an `Rds` or similarly binary compressed R-compatible file) of prepared counts data that has been used to derive research findings reported in publications. This finalized data object should include any normalization, integration, transforms, etc. that have been applied to the gene counts along with a critical cell metadata as outlined at [Supplemental Standards: Single-cell Metadata](https://help.arkportal.org/help/ark-portal-data-standards.md#Single-cell-Metadata).

\*<https://ark-portal.github.io/data_model/docs/attributes/dataLevel.html>

\*\*Data contributors are also required to provide standardized metadata conforming to the ARK Portal Data model. Templates will be provided to guide the collection of these critical metadata.

## CITE-seq data standards

|           **Folder Name**           |                             **Data Description**                             | **Expectation** |                                                            **File Formats** ^**a**^                                                             | **Data Level\*** |
|-------------------------------------|------------------------------------------------------------------------------|-----------------|-------------------------------------------------------------------------------------------------------------------------------------------------|------------------|
| `fastq_files/GEX_fastq`             | scRNA-seq raw fastq files                                                    | required        | gzipped fastq files                                                                                                                             | 1                |
| `fastq_files/feature_barcode_fastq` | feature barcode raw fastq files                                              | required        | gzipped fastq files                                                                                                                             | 1                |
| `CellRanger_counts` or `raw_counts` | raw gene and proteins counts                                                 | preferred       | compressed tar archive (e.g., `.tgz`) or `.h5`file - these are readily available after running Cell Ranger counts on 10x Genomics-derived data. | 3                |
| `processed_data`                    | processed aggregated gene counts                                             | preferred       | AnnData (as an `h5ad`file) or SeuratObj (as an `Rds` or similarly binary compressed R-compatible file)                                          | 4+               |
| `metadata`                          | target panel^b^, standardized metadata\*\*, user-defined metadata (optional) | required        | a tabular file formats (e.g., `csv`, `xlsx`)                                                                                                    | N/A              |

^a^File name conventions are described at [Supplemental Standards: File Names](https://help.arkportal.org/help/ark-portal-data-standards.md#File-Names)

^b^For more details about target panels see [Supplemental Standards: Target Panel](https://help.arkportal.org/help/ark-portal-data-standards.md#Target-Panel).

CITE-seq is a multi-modal data type that simultaneously profiles transcript and target protein abundance at single-cell resolution. Protein targets are profiled by sequencing barcodes contained within oligonucleotides conjugated to antibodies that bind to proteins of interest - where each barcode is uniquely associated with a specific protein. These antibody-derived barcodes are sometimes referred to as antibody-derived tags (ADT) or as feature barcodes. The ARK Portal used the latter terminology.

The protein abundance libraries are created and sequenced as distinct libraries from the scRNA-seq libraries and are treated as a distinct `assay` type within the ARK Portal data model. Specifically, the ARK Portal classifies these libraries under the 'feature barcode sequencing' assay. This is distinct from other antibody-derived barcode methods like hash-tag oligos that are used to demultiplex libraries made of pooled cell suspensions and which do not target specific proteins for the purpose of quantifying protein abundance.

At a minimum, all CITE-seq data contributions to the ARK Portal should include the raw fastq files for both the scRNA-seq libraries and the feature barcode sequencing libraries. We additionally request that contributors provide raw gene and protein counts, either aggregated or split by library/sample, and a final aggregated AnnData (as an `h5ad`file) or SeuratObj (as an `Rds` or similarly binary compressed R-compatible file) of prepared counts data that has been used to derive research findings reported in publications. This finalized data object should include any normalization, integration, transforms, etc. that have been applied to the gene counts along with a critical cell metadata as outlined at [Supplemental Standards: Single-cell Metadata](https://help.arkportal.org/help/ark-portal-data-standards.md#Single-cell-Metadata).

\*\*Data contributors are also required to provide standardized metadata conforming to the ARK Portal Data model. Templates will be provided to guide the collection of these critical metadata.

## Supplemental Standards

### File Names

#### Single specimen files

In the tables above, File Formats indicates the expected format and extension of data files. Here we describe conventions regarding information to include in your file names. The examples below use fastq files from sequencing based experiments to demonstrate ARK Portal file name conventions, but the convention is applicable to many other file types as well.
> TL;DR - if a data file contains data pertaining to a single specimen then the `biospecimenID` should be included in the file name. If a file contains pooled data, particularly raw data files, then the file name should include the corresponding string/variable corresponding to that pool. This variable will differ depending on the assay. For example, pooled sequencing libraries should use the `libraryID`. Olink level 1 files should include the `plateID`, barcoded and multiplexed FCS files should include the `sampleProcessingBatch` or `dataCollectionBatch`, etc.

For single-specimen libraries, i.e., libraries consisting of only a single sample, fastq files should include the [biospecimenID](https://ark-portal.github.io/data_model/docs/attributes/biospecimenID.html) of that sample. All fastq files should include the read (R) or index (I) label and note the flow cell lane that the library was sequenced on as this is necessary for indicating libraries that were sequenced across multiple lanes, e.g.,
> RASLE_000001_L001_R1_001.fastq.gz

Where `RASLE_000001` is the biospecimenID, `L001` is the flow cell lane, and `R1`is the read of the library fragment sequenced in the file. More details on Illumina fastq file naming convention is available at [BaseSpace Naming Convention](https://support.illumina.com/help/BaseSpace_Sequence_Hub_OLH_009008_2/Source/Informatics/BS/NamingConvention_FASTQ-files-swBS.htm).
> While not common, there are some scenarios in which a library may be sequenced across multiple flow cells. In these cases it is important to create and assign distinct batch labels that distinguish between these runs. This can be the flow cell ID, a simple letter or number code, etc. This label should then be included in the fastq file name and will also be captured in the corresponding ARK Portal Assay Metadata Template.

For multi-modal assays where multiple libraries are derived from the same biospecimen, e.g., CITE-seq, the fastq file names should follow the above examples with the addition of a short abbreviation distinguishing libraries for each assay type. The table below outlines the different abbreviations that should be used to distinguish between different libraries in a multi-modal experiment:

Where `abbreviation`is appended to the beginning of the fastq file name  

| **Abbreviation** |                 **Assay**                  |              **Example**               |
|------------------|--------------------------------------------|----------------------------------------|
| GEX              | scRNA-seq (GEX = **g** ene **ex**pression) | GEX_RASLE_000002_L001_R1_001.fastq.gz  |
| FB               | **f** eature **b**arcode sequencing        | FB_RASLE_000002_L001_R1_001.fastq.gz   |
| VDJ              | V(D)J sequencing (TCR + BCR)               | VDJ_RASLE_000002_L001_R1_001.fastq.gz  |
| TCR              | V(D)J sequencing (TCR only)                | TCR_RASLE_000002_L001_R1_001.fastq.gz  |
| BCR              | V(D)J sequencing (BCR only)                | BCR_RASLE_000002_L001_R1_001.fastq.gz  |
| ATAC             | ATAC-seq                                   | ATAC_RASLE_000002_L001_R1_001.fastq.gz |

#### Multispecimen libraries

For multispecimen libraries the file names should use the [libraryID](https://ark-portal.github.io/data_model/docs/attributes/libraryID.html), [plateID](https://ark-portal.github.io/data_model/docs/attributes/plateID.html), [slideID](https://ark-portal.github.io/data_model/docs/attributes/slideID.html), etc. in place of the biospecimenID.

### Single-cell Metadata

Cell-level metadata is a critical component of single-cell and single-nucleus datasets. Contributors are asked to include a minimal set of standardized cell-metadata to streamline data reuse and support a more harmonized metadata infrastructure for ARK Portal data:

* [biospecimenID](https://ark-portal.github.io/data_model/docs/attributes/biospecimenID.html) - by including this ID, each cell will be connected to the associated metadata collected via ARK Biospecimen metadata templates.

* [cellOntologyID](https://ark-portal.github.io/data_model/docs/attributes/cellOntologyID.html) - (if cell type annotations are included) corresponding to the predicted cell type. The Cell Ontology (CL) is a structured, controlled vocabulary for cell types and provides a set of unique identifiers for specifying cell types.You can explore the cell ontology at <https://bioportal.bioontology.org/ontologies/CL?p=summary>.

Any additional cell-level metadata fields should be defined in an accompanying dictionary to clearly document what information/data is also captured in these tables.

To learn more about "FAIRification" efforts for single-cell data please visit <https://sc-fair.org/> and <https://github.com/chanzuckerberg/single-cell-curation/tree/main>.

### Target Panel

To ensure transparency and to support future ARK Portal developments certain datatypes will require the submission of a "target panel" file that details all the molecules (e.g., proteins) targeted and profiled in an assay. At a minimum this should be a tabular file format (e.g., `csv`, `xlsx`) (PDFs will not be accepted) that lists all targets using established unique identifiers, for example [Uniprot IDs](https://www.uniprot.org/) or [HGNC](https://www.genenames.org/) approved gene symbol. This file is often be readily available from manufacturers of pre-defined kits.

### Multiplexed Experiments

Multiplexing samples can be advantageous for increasing throughput and cutting costs. Several methods exist depending on the particular assay. Several assay metadata templates require the 'demultiplexMethod' field if your raw data is multiplexed - in these scenarios it is important to collect additional metadata that facilitates demultiplexing of this data.

Below are additional standards for demultiplexing metadata.  

|     **demultiplexMethod**      |        **assay(s)**         |                                                              **additional data/metadata**                                                              |
|--------------------------------|-----------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------|
| genotypic demux                | sequencing-based            | vcf file used for genotypic demux, output of demux software run, code documenting demux run                                                            |
| hashtag oligos, sample barcode | sequencing-based, cytometry | a plain-text, delimited file indicating the barcode-to-biospecimenID mappings                                                                          |
| slide array map                | spatial imaging             | for each slide, an image or grid-based document clearly indicating the biospecimenID corresponding to specific tissue sections or regions of interest. |

## Resources

### ARK Portal Data Model and Dictionary

The ARK Portal Data Model Dictionary is hosted online at <https://ark-portal.github.io/data_model/>. This site is built directly from the data model files hosted in a public repository on GitHub at <https://github.com/ARK-Portal/data_model/tree/main>. All are welcome to review, submitt issues, and contribute to the ARK Portal data model.

---
language: "en"
---
# Access \& Attribution

The data available in the ARK Portal would not be possible without the participation of research volunteers and the contribution of data by collaborating researchers. Data are made available as Open- or Controlled- Access, where individual-level Human data is Controlled Access. Access to any data requires the data available in the ARK Portal would not be possible without the participation of research volunteers and the contribution of data by collaborating researchers. Data are made available as Open- or Controlled- Access, where individual-level Human data is Controlled Access. Access to any data requires the [registration of a Synapse account](https://www.synapse.org/#!RegisterAccount:0) and agreement to the [Synapse Terms and Conditions of Use](https://help.synapse.org/docs/Synapse-Governance.2004255211.html#SynapseGovernance-SynapseTermsandConditionsofUse), and the [Attribution of funders and data contributors](https://help.arkportal.org/help/data-use-certificate.md#Attribution).

## Data Access

### Open Access

Aggregate data which are data combined from several individuals is treated as Open Access.

### Controlled Access

While all data shared through the ARK Portal falls under the principles of Open Data, individual level (any file that has values for an individual), human data is Controlled Access Data and requires the submission of a Data Use Certificate (DUC). You need to update the DUC to add any new team members, and you must renew the DUC each year as necessary. All instructions are found below.

#### To submit a new DUC:

1. Write a description of your proposed research use, known as the Intended Data Use statement (IDU). The IDU should be 500 words maximum (in English) and should include the following:

   * **Scientific Purpose**: What are the main scientific goals of your research and hypotheses? How does the research serve the public good or advance scientific knowledge?

   * **Data Analysis Plan and AI Use**: How will you analyze the data? Describe your planned analytical methods and statistical approaches. Specify whether you will develop new AI/machine learning models, refine existing models, or apply pre-trained published models.

   * **Data Sets to be Used**: What specific studies/datasets are you requesting? Justify why these data are necessary and critical to meet your research objectives.

   * **Computing Environment:** Where will you store and analyze the data (i.e. secure enclaves, virtual private cloud environments, or local systems)? What security measures (e.g., encryption, multi-factor authentication, and compliance certifications such as NIST, FISMA, or ISO/IEC) are in place to keep the data secure?

   * **Data and Safety Monitoring Plan:** What methods will you use to protect individual privacy, prevent model memorization of individual data, and limit data leakage?

   * **Outputs vetting strategy**: What research outputs will you generate (models, statistical results, publications, visualizations) and their sensitivity level? How will you ensure that the outputs cannot be used to re-identify individuals through techniques such as reverse-engineering, small cell sizes, or variable combinations?

   * **Sharing Plan**: Do you plan to use the models or research outputs for commercial purposes now or in the future? Clarify how data use terms will be preserved in downstream uses. (Note that models trained on the controlled-access data and their parameters are considered data derivatives that can only be shared under controlled access through the original data repository.)

2. [Click here](https://www.synapse.org/Synapse:syn61813193) to access the DUC form, which you can download and print. Note that no changes are allowed to the text of the DUC.

3. Using the DUC, gather Synapse usernames and signatures from all the team members at your institution who will need access to the data.

4. Have an official at your institution review and sign the DUC. A signing official is someone from your organization who can speak to your affiliation and has good standing within the organization. You cannot serve as your own signing official.

5. Scan the completed DUC.

6. [Visit the ARK Portal DUC page](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY) and click "Request Access" to complete the data access request process, which includes submitting your IDU from step 1, and uploading the signed DUC form.

Expected turnaround is within one week of DUC submission. Once approved, data may be downloaded and accessed for one year.

#### To add new team members to an existing DUC:

\*Note that only the submitter of the DUC will have the ability to add new team members to an existing DUC.

Please be aware that if you submit a DUC request on behalf of your research group, you will be responsible for submitting all changes, renewals, and progress reports on behalf of that group. This responsibility cannot be transferred from the submitter to another member of the research group. Furthermore, Sage Bionetworks cannot submit any updates or progress reports on behalf of the submitter. If the submitter is unable to make the required updates, renewals, or progress reports, a new DUC request must be completed by another individual in the group. This individual will then maintain the responsibilities of the submitter.

1. [Visit the ARK Portal DUC page](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY).

2. Click **Update Request**.

3. Download the existing DUC form

4. Check the document version of your existing DUC form to ensure that it matches the most current DUC template document version; the most current version of the DUC template can be found [here](https://www.synapse.org/Synapse:syn61813193).

5. If your existing DUC is using an older version of the document, you must complete and submit the latest version of the DUC; you cannot amend an outdated version of the DUC to add new team members.

6. Gather Synapse usernames and signatures from all the new team members at your institution who will need access to the data.

7. Scan the completed DUC.

8. Complete the online DUC request process by [visiting the page with conditions for AMP RA.SLE data use](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY).

#### To complete the annual renewal:

The request submitter will receive two email reminders to renew data access before the expiration date. Follow the instructions in the reminder email or click on the link provided to the [ARK Portal DUC](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY). Click **Update Request** to begin the access renewal. Please ensure you update the following in your renewal request:

1. Remove any collaborators who no longer need data access from both the Synapse access request and the DUC.

2. Add any new collaborators to both the access request and the DUC. Ensure they have signed the DUC.

3. Update your Intended Data Use statement to reflect your progress since your last access request.

4. Ensure you are submitting the correct version of the DUC. The most current DUC template will be linked in the [request area](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY). It is ok to resubmit your DUC from the previous year as long as the document version is current and the requestors list is up to date.

You will receive an email within two weeks of request submission indicating whether your renewal has been approved or rejected. Once approved, data access will renew for everyone listed within the submission.

#### To request additional DUC support:

If you still have questions or issues related to a DUC, please contact our Access \& Compliance Team (ACT). You can use any of the following methods to contact the ACT:

* Use the [Access \& Compliance Team portal](https://sagebionetworks.jira.com/servicedesk/customer/portal/8)

* Email the team at [act@synapse.org](mailto:act@synapse.org)

## Attributions and Acknowledgments

All data use must be acknowledged in publications. Acknowledgment statements will vary by ARK Portal Program. You agree to this through the Clickwrap terms included as part of the ARK data access request process.

### AMP® RA/SLE

[AMP RA.SLE data use request](https://www.synapse.org/#!AccessRequirements:ID=syn38805284&TYPE=ENTITY)

"The results published here are in whole or in part based on data obtained from the ARK Portal ( <https://arkportal.synapse.org/>). This work was supported by the Accelerating Medicines Partnership® Rheumatoid Arthritis and Systemic Lupus Erythematosus (AMP® RA/SLE) Program. AMP® is a public-private partnership (AbbVie Inc., Arthritis Foundation, Bristol-Myers Squibb Company, Foundation for the National Institutes of Health, GlaxoSmithKline, Janssen Research and Development, LLC, Lupus Foundation of America, Lupus Research Alliance, Merck \& Co., Inc., National Institute of Allergy and Infectious Diseases, National Institute of Arthritis and Musculoskeletal and Skin Diseases, Pfizer Inc., Rheumatology Research Foundation, Sanofi and Takeda Pharmaceuticals International, Inc.) created to develop new ways of identifying and validating promising biological targets for diagnostics and drug development Funding was provided through grants from the National Institutes of Health (UH2-AR067676, UH2-AR067677, UH2-AR067679, UH2-AR067681, UH2-AR067685, UH2- AR067688, UH2-AR067689, UH2-AR067690, UH2-AR067691, UH2-AR067694, and UM2- AR067678)."

### UMass V-CoRT

"The results published here are in whole or in part based on data obtained from the ARK Portal (<https://arkportal.synapse.org/)> contributed by Dr. Manuel Garber and John Harris (University of Massachusetts) from `<dataset DOI>` as described in publication `<publication DOI>`. The data generation was funded by NIH grant 5P50AR080593 and NIAID grant 5U01AI176310."

**Please be sure to cite both the associated publication DOI and dataset DOI.**

---
language: "en"
---
# Getting Started

Welcome to the ARK Portal docs site!

This site exists to help orient you to the portal and use it with ease to make the most of your experience.

What's here for you?

* Read about the background of the portal and how it fits into the bigger data-sharing picture

* Get familiar with the basic structure and components of the portal and how data is stored

* Learn how to navigate the portal to explore and access data

## What do you need to get started?

So, before you dive in, here's a breakdown of what you need to get started using the portal:

### Exploring data

Thanks to the [open science](https://help.arkportal.org/help/about-data-sharing) principles that the portal is built on, anyone can browse its content! Nothing needed to explore---just go right ahead (and remember to use this site for help).

### Downloading data

If you want to download data, you will need to [register for a Synapse account](https://www.synapse.org/#!RegisterAccount:0) and, in doing so, agree to the [Synapse Terms and Conditions of Use](https://help.synapse.org/docs/Synapse-Governance.2004255211.html#SynapseGovernance-SynapseTermsandConditionsofUse).

Synapse is the backend and datastore of the ARK Portal. You can read more[about Synapse here](https://www.synapse.org).

Some data on the portal is controlled access data and requires an additional submission of a [Data Use Certificate (DUC)](https://help.arkportal.org/help/data-use-certificate). We typically try to complete the review within a week, but it's most efficient if you submit the certificate as soon as possible.

**Please note - manually downloading data through the Synapse website is capped at 5 GB or 100 files per download.** Larger downloads can be performed using the Synapse Client. More details are available at <https://docs.synapse.org/synapse-docs/api-clients-and-documentation>.

### Using data

Data on the portal is available for [General Research Use (GRU)](https://www.genome.gov/about-nhgri/Policies-Guidance/Genomic-Data-Sharing/frequently-asked-questions/NHGRI-specific), with no embargo on publications. If you plan to use any data downloaded from the portal in any publications, whether it is open access data or controlled access data, you must include an official acknowledgement of the data contributors. Learn how to find and use pre-written acknowledgement statements [here](https://help.arkportal.org/help/data-use-certificate#DataUse&Acknowledgement-Acknowledgement).

---
language: "en"
---
# Navigating the Portal

Let's go over the basic structure of the portal so you can navigate it successfully. Landing on the portal, you'll find a menu along the top with different tabs to choose from:

**About, Data Access, Explore** , **News** , and **Help**.  
![Screen Shot 2022-12-09 at 1.58.24 PM.png](https://help.arkportal.org/__attachments/a_14fda4655c634eb3ff07df0d51a26d8c28bbae554a5675fd8db70d7abb1b5df8/Screen%20Shot%202022-12-09%20at%201.58.24%20PM.png?cb=ec80f6dbb535b8fec0be8444706dbaae)

Let's go through each one so you understand how to navigate the portal in a way that best suits your needs.

## [About](https://arkportal.synapse.org/About)

This is information about who funds and supports the Portal.

## [Data Access](https://arkportal.synapse.org/Data%20Access)

Although anyone can view data readily from the portal (thanks to our [open data sharing model](https://help.arkportal.org/help/about-data-sharing)), you need a Synapse account in order to download data.

Data held in the portal falls into two categories: Open and Controlled Use. While Open data is available for all registered Synapse users without limitations, Controlled Use data is available to registered, certified, or validated Synapse users that fulfill specific requirements for data access, including a signed Data Use Certificate and Intended Data Use Statement.

Find instructions to register for a Synapse account and to submit a DUC on the the Data Access page

## Explore

The **Explore**tab isn't a page on its own, but a menu of subtabs/pages to choose from:  
![Screen Shot 2022-12-09 at 2.08.15 PM.png](https://help.arkportal.org/__attachments/a_095199576efa4ab347c0e08e39f6138d1593c58b445cecb730f18fb28864c947/Screen%20Shot%202022-12-09%20at%202.08.15%20PM.png?cb=c68bc81afc7806a70f60b9222ebb9cff)

These items correspond to the different ways in which you can find information about and filter and view data.

For example, the [**Programs**](https://arkportal.synapse.org/Explore/Programs)page lists all research programs, so you can read about and select one of interest. A Program may have one or multiple[**Projects**](https://arkportal.synapse.org/Explore/Projects) that describe further details of the Program. Data in the Portal are grouped into **Datasets** , which are linked back to the Project they were generated through. There are 2 types of Datasets ([**Collections**](https://arkportal.synapse.org/Explore/Collections)). Experimental Data - which is are files related to specific assays, and/or diseases, and Publications - which are files used in specific manuscripts. The [**All Data**](https://arkportal.synapse.org/Explore/All%20Data) tab allows for exploration of all files in the portal irrespective of which Program, Project, or Dataset the files are associated with.  
![ARK Portal Docs Figures.jpeg](https://help.arkportal.org/__attachments/a_a271d4b648499c885395996bdebae379782c675fe52fd6c9d60be71f2f0b9bd3/ARK%20Portal%20Docs%20Figures.jpeg?cb=2c23a82680d668364e59f97d6d674357)

## [News](https://news.arkportal.org/)

The **News**tab links to a page providing data release notes.

## [Help](https://help.arkportal.org/help/)

This tab brings you to this help site!

---
language: "en"
---
# Sending Data to Cloud Analytical Platforms

## Overview

The ARK Portal provide direct integration with cloud-based analytical trusted-research environments, allowing you to seamlessly transfer datasets for computational analysis without manual downloads. Supported platforms include:

* [**Cavatica**](https://docs.cavatica.org/docs/getting-started) - A cloud-based data analysis platform powered by Seven Bridges

* [**Terra.bio**](https://support.terra.bio/hc/en-us) - A biomedical research platform by The Broad Institute

![Screenshot 2025-08-21 141019.png](https://help.arkportal.org/__attachments/a_9b69993a16c0a914481822eed34a84401d6ca60f8e78a233f52b0ed6ccdf310c/Screenshot%202025-08-21%20141019.png?cb=c414f6962e0ce351897a5c7427439212)

## Prerequisites

Before sending data to an analysis platform:

1. **Create an account** on your chosen cloud workbench (i.e., Cavatica or Terra)

2. **Link your account** to the data portal through your user profile settings

3. **Accept the required terms of use** for both the portal and analytical platform

4. **Verify data access permissions** - You must have appropriate access to the datasets you want to transfer. Visit the ARK Portal [Access \& Attribution help documents](https://help.arkportal.org/help/data-use-certificate) for more details.

*** ** * ** ***

## How to Send Data

### Step 1: Select Your Data

Navigate to the **Explore** or **Data** section of the portal and use filters to find datasets of interest. You can select:

* Individual files

* Entire datasets

* Multiple files across different studies

Add items to your selection using checkboxes or an "Add to Cart" feature.

### Step 2: Initiate Transfer

1. Click the **"Send to \[Platform Name\]"** button (location varies by portal - typically in the top navigation, cart, or file browser)

2. Choose your destination \& sign in to your workbench account when prompted

3. Grant permission for the portal to access your workbench account

### Step 3: Confirm Transfer

1. Review your selected files and destination

   1. **For Cavatica:** Select an existing project from a dropdown of your current projects or create a new project

   2. **For Terra:**Select an existing workspace or provide a project name to create a new workspace

2. Click **"Confirm"** or **"Export"** to begin the transfer

3. Wait for the confirmation message - transfers typically complete within seconds to minutes depending on dataset size

4. You can close the portal window and transfer continues in the background

### Step 4: Access Your Data

1. Navigate to your analysis platform account

2. Open the destination project

3. Files will appear in the project's data section, organized by study or file type

4. Metadata and file manifests are typically included automatically

*** ** * ** ***

## Important Notes

### DRS (Data Repository Service) Connections

DRS is a standardized API that enables secure, direct data access between repositories and analysis platforms. When a platform has "DRS connection," it eliminates the need to manually download and upload data. The analytical platform can directly access files from Synapse with proper permissions.

You only need to authorize the connection once, then data appears automatically in your analysis workspace without manual file transfers. The authorization must be refreshed periodically to maintain security, at which point you'll be prompted to reconnect.

### Key Advantages

* ++No physical copying++: The platform receives secure access to files stored in cloud buckets without waiting for downloads or duplicating data.

* ++Immediate availability++: Files are accessible in your project immediately after transfer, leading to faster workflow setup.

* ++No cost for transfer or storage++: Sending data from Synapse to integrated platforms does not incur egress or storage fees. The alternative is self-managed compute environments that require manual data transfer using command-line tools or APIs.

* ++Original data remains++: Portal data is not removed or affected by the transfer, with better data provenance tracking.

* ++Collaborative \& cost-effective++: Easily share analysis and results with team members in a secure environment and only pay for the comptuational resources you use.

### Platform-Specific Information

**Cavatica:**

* Organize data into projects that can be shared with collaborators

* Supports CWL and WDL workflows

* Includes pre-loaded analysis tools and pipelines

* [CAVATICA Support: Import from a DRS server](https://docs.cancergenomicscloud.org/docs/import-from-a-drs-server)

* [++CAVATICA Support: Start a bulk import job++](https://docs.cavatica.org/reference/start-a-bulk-import-job)

**Terra:**

* Data is transferred to workspace data tables

* Supports WDL workflows through Cromwell

* Integrates with Google Cloud Platform

* [Terra Support: How to access data with DRS URIs](https://support.terra.bio/hc/en-us/articles/6635247495579-How-to-access-data-with-DRS-URIs)

* [Terra Support: Interactive Analysis apps on Terra](https://support.terra.bio/hc/en-us/articles/5400629590427-Interactive-Analysis-apps-on-Terra)

### Authentication \& Security

* **Reconnect if needed** - Authentication tokens may expire; re-link your account in portal settings if transfers fail

* **Permission-based** - Only data you have authorized access to can be transferred and is verified before transfer occurs

* **Audit trails** - Transfers are logged for compliance and reproducibility

*** ** * ** ***

## Troubleshooting

**Problem:** "Send to Platform" button is disabled or unavailable

**Solution:** Ensure you have linked your platform account in your portal profile settings

**Problem:** Transfer fails or times out

**Solution:** Check that your platform account is active and authentication hasn't expired; re-authenticate in portal settings

**Problem:** Files don't appear in destination project

**Solution:** Refresh the platform interface; verify you selected the correct project; check that files completed transfer in the activity log

**Problem:** Don't see your desired platform option

**Solution:** Confirm the platform is supported by your specific data portal; check if additional access requests are required

*** ** * ** ***

## Additional Resources

* If you have questions, suggestions, or feedback about the ARK Portal, please contact us through the [ARK Portal Service Desk](https://sagebionetworks.jira.com/servicedesk/customer/portal/11).

* For general question or issue related to Synapse please contact through the [Synapse help desk](https://sagebionetworks.jira.com/servicedesk/customer/portal/9)

