Statistics

Research Data Sharing Statistics: Adoption, Barriers, and Infrastructure

Research data sharing statistics covering open-data support, clinical research practices, barriers, infrastructure, policy, and estimated economic value.

Research data sharing is widely supported, but support does not always translate into repository deposits, reusable metadata, or easy access. The figures below show how sharing practices, clinical-research infrastructure, perceived barriers, policy expectations, and estimated economic benefits vary across surveys and reports.

Contents

Open-data support and research outputs

The State of Open Data survey reports strong support for making research data available. In 2023, 89% of respondents said they made their data publicly available. In 2025, 81% of surveyed researchers supported open data. These are related but different measures: one describes reported public availability, while the other describes support for open data.

Awareness of FAIR data principles also increased from 15% in 2018 to 40% in 2025. FAIR refers to data that is findable, accessible, interoperable, and reusable. The survey therefore shows a substantial rise in awareness across the 2018–2025 period, but the figures do not establish that every researcher applying FAIR principles also shares data publicly.

The same source recorded growing use of artificial intelligence in data workflows. Active use of AI for data processing rose from 22% in 2024 to 32% in 2025, while AI use for metadata creation increased from 16% to 25% over the same period. These figures measure reported use, not the quality or completeness of the resulting metadata.

The OECD’s synthesis of the ISSA2 study provides a different view of research production. On average, 67% of scientific production resulted in new data or code in ISSA2, published in 2020. About 24% resulted in both new data and new code. About 40% of newly generated data or code was stored in repositories on average. The production of data and code was therefore common in the study, while repository storage covered a smaller share.

How data is stored, shared, and reused

ISSA2 found that repository or journal-supporting-material sharing covered about 45% of authors’ data, compared with about 20% for code. The comparison suggests that data and code did not move through the same sharing channels. It also describes coverage of authors’ outputs in that study rather than a universal rate for all research disciplines.

Repository storage was particularly limited in some fields. Only about 20% to 25% of data or code was stored in repositories in sociology, psychology, and business and management. A fee was required to access shared data in about 12% of ISSA2 cases. That access figure is important when interpreting the word “shared”: deposited material may still have conditions or costs attached to access.

MeasureReported figurePeriod or source context
Scientific production resulting in new data or code67%ISSA2, published 2020
Scientific production resulting in both new data and codeAbout 24%ISSA2, published 2020
Newly generated data or code stored in repositoriesAbout 40%ISSA2, published 2020
Authors’ data covered by repository or supporting-material sharingAbout 45%ISSA2, published 2020
Authors’ code covered by repository or supporting-material sharingAbout 20%ISSA2, published 2020
Shared data requiring an access feeAbout 12%ISSA2, published 2020

An earlier Wiley Open Science Researcher Insights survey, from 2016, found that 69% of researchers said they shared data. Among those data sharers, increasing research impact and visibility motivated 39%, public benefit motivated 35%, and transparency and reuse motivated 31%. Journal requirements motivated 29%. Because these findings are from 2016, they should be read as historical survey evidence rather than a current estimate.

Clinical-research sharing practices

Clinical research shows both substantial exchange and uneven documentation. In a 2019 survey of 174 clinical-research respondents, 61% had used data generated by other researchers, while 77% had shared their own data in some way. Among clinical researchers who had shared data, 76% had shared directly with collaborators, 65% had shared with an external team, and 53% had deposited data in a repository or listed it in a catalogue.

The ways researchers share are therefore broader than repository deposit alone. Direct collaboration and external-team sharing were more common among data sharers than repository or catalogue listing in this survey. At the same time, 25% of all 174 respondents said their latest publication had no data-sharing statement.

Platform awareness was also incomplete. Among clinical-trial researchers, 54% were unaware of any listed clinical-trial-specific platform. Among all clinical-research respondents, 18% were unaware of any listed general platform, and 41% had not heard of any of the 20 listed platforms. These results come from the 2019 survey and reflect awareness of the platforms included in that survey.

Barriers and researcher concerns

Researchers reported concerns that affect whether data can be shared and under what conditions. In the 2016 Wiley survey evidence summarized by the OECD, intellectual-property or confidentiality concerns were cited by 50% of researchers as a reason to hesitate. Ethical concerns were cited by 31%; concern about misuse or misinterpretation by 23%; and concern about being scooped by 22%.

Clinical-research respondents emphasized operational and governance difficulties. In the 2019 survey, lack of funding was rated challenging by 88%, existing informed consent by 87%, country-specific data protections by 82%, and limited knowledge of data platforms by 82%. Lack of technical expertise was challenging for 81%, while variable access or ethics requirements were challenging for 80%.

Clinical data-sharing challengeRespondents rating it challenging
Lack of funding88%
Existing informed consent87%
Country-specific data protections82%
Limited knowledge of data platforms82%
Lack of technical expertise81%
Variable access or ethics requirements80%

These percentages describe perceived challenges, not failures in every individual project. They also show why a policy requiring sharing may not be enough by itself: researchers may still need funding, consent language, legal interpretation, technical support, and a suitable access model.

Infrastructure and practical access

Access to practical resources was uneven among the clinical-research respondents surveyed in 2019. Fifty-nine percent had access to data-management tools, 49% had access to a repository, and 40% had access to a secure analysis environment. Support for anonymisation and curation was each available to 28%, while catalogue-entry support was available to 14%.

Thirteen percent had access to none of the six listed resources. The six resources were data-management tools, a repository, a secure analysis environment, anonymisation support, curation support, and catalogue-entry support. The result indicates that resource availability extended beyond the question of whether a repository existed: privacy protection, analysis environments, and description support were separate parts of the sharing infrastructure.

The infrastructure figures also help explain the gap between willingness and execution. A researcher can support open data yet lack a secure environment for sensitive records, help preparing anonymised files, or assistance with catalogue metadata. The survey does not measure the quality or capacity of each resource, so access should not be interpreted as guaranteed suitability for every dataset.

Policies, incentives, and compliance

Clinical-research respondents expressed strong expectations of funders. In the 2019 survey, 84% agreed that funders should provide funding for data management and sharing, and 76% agreed that funders should provide guidance on where and how to share. Half agreed that funders should make data sharing a grant condition, while 47% agreed that funders should consider data sharing in grant applications as credit.

These results distinguish support from enforcement and recognition. Funding and guidance received higher agreement than making sharing a grant condition or awarding grant credit. The survey therefore points to several policy tools rather than one universal incentive.

NIH dbGaP statistics provide a separate measure of controlled-access administration. Through July 1, 2018, NIH had approved 42,292 dbGaP data-access requests and identified and managed 38 policy-compliance violations among those requests. NIH reported that the 38 violations represented 0.1% of approved data-access requests. These are historical NIH administrative statistics for that date and should not be treated as a general compliance rate for all research data repositories.

Estimated economic value of data access

The OECD’s 2019 report synthesis reviewed studies that estimated economic benefits from public-sector data access and sharing at 0.1% to 1.5% of GDP. Studies including private-sector data estimated benefits of 1% to 2.5% of GDP, while a few studies estimated benefits as high as 4% of GDP including private-sector data.

These figures are estimates reported in studies reviewed by the OECD, not a single measured global return. Their ranges differ by the data included and the modeling approach. They are best used to describe the scale of possible economic value associated with data access and sharing, rather than as a forecast for a particular country, institution, or research project.

Written by

scientifist.com Editorial Team

Editorial team

scientifist.com publishes practical how-to guides and educational articles with clear steps and useful context.