Twenty years ago, data-sharing policies mostly asked researchers to promise that files would be available. Today, sharing is built into the machinery of research funding. The US National Institutes of Health moved from its 2003 policy for selected, larger grants to a 2023 Data Management and Sharing Policy covering nearly all NIH-supported research that generates scientific data. The European Commission has made open science and FAIR data central to Horizon Europe, while UK Research and Innovation is developing a harmonised research data policy. The direction is clear: sharing is moving from aspiration to expectation.
That is substantial progress. But a requirement to deposit data is not the same as an incentive to make data genuinely reusable. A file can be technically available yet impossible to interpret; a repository record can lack documentation, provenance or a workable licence; and a data-availability statement can say “available on request” even when requests are rarely answered. Compliance can establish a floor. It cannot, by itself, create a culture in which careful stewardship is rewarded.
The policy story can be read in three stages. The first made sharing permissible and encouraged it. The second made planning, deposition and access conditions more explicit. We are now entering a third stage: recognition. The central question is no longer only whether data were shared, but whether the sharing was useful—and whether the people who did the difficult work receive credit.
Several parts of the research system are approaching that question from different directions. The FAIR Principles gave the community a vocabulary for making data findable, accessible, interoperable and reusable. FORCE11’s data-citation principles established that datasets should be treated as legitimate, citable research outputs. DataCite supplies persistent identifiers and metadata that allow datasets to enter the scholarly record. Make Data Count has developed infrastructure for standardised views, downloads and citations, as well as a Data Citation Corpus that links datasets to articles that use them.
This infrastructure matters because conventional citation indexes were designed around papers, not the complicated life of data. A dataset may be cited through a DOI, mentioned only by an accession number, combined with dozens of other resources, or reused in software and artificial-intelligence (AI) workflows that never generate a conventional citation. Views and downloads capture attention, but not necessarily scientific use. Direct citations capture explicit acknowledgement, but not the full downstream influence of a resource. No single signal is sufficient.
A parallel movement is changing what counts as a research contribution. The CRediT taxonomy makes “data curation” a named role alongside conceptualisation, analysis and writing. The Declaration on Research Assessment (DORA) asks institutions to consider datasets and software, not just publications, and to combine quantitative evidence with qualitative judgement. The Coalition for Advancing Research Assessment (CoARA) similarly calls for recognition of diverse outputs, practices and activities. In the UK, the Résumé for Research and Innovation allows applicants to demonstrate a wider range of contributions than a traditional publication-centred CV.
These reforms address a different failure from repository infrastructure. Persistent identifiers can make a dataset visible, but they do not ensure that a hiring panel, promotion committee or funder will value the work behind it. Conversely, narrative CVs can invite researchers to describe data stewardship, but evaluators still need trustworthy evidence. Recognition requires both: infrastructure that records use and assessment systems prepared to interpret it.
The current generation of experiments is beginning to connect those layers. The NIH Data Sharing Index, or S-index, Challenge asks how to recognise high-quality sharing that creates downstream value rather than simply counting deposits. Our finalist project, theSindex.org, examines whether the wider citation graph built on data reuse can reveal influence that direct citations miss. We are also testing a FAIR Agent that helps authors strengthen data-availability statements and repository links before publication.
Three lessons are emerging from this landscape. First, data quality is relational. Reusability depends on the prospective user, the scientific question, the documentation and the standards of a field. A pristine file that cannot be linked to methods or variables may be less valuable than a modest dataset with excellent provenance. Metrics should therefore distinguish availability from stewardship quality and actual reuse.
Second, impact is delayed and field-dependent. Some datasets are reused within months; longitudinal cohorts, reference atlases and rare-disease collections may reveal their value over decades. Citation cultures also vary sharply across disciplines. A universal threshold would advantage fast-moving, densely citing fields and undervalue specialised resources. Any comparison must be contextualised by field, dataset age, scale and access model.
Third, “open” cannot always mean unrestricted. Genomic, clinical and Indigenous data may require consent-based governance, controlled access or community authority. A responsible system should recognise well-governed access and documented reuse without penalising researchers for protecting participants. Counting downloads from an open repository and approved analyses in a secure environment as if they were equivalent would produce the wrong incentives.
Different actors can act now. Funders can pay the real costs of curation and examine reuse after a grant ends. Universities can add datasets, software and stewardship to promotion criteria. Publishers can require formal data citations rather than burying accessions in prose. Repositories can expose standardised, open usage metrics. Researchers who reuse data can cite the dataset itself and describe how it contributed. None of these changes is sufficient alone; together they turn sharing from an administrative endpoint into a visible part of research.
The first two generations of policy helped get data out of laboratories and onto shared infrastructure. The next must make good sharing legible, useful and creditable. If public funders require data stewardship, research assessment should recognise it—not through one number, but through evidence that others could find the data, understand it and build on it.
Disclosure: I lead one of the seven finalist teams in the NIH S-index Challenge and am involved with theSindex.org and its FAIR Agent.
Copyright © 2026 Kuan-lin Huang. Distributed under the terms of the Creative Commons Attribution 4.0 License.