Skip to content

When Data Meets Review: Reflections on PREreview’s First Open Dataset Review

Photo by Gabrielle Henderson / Unsplash

Over the past decade, Open Science has transformed the way we create, share, and evaluate research. Among the many domains of Open Science (Bertram et al. 2023), two have received growing attention: open data and open peer review.

Yet these two worlds do not often meet.

We regularly discuss the importance of sharing data to support transparency, reproducibility, and reuse (Wilkinson et al. 2016). We also explore new approaches to peer review that are more open, inclusive, and collaborative (Ross-Hellauer et al. 2023). But we rarely stop to ask what happens when datasets themselves become the subject of review.

Recent initiatives have begun exploring ways to connect open data and open peer review (Modular Peer Review; Sansing and Wilkinson 2025). One notable development came in  November 2025, when PREreview introduced support on their platform for reviewing datasets deposited in Dryad, an open-access general-purpose data repository. It is in this context that the FORCE11 PREreview Club was invited to pilot a collaborative and structured dataset review process. Our review of the dataset, “Leverage metrics to drive data sharing at the Science journals” by Valda Vinson and Lauren Kmec (2025), is now posted on PREreview and archived on Zenodo (Akuma et al. 2026). It is the first open review of a dataset shared on PREreview.

In this blog post, we reflect on that experience, discuss what we learned, and consider the opportunities of dataset review as an emerging scholarly practice.

Reviewing More Than Results

The Vinson and Kmec dataset (2025) focuses on data and code-sharing practices in articles published by Science. At first glance, it seemed like an ideal case study: data deposited in an open repository, documentation available for users, and a topic closely connected to Open Science itself.

As our discussion unfolded, however, it became clear that reviewing a dataset involves different questions from those we typically ask when reviewing a preprint or a journal article. Rather than focusing solely on findings or conclusions, we found ourselves examining the processes behind the data. How were key variables constructed? How transparent were the classification methods? What assumptions shaped the dataset? How reproducible was the workflow? And what limitations should future users keep in mind when reusing the data?

Very quickly, the conversation moved into topics such as FAIR and CARE principles, metadata quality, methodological transparency, provenance, responsible reuse, and the implications of using automated systems to measure Open Science practices.

In a sense, the review became a metascientific exercise: a group of researchers evaluating a dataset designed to measure behaviors associated with Open Science itself. What began as a discussion about a dataset soon became a broader conversation about trust, transparency, and how we assess research outputs beyond traditional publications.

A Conversation Across Research Communities

One of the most rewarding aspects of the experience was the diversity of perspectives involved in the review. The review brought together researchers based in India, Mexico, Türkiye, the United Kingdom, and the United States, with backgrounds spanning scholarly communication, research assessment, bibliometrics, data management, and Open Science.

Although we were all reviewing the same dataset, we examined it through vastly different lenses. Some participants focused on reproducibility and workflow validation. Others paid closer attention to metadata quality, cross-publisher comparisons, methodological transparency, or the broader ethical and geopolitical implications of data reuse.

 These different and diverse perspectives enriched the discussion. Together, they helped us develop a more nuanced understanding of the dataset, its strengths, and its limitations than any one of us could likely have achieved alone.

This collaborative process also highlighted one of the strengths of open, collaborative peer review. When reviewers bring different experiences, disciplinary backgrounds, and regional perspectives to the table, evaluation becomes more than quality control—it becomes a form of collegial discourse and collective learning.

Open Data Is Not the End of the Story

Perhaps the most important lesson for us was seeing how open review can help bridge the distance between data producers and data users.

One of the greatest challenges in data reuse is not simply gaining access to data, but understanding the context in which the data were created, the decisions that shaped them, and the limitations that accompany their use. Datasets rarely speak entirely for themselves. Even well-documented data often leave room for questions about methods, assumptions, and interpretation.

Open review creates a space where those questions can be raised collectively, transparently, and constructively. It also provides an opportunity for data producers and data users to engage in dialogue, helping to clarify uncertainties, identify issues that may affect reuse, and strengthen future versions of a dataset. 

The experience also reminded us that making data available is not the same as making data understandable, reusable, or trustworthy for every purpose. Openness requires infrastructure, but it also requires documentation, clear licensing, robust metadata, methodological transparency, and ongoing dialogue with the communities that may eventually reuse the data. As datasets evolve through updates, corrections, and new versions, community feedback can play an important role in supporting their long-term reusability.

In other words, depositing a dataset in a repository is not the end of the story.  Rather, it marks the beginning of a new phase for data to become part of the broader scientific discourse. Once shared, datasets can be scrutinized, debated, and critiqued, just like preprints and journal articles. Through review and reuse, datasets can contribute to the refinement of existing knowledge, the generation of new insights, and the advancement of future research.

This experience also reminded us that Open Science extends beyond open access to publications. While open access remains one of its most visible achievements, practices such as open data, open peer review, and responsible data reuse are equally important for building a more transparent and collaborative research ecosystem.

Open Data and Open Review: Better Together

Our open dataset review pilot for PREreview does not answer every question about how research data should be evaluated. But it does demonstrate something important: open data and open review are not separate practices.

When combined, they create new opportunities to strengthen transparency, accountability, and reuse across the research ecosystem. Open data makes reuse possible. Open review helps us understand what we are reusing, how it was created, and what limitations we should keep in mind. Together, they move us closer to a vision of Open Science that extends beyond access alone—one that values not only the availability of research outputs, but also the conversations, feedback, and collective stewardship that help make those outputs meaningful and reusable.

Finally, we would like to thank Jennifer M. Miller for her thoughtful coordination of the review and collaborative writing process, as well as Daniela Saderi and the PREreview team for creating and continuously improving spaces where open and community-driven approaches to research assessment can flourish.

Perhaps that is the most valuable lesson from this experience: opening data matters, but opening the conversation around data can be just as transformative.

References

  1. Akuma, I., Dogan, G., Spick, M., Rogel-Salazar, R., & Li, X. (2026, June 5). Structured PREreview of “Leveraging metrics to drive data sharing at the Science journals.” https://prereview.org/reviews/1d1731cb-e9b9-4bb4-b636-98c235379cc6
  2. Bertram, M. G., Sundin, J., Roche, D. G., Sánchez-Tójar, A., Thoré, E. S. J., & Brodin, T. (2023). Open science. Current Biology: CB, 33(15), R792–R797. https://doi.org/10.1016/j.cub.2023.05.036 
  3. Modular Peer Review. (n.d.). Continuous Science Foundation. Retrieved June 6, 2026, from https://continuousfoundation.org/peer-review 
  4. Ross-Hellauer, T., Bouter, L. M., & Horbach, S. P. J. M. (2023). Open peer review urgently requires evidence: A call to action. PLoS Biology, 21(10), e3002255. https://doi.org/10.1371/journal.pbio.3002255 
  5. Sansing, C., & Wilkinson, C. (2025, October 15). Now you can review datasets on PREreview.org. PREreview Blog. https://content.prereview.org/now-you-can-review-datasets-on-prereview-org/ 
  6. Vinson, V., & Kmec, L. (2025). Leveraging metrics to drive data sharing at the Science journals [Dataset]. Dryad. https://doi.org/10.5061/DRYAD.ZKH1893QT 
  7. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. https://doi.org/10.1038/sdata.2016.18 

Copyright © 2026 Rosario Rogel-Salazar, Xiuqi Li, Güleda Doğan. Distributed under the terms of the Creative Commons Attribution 4.0 License.

Comments

Latest