Implementation of a Structured Data Management Template for Information Transfer Between Contracting Organizations and their Partners

By: Michael J. Barrett,ab Miu-Ling Lau,ab Peter Tattersall,ac Brian A. Taylor,ad John Dzmil,b Andrew Lennard,ae Gabrielle Barney,af and Brian Goodag

Abstract

Our shared vision across the pharmaceutical/biopharmaceutical industry is to enable frictionless data sharing among collaborators and efficient data transfer to regulatory agencies, adhering to industry-standard protocols such as FHIR (Fast Healthcare Interoperability Resources). To realize this vision, we emphasize the necessity of converting unstructured data - commonly found in PDF reports and unstructured spreadsheets - into a standardized format. While systems are being developed for automated, paperless data exchange between collaborators (e.g. Accumulus Synergy), and standard data framework is being developed for data capture (e.g. Allotrope Foundation), this manuscript presents a framework for structured data management aimed at facilitating seamless information exchange between pharmaceutical manufacturers and their Contract Research Organizations (CROs) and Contract Manufacturing Organizations (CMOs).

We propose a structured data collection template that enhances data organization and streamline transfer processes. This template is designed to accommodate various data and metadata types, ensuring that data are stored in a consistent format. By integrating multiple sheets, we can apply standardized logic to generate outputs suitable for scientific discussion and reporting.

This paper discusses the motivation behind the template, its foundational principles, structural design, data transfer methodologies, and implementation strategies within and between organizations. Furthermore, we explore future advancements in data management and the potential impact of widespread adoption of this structured data management template on the pharmaceutical industry. The template and supporting material can be found at https://iqconsortium.org/.

Introduction

Pharmaceutical pipelines are becoming more diverse, with more prevalent use of Contract Research Organizations (CRO) and Contract Manufacturing Organizations (CMO). This trend necessitates the efficient integration of data from various sources for informed decision-making, data trending (e.g. stability, batch release, process development and characterization and DOEs) and/or regulatory filings. One of the largest barriers to efficient transfer of data are reports that are commonly inconsistent and lacking in structure.1 Data reports used to transfer raw data from CROs and CMOs to their contracting partner are commonly unstructured and inconsistent, (Figure 1), which makes automated data parsing very difficult. Such reports commonly vary between CRO to CRO, project to project, and even analyst to analyst. These variations, combined with lack of structure in these reports, make the automated consumption of data nearly impossible. Additionally, raw data generated from various sources undergo multiple rounds of transcription and review, often leading to resource inefficiencies and a heightened risk of human error (Figure 2). The ideal scenario for data transfer involves a fully automated data pipeline, where the data are reviewed at point of generation with a human setting the context, further contextualized as appropriate, stored in a manner that maintains its integrity and then consumed (Figure 3). This streamlined process would facilitate direct utilization of source data, and tailoring to the specific requirements of the final documentation, such as in the case of FHIR’s proposal for an idealized state for end-to-end data transfer (Figure 4).2,3

628058-1.jpg

Figure 1. Inconsistent and unstructured nature of common reports. Reports delivered from CROs and CMOs to their partner companies commonly have 1.) blank cells, 2.) nested tables, 3.) merged or split cells, 4.) inconsistent structure, and/or 5.) separate tables, all which contribute to the difficulties of the consumption of report data from CROs and CMOs.

628058-2.jpg

Figure 2. Current state of managing CRO data with different data structures, multiple transcription steps and inconsistent storage require repetitive checks.

628058-3.jpg

Figure 3. Ideal state of CRO data flow with standard data structure, validated transfer and centralized data storages.

628058-4.jpg

Figure 4. Ideal state of data pipelines as envisioned by FHIR3 (flames representing the FHIR HL7 standards). The relationships between CRO/CMO and their contracting partners are designated by the blue-dashed outline. Proposed in this manuscript is a data transfer paradigm to allow structured transfer of data before the investments to build APIs between test labs and partners internal systems can be made.

Leveraging highly organized and machine-readable structured content would enable efficient storage, processing, and direct integration into reports and filings. Properly validating the transfer would eliminate the need for manual verification, thus saving time, reducing errors, and improving data accessibility.4 Additionally, well-structured data simplifies system management, optimizes storage, enhances automation, and supports advanced analytics and scientific research throughout all development stages.5 To accomplish this, connectors between each step would need to be developed independently and often at different times. While the connectors and data transformers (that support interoperability of different data sources) downstream of the partners internal systems remain largely in the contracting partners’ (sponsors’) control, the data transmission from third party testing labs and manufacturing requires a partnership to build the connections. With the large number of relationships between contractors and their partners, building Application Programing Interfaces (APIs) for all relationships will cost both time and money, with the potential to delay the readiness of downstream systems. To realize the full benefits of a data pipeline sooner and to realize full digital benefits from slower moving relationships, a standardized data transfer approach as a bridging solution is proposed for the less mature digital relationships between CRO/CMOs and their partners. A common template across analytical groups, CMOs, and CROs is proposed to reduce the number of unique formats, simplify data entry, lower costs, and minimize errors. This approach accelerates data accumulation, supports timely scientific decisions, and speeds regulatory submissions. Standardized data also facilitates sharing and collaboration across organizations and integrates seamlessly with analytical tools (i.e. visualization and statistical software) for efficient analysis and reporting. Additionally, this would eliminate the bottleneck of the individual CRO/CMO relationships and allow contracting partners to realize the full benefits of an end-to-end data pipeline before direct transfer of data can be realized.

In the pharmaceutical sector, the relationships and digital capabilities between Contract Research Organizations (CROs), Contract Manufacturing Organizations (CMOs), and their partners vary widely. Some partnerships benefit from advanced digital infrastructures that enable smooth, system-to-system data exchanges, while others still depend on traditional paper-based reporting. This disparity is influenced by factors such as company size, digital maturity, and contractual terms. As a result, a partner may receive data electronically from one CRO but only paper reports from another. The lack of standardized data formats and clear guidelines for collecting and sharing experimental data creates inconsistencies that hinder efficient data management, reduce data quality, and limit reuse. With thousands of CRO/CMO partnerships in play, these challenges multiply. To overcome these barriers and move toward a fully integrated data pipeline, a standardized data template would allow any CRO or CMO to deliver data in a consistent format, simplifying data parsing and integration for partners and third parties, and accelerating the transition to seamless, end-to-end data sharing.

Principles

The template is designed to collect summarized experimental results that serve as a Project Team’s Primary Data set for review, and transfer to other systems for additional analysis, and reporting in company specific systems, formats and reports with data integrity maintained. FAIR data and Lean Process principles were applied in designing the templates.

FAIR data– Findable, Accessible, interoperable, and Re-usable.6

Findable is interpreted as uniquely referencing data to quickly locate and link related data stored in different places. A single Primary master data set file must be maintained, locatable by all, to ensure project teams members have access to the most up to date results and are using the same version of results to make project decisions. This primary dataset would act as the ‘source of truth’ when referenced to downstream use of the data.

Accessible is interpreted as storing data in a software system and file format widely available to all companies and the retrieval process from a system easy to set up. There are very few systems available to all companies that are familiar and accessible by the relevant personnel.

Interoperable is interpreted as easily moving data between different systems in the present and future. The template design should reflect that humans and machines will interact with the data. The captured structured data follows good spreadsheet data entry practices and minimizing LEAN Methodology waste types. Template design aims to minimize data pre-processing time before moving or re-using data.

Re-usable is interpreted as sufficient data stored to provide context and confidence in the experimental work done to use results at the time and re-use in the future. Data needs capturing to define how experiments were carried out to conclude that they were done as intended and sufficient information provided to give context to what the results mean.

628058-5.jpg

Figure 5. Qualitative availability and familiarity with selected software. While some software is widely available to the public, support within organizations may limit their availability for research and development.

To meet the FAIR data principles, a universal template must be both familiar and available to the broadest audience possible (Figure 5). We have chosen a spreadsheet-based template to meet these criteria.

Template Structure

The template (See Figure 6) consists of a spreadsheet with multiple tabs. Its organization was designed to balance human readability and data entry, with machine readability by maintaining a structured format. The first 3 tabs include (1) “Study_Metadata” (2) “Sample_Run_Metadata” and (3) “Specification_Metadata”. These allow the user to capture pertinent context for the result data while minimizing repeat data entry. The study metadata tab captures information that is static throughout the entire study, such as sample or run identifier, project name, study name, stability timepoint and stability storage temperature. The “Sample_Run_Metadata” tab captures information specific to the run being conducted. Depending on the study design and input preferences of the users, some information could be captured on either the “Sample_Run_Metadata” tab or the “Study_Metadata” tab. The tabs would be formatted by the contracting partner, eliminating duplicate columns, so that the metadata would only be captured on one tab. The information, number of columns, or other variables captured in these tabs can be expanded to provide additional experimental context for different study types (i.e. stability, robustness DOE, forced degradation). The “Specification_Metadata” tab captures information about the specification and acceptance criteria.

The next set of tabs captures sample/test specific information, such as the testing references, dates, and methods used. The tabs “Testing_Reference”, “Testing_Dates”, and “Testing_Methods” are given as examples. The result data tabs can be organized in either a (1) Long or (2) Wide format, which could be selected at the contracting partner based on relationship, project needs, or partner capabilities. Both formats are pivoted tables, grouped by a sample identifier for ease of data entry and will require basic manipulation (Transpose, unpivot, merge) for data ingestion. Additionally, both were designed to be a simple transpose of one another to allow for more robust interchangeability. The tables capture the result, name of the result, and test performed.

To link the data across tabs, key columns are utilized:

  • “Study Number” link metadata across the “Sample_Run_Metadata” and “Study_Metadata” tabs
  • “Specification Revision” link metadata across the “Sample_Run_Metadata” and “Specification_Metadata” tabs
  • “Test” and “Reportable Measurement” link metadata across the “Results” and “Specification_Metadata” tabs
  • “Sample Identifier” links data across the “Results,” “Sample_Run_Metadata,” and the sample/test specific information tabs (such as “Testing_Reference,” “Testing_Dates,” and “Testing_Methods”)

The templates are designed as a guide, where additional fields (such as tests, conditions, or other information) can be added to capture data as needed, as long as key columns and recommended fields to aid in data uptake are captured, as a well-designed uptake protocol would be able to upload the information into relevant databases as well as add the appropriate ontology to the data.

628058-6.jpg

Figure 6. Example of spreadsheet configuration

The current version of template along with a detailed user guide and other supporting material is currently being archived with the IQ consortium at https://iqconsortium.org/.

Data Transfer

Data transfer between CROs, CMOs, and their partners must address the wide variation in digital capabilities and ensure compliance, data integrity and usability. The lack of standardized data formats and clear guidelines for collecting and sharing experimental data creates inconsistencies that hinder efficient data management, reduce data quality and limit reuse. To manage these challenges, four potential approaches support controlled and secure data exchange. Any approach utilized, must meet internal standards and procedures for the data’s intended use.

  • Direct electronic signoff of spreadsheets: Utilize native spreadsheet features that allow electronic signatures or digital certificates (e.g., DocuSign could potentially fulfill that role if permitted by procedures of all parties) to approve data directly within the file. This method ensures the spreadsheet remains in a digital format that can be easily consumed by downstream applications without the need for data extraction or parsing. The signatory needs to meet industrial requirements.7
  • Access to native files in compliant document repositories: Store locked, official versions of documents in secure, shared repositories that support native file formats (e.g., spreadsheets, word processing files). Partners can be granted controlled access to view these files directly, ensuring data integrity and traceability.1
  • Generation and signing of PDFs with OCR conversion: Produce signed PDF reports as official documents, then apply Optical Character Recognition (OCR) technology to convert the PDFs back into editable and searchable digital data. OCR technology has advanced significantly, enabling accurate recognition of diverse layouts, formats, and special characters.8
  • Dual document delivery with signed PDFs and CSV files: Provide a signed PDF as the official, regulatory-compliant record alongside a CSV file containing the raw data marked “for information only.” This approach supports non-GMP activities such as trend monitoring and root cause analysis while maintaining strict control over official data. The result of the copy process should be verified either automatically by a validated process or manually to ensure that the same information is present —including data that describe the context, content, and structure — as in the original.9

These methods, governed by internal policies and regulatory requirements, help maintain data accuracy, security and accessibility throughout the transfer process, addressing the complexities introduced by thousands of CRO/CMO partnerships.

Use of Template

Once the structured data is transferred into the Company’s environment and ingested into the Company’s data lake, it can be combined with other internal datasets for comprehensive analyses and reporting. The meta-data from the CSV file provides the context of the Results data and will enable combining of CRO/CMO with internal data. Additionally, a standardized format enables the development of validated automated tools for downstream data transfer, analytics and reporting, significantly reducing the need for manual reviews.

To successfully implement these standardized structure templates, the company must secure a strong partnership and prioritize training for staff to ensure smooth adoption and integration into existing workflows. Collaboration with the IT department is essential for developing utility tools for data merging and transformation, ensuring seamless integration with internal systems and sustainability when system changes or upgrades take place. Engaging with CROs and CMOs to promote the adoption of these templates will create a collaborative environment that benefits all stakeholders. Finally, the company should regularly review and update the templates based on feedback and evolving industry standards to maintain their relevance and effectiveness. By leveraging standardized structure templates, the company can significantly enhance its data management processes, improve collaboration with partners, and position itself for future growth and innovation in the biopharmaceutical industry. The template has been demonstrated to import data for use in downstream processing.

Future Goals

Looking forward, the long-term goal is to eliminate physical documentation entirely by consolidating diverse data sources within cloud-based enterprise data lakes (EDLs). These platforms will organize data into semantically structured blocks using consistent nomenclature and metadata, allowing flexible extraction and presentation of information in any desired format. Automated Structured Content Data Management (SCDM) will enable real-time data sharing among stakeholders, fostering greater collaboration and regulatory compliance, ultimately accelerated regulatory dossier submission readiness.

Recognizing that full adoption across stakeholders of these advanced technologies will take time, the standardized data reporting templates described above are proposed. This approach introduces document-based structured content data management as an interim solution, supporting gradual migration from paper-based exchanges to fully digital, cloud-enabled workflows among manufacturers, CROs, CMOs, and other partners. As digital exchanges are brought online between CROs/CMOs and their partners,10 it is expected that the templated solutions will be phased out.

The data reporting templates presented here respond to the urgent need for a harmonized minimum standard that supports structured data sharing across the pharmaceutical ecosystem. Designed to integrate seamlessly with existing data pipelines, these templates facilitate direct transfer of information from CROs and CMOs to their partner organizations, including connections to LIMS and raw data sources such as chromatographic instruments.

As part of this initiative, a collaborative technology build is underway with a third party to demonstrate proof-of-concept capabilities encompassing data collection, aggregation, visualization, and reporting. Supported by the Enabling Technology Consortium (ETC), this initiative is already delivering tangible benefits, including streamlined data transfer, enhanced access to well-organized datasets, and improved data integrity - particularly for analytical stability data and manufacturing process transfers.

Although the initial template version incorporates a limited set of taxonomies and lacks a fully developed ontology, its standardized structure lays out the foundation for future ontology integration. This will enable alignment with established data frameworks like those from the Allotrope Foundation, thereby increasing the templates’ versatility and effectiveness.10,11 Moreover, this approach supports the development of reusable analytics and reporting tools and prepares the ground for adopting industry standards such as FHIR for data exchange.3,12

Conclusion

Implementing a structured content data management template offers numerous benefits for businesses, including improved data organization (Findable, Interoperable), simplified data transfer (Accessible), and support for electronic filings (Reusable). By adhering to the template’s principles, and utilizing a spreadsheet as the primary tool, companies can achieve consistent and efficient data management. The template’s flexibility and potential for widespread adoption (e.g. use in stability, batch release, process development and characterization and DOEs) opens the door for the development of advanced software solutions for data transfer and import. Working as a consortium across industry helps to simplify the process of all parties involved. Individual CRO and CMOs would be able to use the same format, regardless of their partner. This would allow for the technicians and bench scientists that populate and review the template, to enter the data in the same format, thereby reducing the potential of error. Contracting partners would be able to reduce the resource burden associated with introducing data templates to the CRO and CMOs when working with multiple partners. Adoption of the same CRO/CMO data transfer solution across the industry would simplify the process for all parties involved. By embracing this templated tool companies would be more enabled to effectively manage structured data and gain actionable insights for informed decision-making. The proposed solution provides a ‘level playing field’ across stakeholders as a minimum standard of structured data transfer, while more direct cloud-based solutions are developed.

Disclaimer

Positions in this manuscript are those of the authors and are not official opinions of their respective companies.

Acknowlegments

  • The IQ team, IQ Board of Directors, Tony Mazzeo (BMS)
  • This manuscript was developed with the support of the International Consortium for Innovation and Quality in Pharmaceutical Development (IQ, www.iqconsortium.org). IQ is a not-for-profit organization of pharmaceutical and biotechnology companies with a mission of advancing science and technology to augment the capability of member companies to develop transformational solutions that benefit patients, regulators and the broader research and development community

References

  1. Blumzon CFI, Pănescu A-T. Data Storage. In: Bespalov A, Michel MC, Steckler T, eds. Good Research Practice in Non-Clinical Pharmacology and Biomedicine. Springer International Publishing; 2020:277-297.
  2. Ahluwalia K, Abernathy MJ, Beierle J, et al. The Future of CMC Regulatory Submissions: Streamlining Activities Using Structured Content and Data Management. J Pharm Sci. May 2022;111(5):1232-1244. doi:10.1016/j.xphs.2021.09.046
  3. (HL7) HLSI. Fast Healthcare Interoperability Resources (FHIR) Release 5. Accessed 19 September, 2025, https://www.hl7.org/fhir/
  4. Dempsey KL, Pillitteri VY, Regenscheid A. Managing the Security of Information Exchanges. Special Publication (NIST SP), National Institute of Standards and Technology, Gaithersburg, MD; 2021.
  5. Hu H, Wen Y, Chua TS, Li X. Toward Scalable Systems for Big Data Analytics: A Technology Tutorial. IEEE Access. 2014;2:652-687. doi:10.1109/ACCESS.2014.2332453
  6. Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data. 2016/03/15 2016;3(1):160018. doi:10.1038/sdata.2016.18
  7. FDA Guidance on Electronic Records and Signatures (21 CFR Part 11). U.S. Food and Drug Administration. Accessed September 19, 2025, https://www.fda.gov/regulatory-information/search-fda-guidance-documents/part-11-electronic-
    records-electronic-signatures-scope-and-application
  8. Corp J. Optical Character Recognition (OCR): Emerging Trends and Future Applications. The Identity and Beyond Blog blog. November 13, 2024, 2024. https://www.jumio.com/optical-character-recognition-trends-and-applications/
  9. Guideline on Computerised Systems and Electronic Data in Clinical Trials European Medicines Agency. Accessed September 19, 2025, https://www.ema.europa.eu/
    en/documents/regulatory-procedural-guideline/guideline-computerised-systems-and-
    electronic-data-clinical-trials_en.pdf
  10. Kayser H, Lau M-L. Growing value of data standardization: Allotrope Foundation Connect Workshop Proceedings. Drug Discovery Today. 2024/06/01/ 2024;29(6):103988. doi:https://doi.org/10.1016/j.drudis.2024.103988
  11. Gardiner S, Haynie C, Della Corte D. Rise of the Allotrope Simple Model: Update from 2023 Fall Allotrope Connect. Drug Discovery Today. 2024/04/01/ 2024;29(4):103944. doi:https://doi.org/10.1016/j.drudis.2024.103944
  12. Beierle J, Algorri M, Cortés M, et al. Structured content and data management—enhancing acceleration in drug development through efficiency in data exchange. AAPS Open. 2023/05/08 2023;9(1):11. doi:10.1186/s41120-023-00077-6

aNext Generation Data Analytics Working Group, International Consortium for Innovation & Quality in Pharmaceutical Development

bMerck & Co., Inc., Rahway, NJ, USA

cBMS Bristol Myers Squibb, New Brunswick, NJ, USA

dAstraZeneca, Macclesfield, UK

eAmgen Ltd, Uxbridge, UK

fEli Lilly and Company, Indianapolis, IN, USA

gRayzeBio, a Bristol Myers Squibb Company

Subscribe to our e-Newsletters
Stay up to date with the latest news, articles, and events. Plus, get special
offers from American Pharmaceutical Review delivered to your inbox!
Sign up now!

  • <<
  • >>

Join the Discussion