SlideShare a Scribd company logo
Stewarding Big Data: Perspectives on Public Access
   to Federally Funded Scientific Research Data

 Big Data and Big Challenges for Law and Legal Information
                   Georgetown Law Library
                      January 30, 2013



                    William G. LeFurgy
                    Library of Congress
                         @blefurgy
My Perspective on Big Data
             Stewardship
• Realizing full potential from big data depends
  keeping it accessible over time
• Accessibility depends on life cycle management,
  most especially preservation
• Advocate for collaborative, distributed model
• Understand that “stewardship” has a different
  meaning for many data creators
White House RFI Input
             Instructive
• Request for Information on Public Access to
  Federally Funded Scientific Research Data, Nov.
  2011
• Interested individuals and organizations to
  provide recommendations on approaches for
  ensuring long-term stewardship and
  encouraging broad public access
• Input provided to inform development of agency
  policies and standards for managing big data
Summary of Responses

• 118 individual responses
  – 50% from academic research departments,
    professional organizations
  – 35% from libraries, repositories and allied
    organizations
  – 10% from publishers and commercial organizations
  – 5% other
• Excellent (unstructured!) data set to analyze
  current thinking on big data stewardship
Top-Level Policy Recommendations

• Remarkable degree of congruence among
  comments
  – Broadly allocate adequate resources for data
    stewardship
  – Extend a collaborative national digital stewardship
    infrastructure
  – Institute and enforce a data preservation mandate
  – Strongly encourage policies to support secondary
    use, respect for data
• But… conflicted about IP, copyright, privacy
Need: Resources
• Funders to include money in awards for data
  stewardship
• Need cost models, other guidance for estimating
  data life cycle costs
• Allocate expanded resources to support national
  data repositories
Need: National Digital Stewardship
          Infrastructure
• Leverage current institutional efforts to
  define best practices, tools, services
• Extend community of practice for data
  stewardship through collaborative action
  across disciplines
• Develop a skilled workforce with data
  stewardship expertise
Need: A Data Preservation Mandate
• Incentivize grant applicants to make realistic
  plans for data
   – Stronger data manager requirements in application
     process
   – Tie future awards to demonstrated success with data
     stewardship
   – Enable direct support of PIs by data stewardship
     specialists
Support: Secondary Use, Respect for Data

• Broadly apply a citation mechanism for
  data sets (e.g., DataCite, DOIs)
• Criteria for evaluating grant applications
  tied to secondary use of data
• Give equal credit for publishing articles
  and data sets
• Develop robust metrics to track data
  publication and use
Muddled Picture for IP
• Opinions diverge about role of copyright,
  patents, etc., in regard to research data
  – Commercial interests see IP as critical
  – Many data users favor Creative Commons or public
    domain approach
  – Data creators fall between these positions
• A significant degree of concern raised regarding
  privacy in connection with IRB, personal data
Next Steps

• Two interagency working groups within the
  National Science and Technology Council
  reviewing recommendations
• Groups will develop science agency policies for
  data dissemination and stewardship
• Potential for major change, as policies may have
  association with funding from the Federal
  science agencies
Websites
Request for Information: Public Access to Digital Data Resulting From
Federally Funded Scientific Research, http://ow.ly/ePB93
Your Comments on Access to Federally Funded Scientific Research
Results, http://ow.ly/ePBb9
National Science and Technology Council, http://ow.ly/h87Li

More Related Content

PPT
Overview of Emerging Requirements for Data Management of Federally Funded Res...
PPTX
Protecting Private Data: Research Data, Data Sharing, and Privacy
PPTX
Overcoming obstacles to sharing data about human subjects
PPTX
Data as a research output and a research asset: the case for Open Science/Sim...
PPTX
RDAP14: DataNet Federal Consortium Update
PPTX
RDAP 16: Perspective on DMPs, Funders and Public Access (Panel 5: DMPs and Pu...
PDF
Open Science and Higher Education in Africa – Status, challenges & opportunit...
PDF
How FAIR is your data? Copyright, licensing and reuse of data
Overview of Emerging Requirements for Data Management of Federally Funded Res...
Protecting Private Data: Research Data, Data Sharing, and Privacy
Overcoming obstacles to sharing data about human subjects
Data as a research output and a research asset: the case for Open Science/Sim...
RDAP14: DataNet Federal Consortium Update
RDAP 16: Perspective on DMPs, Funders and Public Access (Panel 5: DMPs and Pu...
Open Science and Higher Education in Africa – Status, challenges & opportunit...
How FAIR is your data? Copyright, licensing and reuse of data

What's hot (20)

PPTX
Standardising research data policies, research data network
PPTX
2013 ICPSR Data Services
PPTX
Frances Burton on sensitive data
PDF
Connected health cities
PDF
Data Policy for Open Science
PPTX
The African Open Science Platform/Susan Veldsman
PPTX
ESIP Federation: Community-Driven, Collaborative Governance - Carol Beaton Me...
PPTX
Journal research data policy update
PPTX
‘Good, better, best’? Examining the range and rationales of institutional dat...
PDF
Research Week 2014: Tri-council Open-Access Policies and Data Management Plan...
PPTX
A SWOT Analysis of Data Science @ NIH
PPTX
Making Biomedical Research More Like Airbnb
PDF
NIH BD2K DataMed model, DATS
PPTX
RDAP14: Maryann Martone, Keynote, The Neuroscience Information Framework
PPTX
HESA data, describing research activity and #REF2021
PPT
Joy Davidson “Data Management Planning: an introduction” SALCTG June 2013
PPT
Managing sensitive data at the University of Bristol
PDF
Borgman - Privacy, Policy and Data Governance in the University
PPT
Libraries, RDM and e-infrastructure requirements
PPTX
State of open research data open con
Standardising research data policies, research data network
2013 ICPSR Data Services
Frances Burton on sensitive data
Connected health cities
Data Policy for Open Science
The African Open Science Platform/Susan Veldsman
ESIP Federation: Community-Driven, Collaborative Governance - Carol Beaton Me...
Journal research data policy update
‘Good, better, best’? Examining the range and rationales of institutional dat...
Research Week 2014: Tri-council Open-Access Policies and Data Management Plan...
A SWOT Analysis of Data Science @ NIH
Making Biomedical Research More Like Airbnb
NIH BD2K DataMed model, DATS
RDAP14: Maryann Martone, Keynote, The Neuroscience Information Framework
HESA data, describing research activity and #REF2021
Joy Davidson “Data Management Planning: an introduction” SALCTG June 2013
Managing sensitive data at the University of Bristol
Borgman - Privacy, Policy and Data Governance in the University
Libraries, RDM and e-infrastructure requirements
State of open research data open con
Ad

Similar to Stewarding Big Data (20)

PDF
Forschungsdaten-Repositorien: Informationsinfrastrukturen für nachnutzbare F...
PDF
Branding the Stewardship of Big Data
PPT
Overview of Emerging Requirements for Data Management of Federally Funded Res...
PPT
Bloomsbury Conference
PDF
Stewardship data-guidelines- research information network jan 2008
PDF
Forschungdaten-Repositorien - Stand und Perspektive
PPTX
Why manage research data?
PPTX
From Data Sharing to Data Stewardship
PDF
Framework and Roadmap towards an Open Science Infrastructure/Simon Hodson
PDF
Open Science - Global Perspectives/Simon Hodson
PDF
Graham Pryor
PDF
BLC & Digital Science: Mark Hahnel, Figshare
PDF
Gradscicomm Day 2
PPT
Common Ground: a policy framework for open access to research data
PPTX
2013 DataCite Summer Meeting - Closing Keynote: Building Community Engagement...
PDF
I o dav data workshop prof wafula final 19.9.17
PDF
Pampel & Kindling: Repositorien für Forschungsdaten - Infrastrukturen für die...
PPTX
A coordinated framework for open data open science in Botswana/Simon Hodson
PDF
The current challenges of upgrading the infrastructure
PPTX
Introduction to Research Data Management
Forschungsdaten-Repositorien: Informationsinfrastrukturen für nachnutzbare F...
Branding the Stewardship of Big Data
Overview of Emerging Requirements for Data Management of Federally Funded Res...
Bloomsbury Conference
Stewardship data-guidelines- research information network jan 2008
Forschungdaten-Repositorien - Stand und Perspektive
Why manage research data?
From Data Sharing to Data Stewardship
Framework and Roadmap towards an Open Science Infrastructure/Simon Hodson
Open Science - Global Perspectives/Simon Hodson
Graham Pryor
BLC & Digital Science: Mark Hahnel, Figshare
Gradscicomm Day 2
Common Ground: a policy framework for open access to research data
2013 DataCite Summer Meeting - Closing Keynote: Building Community Engagement...
I o dav data workshop prof wafula final 19.9.17
Pampel & Kindling: Repositorien für Forschungsdaten - Infrastrukturen für die...
A coordinated framework for open data open science in Botswana/Simon Hodson
The current challenges of upgrading the infrastructure
Introduction to Research Data Management
Ad

Recently uploaded (20)

PDF
A novel scalable deep ensemble learning framework for big data classification...
PPTX
Modernising the Digital Integration Hub
PPT
What is a Computer? Input Devices /output devices
PDF
How ambidextrous entrepreneurial leaders react to the artificial intelligence...
PPTX
O2C Customer Invoices to Receipt V15A.pptx
PDF
A contest of sentiment analysis: k-nearest neighbor versus neural network
PPTX
OMC Textile Division Presentation 2021.pptx
PDF
Transform Your ITIL® 4 & ITSM Strategy with AI in 2025.pdf
PDF
A comparative study of natural language inference in Swahili using monolingua...
PPTX
Final SEM Unit 1 for mit wpu at pune .pptx
PDF
Enhancing emotion recognition model for a student engagement use case through...
PDF
TrustArc Webinar - Click, Consent, Trust: Winning the Privacy Game
PPT
Module 1.ppt Iot fundamentals and Architecture
PDF
Architecture types and enterprise applications.pdf
PDF
DP Operators-handbook-extract for the Mautical Institute
PDF
Getting Started with Data Integration: FME Form 101
PPTX
The various Industrial Revolutions .pptx
PPTX
MicrosoftCybserSecurityReferenceArchitecture-April-2025.pptx
PDF
1 - Historical Antecedents, Social Consideration.pdf
PDF
Video forgery: An extensive analysis of inter-and intra-frame manipulation al...
A novel scalable deep ensemble learning framework for big data classification...
Modernising the Digital Integration Hub
What is a Computer? Input Devices /output devices
How ambidextrous entrepreneurial leaders react to the artificial intelligence...
O2C Customer Invoices to Receipt V15A.pptx
A contest of sentiment analysis: k-nearest neighbor versus neural network
OMC Textile Division Presentation 2021.pptx
Transform Your ITIL® 4 & ITSM Strategy with AI in 2025.pdf
A comparative study of natural language inference in Swahili using monolingua...
Final SEM Unit 1 for mit wpu at pune .pptx
Enhancing emotion recognition model for a student engagement use case through...
TrustArc Webinar - Click, Consent, Trust: Winning the Privacy Game
Module 1.ppt Iot fundamentals and Architecture
Architecture types and enterprise applications.pdf
DP Operators-handbook-extract for the Mautical Institute
Getting Started with Data Integration: FME Form 101
The various Industrial Revolutions .pptx
MicrosoftCybserSecurityReferenceArchitecture-April-2025.pptx
1 - Historical Antecedents, Social Consideration.pdf
Video forgery: An extensive analysis of inter-and intra-frame manipulation al...

Stewarding Big Data

  • 1. Stewarding Big Data: Perspectives on Public Access to Federally Funded Scientific Research Data Big Data and Big Challenges for Law and Legal Information Georgetown Law Library January 30, 2013 William G. LeFurgy Library of Congress @blefurgy
  • 2. My Perspective on Big Data Stewardship • Realizing full potential from big data depends keeping it accessible over time • Accessibility depends on life cycle management, most especially preservation • Advocate for collaborative, distributed model • Understand that “stewardship” has a different meaning for many data creators
  • 3. White House RFI Input Instructive • Request for Information on Public Access to Federally Funded Scientific Research Data, Nov. 2011 • Interested individuals and organizations to provide recommendations on approaches for ensuring long-term stewardship and encouraging broad public access • Input provided to inform development of agency policies and standards for managing big data
  • 4. Summary of Responses • 118 individual responses – 50% from academic research departments, professional organizations – 35% from libraries, repositories and allied organizations – 10% from publishers and commercial organizations – 5% other • Excellent (unstructured!) data set to analyze current thinking on big data stewardship
  • 5. Top-Level Policy Recommendations • Remarkable degree of congruence among comments – Broadly allocate adequate resources for data stewardship – Extend a collaborative national digital stewardship infrastructure – Institute and enforce a data preservation mandate – Strongly encourage policies to support secondary use, respect for data • But… conflicted about IP, copyright, privacy
  • 6. Need: Resources • Funders to include money in awards for data stewardship • Need cost models, other guidance for estimating data life cycle costs • Allocate expanded resources to support national data repositories
  • 7. Need: National Digital Stewardship Infrastructure • Leverage current institutional efforts to define best practices, tools, services • Extend community of practice for data stewardship through collaborative action across disciplines • Develop a skilled workforce with data stewardship expertise
  • 8. Need: A Data Preservation Mandate • Incentivize grant applicants to make realistic plans for data – Stronger data manager requirements in application process – Tie future awards to demonstrated success with data stewardship – Enable direct support of PIs by data stewardship specialists
  • 9. Support: Secondary Use, Respect for Data • Broadly apply a citation mechanism for data sets (e.g., DataCite, DOIs) • Criteria for evaluating grant applications tied to secondary use of data • Give equal credit for publishing articles and data sets • Develop robust metrics to track data publication and use
  • 10. Muddled Picture for IP • Opinions diverge about role of copyright, patents, etc., in regard to research data – Commercial interests see IP as critical – Many data users favor Creative Commons or public domain approach – Data creators fall between these positions • A significant degree of concern raised regarding privacy in connection with IRB, personal data
  • 11. Next Steps • Two interagency working groups within the National Science and Technology Council reviewing recommendations • Groups will develop science agency policies for data dissemination and stewardship • Potential for major change, as policies may have association with funding from the Federal science agencies
  • 12. Websites Request for Information: Public Access to Digital Data Resulting From Federally Funded Scientific Research, http://ow.ly/ePB93 Your Comments on Access to Federally Funded Scientific Research Results, http://ow.ly/ePBb9 National Science and Technology Council, http://ow.ly/h87Li

Editor's Notes

  • #2: Thanks for having me here today. I’m going to do my best to give you an overview from the perspective of libraries and archives on keeping big data for scholarship and public policy.
  • #3: I like the term “stewarding” to sum up all the activities involved in acquiring, preserving and making available data sets. Stewarding is essential if we as a society are going to see the full potential from big data. It’s a pretty basic proposition: somebody must devote time and effort to keeping data and to helping users access it. If this doesn’t happen, data will be hard to use, scattered and even lost. There are two basic considerations here. Collecting organizations need to concern themselves with the full life cycle of data, from initial creation, through use, to “archiving,” to long-term preservation and access, and The job is bigger than any one organization can handle; the volume and complexity of data require many organizations to work together in new ways.
  • #4: I thought a good way to frame this discussion would be to summarize what a variety of organizations said in response to a recent White House request for information. This request asked for input about ensuring stewardship and encouraging broad public access to federally funded scientific research data. The White House will use the information submitted to draft revised agency policies in connection with big data. This has huge potential. The revised policies could cover requirements for data management tied to billions in funding from the National Science Foundation, the National Institutes of Health and other funding agencies.
  • #5: The White House says they received 118 individual responses, all of which are made available on their website. There’s an interesting mix of respondents. Half came from discipline-specific academic research departments or professional organizations. I’d characterize them as data creators and data users. About a third of the submissions came from libraries, archives and other collecting entities. The rest came from a mix of individuals, publishers and commercial organizations. What we have here is an excellent data set that offers a broad-based snapshot of current thinking on data preservation. The response data set is seriously unstructured, as it is made up of randomly formatted textual documents, but it fairly easy to analyze.
  • #6: I was pleasantly surprised at the degree of congruence among the comments. Nearly everyone enthusiastically agreed that enhanced data stewardship was critical, both to support primary scientific research and broad secondary use by the public. Most submissions explicitly called for increased resources for data stewardship. There was heavy agreement that a distributed national digital stewardship infrastructure was the right vehicle for the infusion of new funding. Apart from money, the comments also aligned in calling for a strong data preservation mandate from funding agencies. The basic idea is that receipt of funding awards should be tied to a clear expectation for long-term data management. Many of you won’t be surprised to hear that there was much less agreement on traditionally thorny topics such as intellectual property, copyright and personal privacy.
  • #7: In terms of a push for increased resources, the comments clustered around three intentions. Individual funding awards should include a dedicated line item for data stewardship There is a need for models for projecting the lifetime cost of keeping data. Funding also needs to be channeled to a national infrastructure, most especially to support a distributed network of data repositories.
  • #8: The focus on a national data infrastructure zeroed in on ideas for extending work that’s already underway in terms of standards, tools, and best practices. There was enthusiasm for boosting the present community of practice for data stewardship, most particularly in a way that bridges different research disciplines. This really makes sense to me. While there is excellent work going on, much of it tends to reside in specialty silos. We need to accept that, at a certain level, data stewardship has a common set of requirements that are best addressed collectively. Related to this is a pressing need for a much expanded work force of data stewards.
  • #9: The need for a data preservation mandate comes down to what economists call “incentivizing.” In other words, if we want better data management, principal investigators have to be properly motivated. This motivation can come in different forms. Funding applications could call for detailed attention to data management. Evaluation of funding awards can be tied to prior demonstrated success with data stewardship And there could be provisions for data stewardship specialists to support PIs.
  • #10: There was strong support for what I characterize as “respect for data,” which is linked to recognizing the broad potential for secondary use. Ideas for enabling this include adoption of a citation mechanism for data sets, such as that offered by the DateCite organization. Related to this was the proposal to give the same credit for providing useful data sets as is now given for published articles. Securing this kind of credit depends on developing a new set of metrics to track data sets and their use.
  • #11: It’s no shock that consensus evaporated when it came to traditional hot-button issues. Commercial interests see control of IP as critical, while data users want relaxed IP barriers. Data creators fall between these endpoints—some want more stringent control, while others see the benefit of wider use. The issue of data privacy, most especially in connection with personal data collected under IRB rules, was strongly voiced by a number of creators, some of whom said that rules essentially barred any secondary use of certain data sets.
  • #12: In terms of next steps, the ball is in the White House’s court. Two interagency working groups are mulling over the comments and will use them to draft new policies governing data stewardship. As I noted earlier, we have the potential for major improvements in how federally funded data is kept and used. But I hasten to add that the outcome is still uncertain. What is clear, however, is that there is a strong consensus among data producers, users and keepers about what should happen.
  • #13: Here is a list of the websites I used in developing this presentation. Thank you.