Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

24 January 2013

DNA as a new method of digital data storage

Quite the most exciting, extraordinary thing that I've read about recently: encoding information in DNA as a means of high density archival data storage. The Guardian has a news article about this: 'Shakespeare's sonnets encoded in DNA' and the original research is reported in an article in Nature 'Towards practical, high-capacity, low-maintenance information storage in synthesized DNA'.

24 January 2012

More thoughts about data

Is managing data a bit like designing a fountain, striking a balance between a contained flow and a deluge?

Last week I attended an event about data management, led by the Digital Curation Centre. There's also a new book out about this topic: Managing Research Data, edited by Graham Pryor and published by Facet.

The session I attended here in Sheffield focused particularly on the importance of effective data management planning and introduced us to the DMP Online tool. It offers an interesting way of building a data management plan, particularly for supporting applications for funding from research councils.

One of the issues about data management which seems to be given different emphasis in different contexts is the question of when not to share or retain data. Are all data equal or are some data more equal than others? For example, in my project I can imagine a possible future use for raw quantitative data from a survey I hope to carry out later this year. However, I don't think the same applies to qualitative interview data. I'm almost tempted to say that these data are just too unique: the process of carrying out the interviews and my subjective participation in them as a researcher is too significant and important an influence on them for them to be truly reusable by others. I also think that even carefully "anonymised" qualitative data, when taking the form of a full transcript, is rarely truly, fully "anonymous".

It's also interesting to think about this in the context of library collection decisions. How do we assess the potential use or usefulness of an item (or a dataset) before adding it to a collection? Would some version of a 20:80 rule apply to large data collections, as it is said to apply to print collections?

Another question which interests me is the division between different types of data. In November it was announced that more government data will be being made freely accessible including, controversially, the potential release of anonymised health data. There's more detail here, but it's interesting that the separation between "research data" and "data potentially useful for research" seems so clearly established. And what about the grey data generated by organisations which are neither public sector, nor research-led - perhaps including social enterprises?

01 March 2011

Digital Curation Centre Roadshow

Today I attended the first day of the DCC Roadshow Sheffield: Institutional Challenges in the Data Decade.

Amongst the participants, librarians were in the minority: only two (of eight) speakers were from libraries; the role libraries are expected or able to play in leading on this issue clearly varies considerably between organisations. The University of Sheffield's Director of Library Services, Martin Lewis, spoke of managing research data as one of the biggest professional challenges facing academic librarians. He and other speakers also emphasised the importance of close cooperation across different departments within an institution.

Dr Liz Lyon (DCC Associate Director) discussed the scale of the data challenge, mentioning a recent special data-focussed issue of Science (11/02/2011), and suggested three useful ways of thinking about data: The presentation also discussed the findings of the report Open science at web-scale regarding transparency in scientific research, the growth of "citizen science" (crowd-sourcing research tasks) and the use of data in predicting outcomes.

Key data management issues include:
  • developing data management policies (increasingly a factor in Research Council funding evaluations);

  • ethical issues in data sharing;

  • data storage (including cloud computing - and the recent HEFCE grant for development).

The presentations which followed provided five case studies of projects managing, using, or providing training about, data. Meik Poschen described the Manchester eResearch Centre MaDAM project working with biomedical researchers. Two presentations covered data management training programmes: Richard Plant discussed the University of York's DMTpsych training for psychology postgraduates and Stuart Jeffrey talked about the DataTrain programme for postgraduate archaeologists and anthropologists at the University of Cambridge. Mark Birkin discussed the NeISS (National eInfrastructure for Social Simulation) project, which facilitates data curation and data use for developing models to predict the impact of policies on populations.

Matthew Herring (Digital Library Officer, York) described different tools for data management in use at the University of York - including YODL (York Digital Library), a multimedia resource currently mainly containing images relating to the university's research. This presentation also provided the quotation which, for me, summed up the day and seems like a good way to finish this post: "openness rocks".