Monday, 25 February 2013

How is a scarf like a dataset?

My teal blue feather and fan scarf

No, it's not a riddle! 

It struck me recently that there's lots of parallels one can draw between the act of creating and describing a dataset and the act of  hand knitting something. (Bear with me on this - it'll make sense, honest!)

The picture above is my scarf. I'm very fond of it. I knitted it myself, and it's warm and comfortable and goes well with a lot of my clothes.

When you're hand knitting a scarf, you take a ball of yarn, and you cast on stitches to make a row, then you keep adding rows until you run out of yarn, the scarf gets to the right length, or you get fed up with knitting.

The yarn in a ball doesn't contain any information or structure, but by the act of putting stitches into it, you're encoding something. In the case of my scarf above, it's a repeating pattern called feather and fan stitch , but it can just as easily be another pattern, or no pattern at all. If you wanted to get really fancy, you can encode all sorts of information into a knitted item - the most famous example of this is Madame Defarge in Dicken's "A Tale of Two Cities", knitting the names of the upper classes doomed to die at the guillotine into a scarf.

(Pushing the analogy a bit far, each stitch could represent a bit in a dataset, with a knit stitch signifying a zero and a purl stitch a one, but in this case that's not so helpful, as I've got yarn overs and knit-two-togethers as well as knit and purl stitches in there.)

My scarf was created by a process of appending- each new row got added to the previous, like a dataset where each new measurement gets appended on to the previous one to make a time series. The scarf has a fixed number of stitches in each row, the same as a dataset where a fixed number of measurements are taken each day. This doesn't have to be the case, I've seen plenty of patterns for scarves out there with variable row lengths. It all depends on the look you want it to have, or what the knitting is supposed to be - you have variable row lengths to shape the sleeves of a jumper, for instance.

Sometimes my data got corrupted. I dropped a stitch, or miscounted the number of knit-two-togethers that I needed to do, and came out with the wrong number of stitches at the end of the row. Usually when this happens you have to pull out the stitches until you get back to the place where you can fix the mistake, and then re-knit the rows you've pulled apart. It can get a bit annoying, especially when you're ripping out perfectly good rows to fix a mistake you hadn't spotted before, which is several rows (and possibly hours of knitting time) below.

I know for a fact that my scarf is not perfect. I've made mistakes there, and I'd feel really uncomfortable having someone scrutinise it and point out all my errors. Thankfully, no one's planning on peer-reviewing my scarf - though they would if I entered it into one of the knitting competitions you sometimes get at village fetes.

Like a dataset, I could have kept adding stitches and rows to my scarf ad infinitum, but there came a point when I actually wanted to wear it, so that meant I had to finish it off (i.e. cast off the stitches and sew in the ends). I could have used it while it was still being knitted (er... maybe as a pot holder, or a lap warmer?) but the knitting needles would have got in the way. It wouldn't have been ideal. Even if I had decided that I didn't want it to be a scarf after all, and was happy with it as a washcloth (a very sparkly one), I still would have had to have cast off and finished it properly, otherwise the first time I used it, it would have pulled apart into a big tangle of yarn. The same is true for datasets - if you're going to use them, you need them to be properly finished off - i.e. a firm definition of what pieces of data you are using, and what pieces you're not.

So, I finished my scarf/dataset, and I can now use it for the purpose for which it was intended - to keep my neck warm in a stylish yet comfortable way. Now what?

Well, I have a lot of scarves. So I need some way of identifying it, storing it, and maybe even reproducing it (when it wears out, or someone wants to make themselves one just like it). In other words, I need metadata about my scarf.

Descriptive metadata is easy. At a very basic level it's things like colour: "teal blue" and what it is: "scarf". But even with something this simple, you still need to have common language to make sure that the descriptors are understood. "Teal blue" makes perfect sense to me, but might not mean anything to someone else, who might think it looks a bit green.

Thankfully, there are other ways of describing the scarf. I can say that it's 200cm long, and 20cm wide, and that it was made from King Cole Haze Glitter DK (the type of yarn), colourway 124 - Ocean, with dyelot 67233. And all those last pieces of metadata, though too specific for general use, do describe the scarf accurately, though not completely, and makes a start at providing the information needed to recreate it.

For recreating the scarf, I need all the metadata about what yarn was used, but I also need the size of the needles I knitted it on (4mm). I need the pattern that I used (18 stitch feather and fan, with a 2 stitch garter stitch border at the edges). I need the number of stitches I cast on (54) and my tension (how tightly I knit in this pattern - 28 rows and 27 stitches for a 10cm by 10cm square). You don't need any of this information to wear the scarf, but it is important to keep it if you want to recreate it!

(As an aside, I didn't keep all the metadata about how I made the scarf and what yarn I used for it written down somewhere, which meant that when I came to write this post, I needed to work it out all over again. In other words, metadata should be collected from the start and stored somewhere safe, regardless of what it's describing!)

I then also need to make sure my scarf is stored correctly when I'm not using it, so it doesn't get lost, or (heaven forbid!) corrupted (i.e eaten by moths or shredded by mice). I also need to be able to tell people where it's stored, so that when I ask my other half to fetch it for me, I can also tell him that it's hanging on the door of my wardrobe.

I want to be able to cite my scarf when I'm talking about it. Mostly, I just do it by saying "my teal blue feather and fan scarf", to distinguish it from the other scarves I have hanging around the place. I could get fancy and assign it a KOI (a Knitted Object Identifier) but most of my handknits are sufficiently distinct that a casual glance can tell which is which from a short description!

And finally, because I've put a lot of time and effort into making my scarf, I'd like to get credit for doing that. Which, for me anyway, is covered when someone says to me "that's a nice scarf" and I can respond with "thanks! I made it myself" and a proud smile.

There's more to this analogy than the special case of one dataset/scarf being created by one single creator, but I'm sure I've bludgeoned you with enough knitting terminology already today. I'm sure I can stretch the analogy further, but that's something for another post!

I'll leave you with a challenge, to think about something you've made yourself, with your own two hands. It can be anything; a nice meal, a garden, a piece of clothing, a lego model, a painting, a piece of furniture. Something that you made yourself and are proud of. Got something?

The way you feel about that thing is the way that dataset creators feel about their data, especially if that dataset has been created through great effort and took a lot of time. Everyone wants acknowledgement and credit for the work that they do. Data creators are no different!








Tuesday, 29 January 2013

Data journals - as soon-to-be-obsolete stepping stone to something better?

Stepping Stones
Stepping Stones by mark 75 

During the PREPARDE project workshop at the International Digital Curation Conference, one of the presenters raised the thought that data journals may just be a temporary phenomena pending better data organisation and credit. (I'm paraphrasing from memory here, so forgive me if I get it wrong!) Their thinking was that we want to make data a first class scientific object and will do so through  data citation, and then we will also want to enhance existing scientific publications with links back to the data they use and associated interactive gubbins. So therefore, data journals, which publish the datasets and a brief paper describing them won't be needed, because you'll either cite the data directly, or have links in an analysis-and-conclusions article to the data.

I'm not arguing with the need for proper data citations, or the benefits they'll give. I also agree that analysis-and-conclusions articles will and should have better links to the data that underlies them. I do think though that there's an awfully big jump between a dataset, in a repository, ready to be cited, and a full analysis-and-conclusions article.

(A brief digression - I know we're piggy-backing on article publication to provide data creators with the credit they deserve for creating the datasets, and this is nowhere near the ideal way of doing it! But that's a subject for another post, so, for today, let's go with the whole data publication thing as a given.)

Let's start with direct citations of datasets. Ok, so you've created your dataset and you've put it in a repository somewhere, and cite it using a permanent id (DOI/ARK/whatever). Using that citation, another researcher can go and find your dataset where it's stored, and will have at least the minimum level of metadata given in the citation (Authors, Title, Publisher, etc.) What the user of the dataset doesn't get is any indication of how useful the data is likely to be (apart from what they can guess through their knowledge of the authors' and repository's reputation), and they may not get any information at all about whether or not the dataset meets any community standards, is in appropriate formats, or has extra supporting metadata or documentation.

This isn't a particularly likely situation for most discipline-based repositories, who have a certain amount of domain knowledge to ensure that community standards are met. But for institutional or general repositories, who may have to cover subject areas from art history to zoology, they simply won't be able to provide this depth of knowledge. So a data citation can easily provide the who, where and maybe the what of a dataset (who created it? where is it stored? what is it - or at least what is it called?) , but doesn't automatically provide any information on the how or the why the dataset was created - which is important for when it comes to judging the quality and reuse potential of the dataset. 

Looking from the other end, analysis-and-conclusions papers tend to be pretty long things, and they often have to describe a lot in terms of the methods used for the analysis. Having to explain the data collection and processing method before you even get to the analysis methods is a pain (even if you'd only have to do it once and would then cite that first paper), but is still an essential part of the paper if the conclusions are to hold up. 

Yes, it will be great to click on a graph and be taken to the raw data that created that plot, but you'd still need to provide metadata for that subset of the dataset (and most repositories only store and cite the full dataset, not subsets). Clicking through to a subset of the data doesn't give the whole picture of the dataset either, what if that particular data subset was cherry-picked to best support the conclusions drawn in the paper?  There's technical issues there, which I'm sure will be solved, but they aren't yet.

It's also about the target audience as well. If I'm looking for datasets that might be useful to me, I don't want to be trawling through pages of analytical methods to find them. Ditto if I'm interested in new statistical techniques, all the stuff about how the data was collected is noise to me. Splitting the publication between data article (which gives all the information about calibrations and instrument set-up and the like) and analysis-and-conclusions and citing the former from the latter seems sensible to me. Not to mention that it might work out quicker to publish two smaller papers than one large one (and would certainly be easier to write and review!)

So I really do think there's a long-term place for data journals, between data citation and analysis-and-conclusions articles. Data articles allow for the publication of more information about a dataset (and in a more human-readable way) than can be captured in a simple metadata scheme and a repository catalogue. Data articles also provide a mechanism for the dataset's community to judge the scientific quality and potential reuse of the dataset through peer-review (open or closed, pre- or post-publication). 

I think a data article is also a sign that the data producer is proud of their data and is willing to publicise it and share it with the community. I know that if I had a rubbish dataset that I didn't want other people using, but had been told by someone important that it had to be in a repository, then I'd be sure to put it somewhere with the minimum amount of metadata. Yes, it could still be cited, but it wouldn't necessarily be easy to use!

There's only one way to find out if data journals are just a temporary stepping stone between data and analysis-and-conclusions articles until data citation becomes common practice and enhanced publications really get off the ground. And that's to keep working to raise the profile of data citation and data publication (whether in a data article, or as a first class part of an analysis-and-conclusions article) so it becomes the norm that data is made available as part of any scientific publication. 

In the meantime, let's keep talking about these issues, and raising these points. The more we talk about them and the more we try to make data citation and enhanced publications happen, the more we're raising consciousness about the importance of data in science. That's all to the good!

Tuesday, 18 December 2012

SpotOn London


This post is a bit "better late than never", even though it's been just over a month since SpotOn London happened. And I'm not too sure what I can actually say about the conference, other than it was amazing, and there was cake!

Actually, no, I can say more than that. SpotOn was an unusual conference for me in that I'm used to the traditional academic conferences where you have people presenting their latest research, followed by a few questions from the floor, and then on to the next thing. SpotOn was a series of sessions which were almost completely panel discussions, where questions from the floor made up the vast majority of the conversation. Add to that format science communicators, tools developers, researchers and policy makers, and you've a potent mix of people to really get the conversations going.

The whole thing was kicked off by Ben Goldacre's gloriously chaotic talk about, um data and randomised trials and stuff, featuring such nuggets of information as that it's possible to buy Uranium off Amazon (but they won't ship it to the UK) and some gleeful choices of words that made me glad I wasn't drinking tea at the time I heard them. 

The second keynote was given by Kamila Markram talking about the publishing process, and drivers for open science, in the context of frontiersin.org - a combination publishing and social networking platform for scientists.

Both keynotes (in fact, all the sessions) were videoed, so I recommend going and having a listen.

I got tapped to sit on the Data Reuse panel, along with Mark Hahnel of Fighare and Ross Mounce, even though my voice was still a bit ropey. Gratifyingly, the session was standing room only, and we covered topics including open data, reuse, credit for making data open, data publication and citation, peer-review of data and impact of data. (If you don't have time to watch the video the storify of the session does a good job of capturing all the main points, and a few asides too!)

The other things that have stuck in my mind (a month or so later) include:
  • The Assessing Social Media Impact session, where we could tell we were making an impact because of the rapid number of spammers targeting the #solo12impact hashtag. (Storify here.)
  • Preaching to the choir in the Incentivising Open Access and Open Science: Carrot and Stick session, where there was plenty of talk about making other people do things to make science open, but precious little about how to do it yourself. I subscribe to the view that it's either important, so we should do it, or it's not, so we shouldn't and should stop talking about it! And with the whole carrot and stick thing - yes, researchers are not donkeys, but they are human, and we are herd animals! Lead by example! (Storify is here.)
  • The ScienceGrrl crowd - flying the flag for female scientists!
  • The fact that of all the badges being given out, the first one to completely disappear was the one saying "Data is the new black"
All in all, a really good conference for meeting new people and getting fired up about all sorts of really cool stuff. I'll be back next year!



Tuesday, 13 November 2012

When science and stories collide

The Story Collider

I'm back in the office today after a wonderfully intense couple of days at SpotOn London 2012 - which I'll be blogging more about in another post. 

But first, I want to talk about the Story Collider - the fringe event which kicked off the whole conference for me, which was held on the Saturday night in the upstairs room of a pub in Camden (not the usual location for scientific shennanigans, to be fair!)

I'm still not entirely sure how I wound up there, hiding at the back of the room, frantically reading and re-reading my notes. Well, yes, I do know how I wound up there. When the email came around to the registered conference attendees asking for storytellers, I took a look at it and thought "that could be interesting - I wonder if my story is appropriate?" And it went from there. The organsisers liked the sound of my story outline, and that was it. I was on the list to tell my tale.

(I was, of course, blithely ignoring the fact that I was due to vanish into the wilds of West Wales the weekend before the show. Oh, and the fact that my voice was somewhat on the croaky side, and not showing any signs of coming back...)

Anyway, the Story Collider is part stand-up comedy, part confessional, and aims to bring together people to listen to and to tell stories about science in their lives. Its format is simple, a half dozen storytellers, talking for about ten minutes each, standing alone on a stage in front of a microphone.

I think I can safely say it was one of the scariest experiences I've had in a long time. I'm no stranger to the stage, but there's a big difference between presenting research (where you can hide behind powerpoint slides and acronyms), or singing songs (where the words are already written and you know them by heart), to standing in front of strangers, telling them about something that actually, really happened to you, and how it made you feel. (The feelings part was the hardest!)

Be that as it may - I did it. I was shaking like a leaf when I got off that stage, but I did it! 

The audience was lovely - only a science crowd would have given me a cheer when I told them how I was finally going to get my dataset published. And I got a lot of laughs, and a lot of really nice comments afterwards too - the ones that stuck in my mind were the ones that said how nice it was to hear a story about the actual trials and tribulations of doing science.

The whole event was recorded, so I'm hoping there'll be podcasts of the show coming out in the not-too-distant future. I'd really like to listen to the other stories that were told that night again, as being second last in the running order meant that I was too distracted by being nervous to give them my full attention!

Many thanks to all the Story Collider organisers for giving me the chance to tell my story, and my fellow story-tellers and the audience for being so supportive, and for laughing and cheering! If you get the chance to go to a Story Collider event, or even talk at one, go for it!

One theme that kept coming back in the discussions at SpotOn London was how much we scientists need to get better at telling stories and talking to people. The Story Collider provides an excellent way of doing just that.

Citing Sensitive Data - workshop report

"Burned" DVD, microwaved to ensure total elimination of private data.
"Burned" DVD, microwaved to ensure total elimination of private data , by NightRStar

On the 29th October, I went to the British Library for a workshop on the topic of managing and citing sensitive data, one of a series of workshops all about data citation.

I won't go into the detail of what was said during the presentations as all the slides are available on-line here, and there's a good blog post summarising the workshop here.

I will take the opportunity to re-iterate what I said in my previous post about how citation doesn't equal open. Though I will expand on it further and say that there needs to be extremely good reasons for keeping data closed when public money has funded its collection (reasons along the lines of patient confidentiality, saving endangered species, etc, not "but I need extra time to write a paper!")

After all the presentations, we were split up into groups, and made to do some work, it being a workshop and all. First of all, we had to come up with some example scenarios for how to cite data given certain access conditions or embargos, and then we had to swap these with another group and try to solve them. This turned out to be a lot of fun, though I did somehow manage to wind up in the group that was threatening to fire people left, right and centre if they didn't behave!

The Yellow group were looking at access conditions for a study where different participants had given different levels of consent. The solutions they came up with were: 1) have an umbrella DOI for the whole dataset with multiple DOIs for the subsets with different access conditions. 2) Have a hierarchical DOI, or 3) have an umbrella DOI linking to subsets. The trade-off here was clarity versus nuance, and it was generally agreed that communities in different disciplines would have to decide the best approach. We also can't draw an inference on a subset of the data without taking the whole dataset into account.

The Red group were looking at embargoed data. First up was "researchers want to gain more research credit". Suggestions included: early deposit, while the embargo still is in play; access by request during embargo; DOI minted on deposit; open landing page in the repository (so people know the data exists, even if they can't access it yet) with end of embargo date on it; and the metadata should be specified on deposit too.

Next the Red group looked at the situation of longitudinal cohort studies which may change and have multi-layered embargoes. Access to variables could be dependent on layers of the dataset, with access to layers potentially increasing in time. The suggestion was to have multiple DOIs for multiple layers, with links between the landing pages to show how the layers fit together.

The Green group also looked at embargoes - specifically the situation where there was retrospective withdrawal of permission for a dataset and the data was embargoed while an investigation took place. (The assumption was that the DOI had already been minted for the dataset.) Suggested action was: retain the same landing page, but add text to it detailing the embargo and the expected date when the investigations would end (compliant with the institution's investigations policy). A user option to register to get notified when the dataset becomes un-embargoed would be a nice thing to have. When the investigation is complete, update the metadata depending on the results. And, at the beginning of the data collection, make sure that the permissions and data policy are set out clearly!

The Blue group were looking at access criteria, in two cases. Firstly was "White rhino numbers and GPS tracking information". The suggestions were: assigning a DOI to the analysed rather than raw data, and apply access conditions to the raw data so as to verify user credentials. The format of the public dataset could be varied, e.g. releasing it as snapshots instead of time series, or delaying the release of the dataset until the death of the tagged rhinos. Some of the rich descriptive data might also be kept back from the DataCite metadata store in order to protect the subjects.

The second scenario the Blue group looked at was animal experiments - medical testing on guinea pigs with photos and survival times. This one was noted as being difficult - though there was agreement that releasing data should be guided by funders and ethics committees. The metadata should not name individuals, and the possibility of embargoing data, or publishing subsets (without photos?) should be investigated.


In the general discussion afterwards it was (quite rightly!) pointed out that it's ok to cite and make available different levels of data (raw/processed) as raw data might well be completely incomprehensible to non-experts. We also had a lot of discussion about those two favourite topics in data citation - granularity and versioning. Happily enough, they'll be the subject of the next workshop, booked for Mon 3rd Dec. 

Friday, 19 October 2012

Why Citation does not equal Open

Open Means Never Having to Say You're Sorry
By cogdogblog. http://www.flickr.com/photos/cogdog/7155294657/in/pool-67039204@N00/ 

Recently I've had a few emails from people expressing concern about data licensing, especially when it comes to assigning DOIs to datasets so they can be formally cited. The assumption seems to be that if a dataset has a DOI assigned to it, the data must therefore be open. This isn't the case.

Citation and open data seem to have got tangled together. Yes, citation is a mechanism for encouraging researchers to make their data open, but it doesn't follow that everything you cite has to be open.

Let's take an example of a journal paper. You can cite a journal paper whether it's open or not, and the citation simply gives information about the paper and where you can find it. The DOI for the paper will take you to a landing page, and the landing page then tells you what restrictions are on the paper (if any). It's commonplace to cite a paper that you have to pay to access - I know I've done it many a time.

Similarly, say you want to cite the Book of Kells (Trinity College Dublin MS 58). That's easy - in fact I've just done it. But for the casual reader to access it, you'd need to travel to Dublin, go to Trinity College Library, pay €9 and look at whatever page happens to be open on display at that particular time. (I'm sure there are more stringent restrictions on researchers who actually want to be able to flick through the pages!)

So, there's plenty of precedent for researchers citing things that aren't open, or are restricted in some way. Data will be no different.

DataCite themselves have accounted for some situations where access to data might be restricted (because of confidentiality issues, embargo periods, etc.) in the publication year in the mandatory properties and also in the date element in the optional properties of the DOI metadata schema.

Publication Year “If an embargo period has been in effect, use the date when the embargo period ends. “

The landing page for a DOI-ed dataset needs to be completely open with the relevant information about why there is restricted access and/or what to do to get full access.

There's a planned JISC-British Library DataCite Workshops, focusing on managing and citing sensitive data, taking place on Monday 29th October in the British Library Conference Centre, which will look in greater detail at exactly these sorts of issues. Registration is still open!

For me, I want to spread the word that you can cite data without having to make it open. Open data is, of course, something to be encouraged wherever possible. But scientists are nervous enough about open data and the possibility of getting scooped, or having legal or IPR issues causing problems. Going for the softly, softly approach of citing data whether it's open or not will allow researchers to get used to the idea of data citation. Once they get credit for their work in creating the datasets, that's when we can show them how much more credit they can get for making them open.

And in a lot of cases, data needs to be restricted for very good reasons (for example protecting patient confidentiality). Penalising the researchers who created those datasets by not allowing them citations because their data can't be make open just seems unfair.

Wednesday, 27 June 2012

DataCite Summer Meeting June 2012


Sand castles by experts in Copenhagen


This is the last one of these posts - as it's the end of my notes from the Talinn/Copenhagen trip. Unfortunately it wasn't the last of the meetings I had to go to; the final one was a CODATA working group on data citation report drafting meeting, which doesn't have any presentation notes, but meant I missed the second half of the DataCite meeting.

Anyway, notes from the DataCite Summer Meeting presentations I did get to see are below: