Tuesday, 18 December 2012

SpotOn London


This post is a bit "better late than never", even though it's been just over a month since SpotOn London happened. And I'm not too sure what I can actually say about the conference, other than it was amazing, and there was cake!

Actually, no, I can say more than that. SpotOn was an unusual conference for me in that I'm used to the traditional academic conferences where you have people presenting their latest research, followed by a few questions from the floor, and then on to the next thing. SpotOn was a series of sessions which were almost completely panel discussions, where questions from the floor made up the vast majority of the conversation. Add to that format science communicators, tools developers, researchers and policy makers, and you've a potent mix of people to really get the conversations going.

The whole thing was kicked off by Ben Goldacre's gloriously chaotic talk about, um data and randomised trials and stuff, featuring such nuggets of information as that it's possible to buy Uranium off Amazon (but they won't ship it to the UK) and some gleeful choices of words that made me glad I wasn't drinking tea at the time I heard them. 

The second keynote was given by Kamila Markram talking about the publishing process, and drivers for open science, in the context of frontiersin.org - a combination publishing and social networking platform for scientists.

Both keynotes (in fact, all the sessions) were videoed, so I recommend going and having a listen.

I got tapped to sit on the Data Reuse panel, along with Mark Hahnel of Fighare and Ross Mounce, even though my voice was still a bit ropey. Gratifyingly, the session was standing room only, and we covered topics including open data, reuse, credit for making data open, data publication and citation, peer-review of data and impact of data. (If you don't have time to watch the video the storify of the session does a good job of capturing all the main points, and a few asides too!)

The other things that have stuck in my mind (a month or so later) include:
  • The Assessing Social Media Impact session, where we could tell we were making an impact because of the rapid number of spammers targeting the #solo12impact hashtag. (Storify here.)
  • Preaching to the choir in the Incentivising Open Access and Open Science: Carrot and Stick session, where there was plenty of talk about making other people do things to make science open, but precious little about how to do it yourself. I subscribe to the view that it's either important, so we should do it, or it's not, so we shouldn't and should stop talking about it! And with the whole carrot and stick thing - yes, researchers are not donkeys, but they are human, and we are herd animals! Lead by example! (Storify is here.)
  • The ScienceGrrl crowd - flying the flag for female scientists!
  • The fact that of all the badges being given out, the first one to completely disappear was the one saying "Data is the new black"
All in all, a really good conference for meeting new people and getting fired up about all sorts of really cool stuff. I'll be back next year!



Tuesday, 13 November 2012

When science and stories collide

The Story Collider

I'm back in the office today after a wonderfully intense couple of days at SpotOn London 2012 - which I'll be blogging more about in another post. 

But first, I want to talk about the Story Collider - the fringe event which kicked off the whole conference for me, which was held on the Saturday night in the upstairs room of a pub in Camden (not the usual location for scientific shennanigans, to be fair!)

I'm still not entirely sure how I wound up there, hiding at the back of the room, frantically reading and re-reading my notes. Well, yes, I do know how I wound up there. When the email came around to the registered conference attendees asking for storytellers, I took a look at it and thought "that could be interesting - I wonder if my story is appropriate?" And it went from there. The organsisers liked the sound of my story outline, and that was it. I was on the list to tell my tale.

(I was, of course, blithely ignoring the fact that I was due to vanish into the wilds of West Wales the weekend before the show. Oh, and the fact that my voice was somewhat on the croaky side, and not showing any signs of coming back...)

Anyway, the Story Collider is part stand-up comedy, part confessional, and aims to bring together people to listen to and to tell stories about science in their lives. Its format is simple, a half dozen storytellers, talking for about ten minutes each, standing alone on a stage in front of a microphone.

I think I can safely say it was one of the scariest experiences I've had in a long time. I'm no stranger to the stage, but there's a big difference between presenting research (where you can hide behind powerpoint slides and acronyms), or singing songs (where the words are already written and you know them by heart), to standing in front of strangers, telling them about something that actually, really happened to you, and how it made you feel. (The feelings part was the hardest!)

Be that as it may - I did it. I was shaking like a leaf when I got off that stage, but I did it! 

The audience was lovely - only a science crowd would have given me a cheer when I told them how I was finally going to get my dataset published. And I got a lot of laughs, and a lot of really nice comments afterwards too - the ones that stuck in my mind were the ones that said how nice it was to hear a story about the actual trials and tribulations of doing science.

The whole event was recorded, so I'm hoping there'll be podcasts of the show coming out in the not-too-distant future. I'd really like to listen to the other stories that were told that night again, as being second last in the running order meant that I was too distracted by being nervous to give them my full attention!

Many thanks to all the Story Collider organisers for giving me the chance to tell my story, and my fellow story-tellers and the audience for being so supportive, and for laughing and cheering! If you get the chance to go to a Story Collider event, or even talk at one, go for it!

One theme that kept coming back in the discussions at SpotOn London was how much we scientists need to get better at telling stories and talking to people. The Story Collider provides an excellent way of doing just that.

Citing Sensitive Data - workshop report

"Burned" DVD, microwaved to ensure total elimination of private data.
"Burned" DVD, microwaved to ensure total elimination of private data , bNightRStar

On the 29th October, I went to the British Library for a workshop on the topic of managing and citing sensitive data, one of a series of workshops all about data citation.

I won't go into the detail of what was said during the presentations as all the slides are available on-line here, and there's a good blog post summarising the workshop here.

I will take the opportunity to re-iterate what I said in my previous post about how citation doesn't equal open. Though I will expand on it further and say that there needs to be extremely good reasons for keeping data closed when public money has funded its collection (reasons along the lines of patient confidentiality, saving endangered species, etc, not "but I need extra time to write a paper!")

After all the presentations, we were split up into groups, and made to do some work, it being a workshop and all. First of all, we had to come up with some example scenarios for how to cite data given certain access conditions or embargos, and then we had to swap these with another group and try to solve them. This turned out to be a lot of fun, though I did somehow manage to wind up in the group that was threatening to fire people left, right and centre if they didn't behave!

The Yellow group were looking at access conditions for a study where different participants had given different levels of consent. The solutions they came up with were: 1) have an umbrella DOI for the whole dataset with multiple DOIs for the subsets with different access conditions. 2) Have a hierarchical DOI, or 3) have an umbrella DOI linking to subsets. The trade-off here was clarity versus nuance, and it was generally agreed that communities in different disciplines would have to decide the best approach. We also can't draw an inference on a subset of the data without taking the whole dataset into account.

The Red group were looking at embargoed data. First up was "researchers want to gain more research credit". Suggestions included: early deposit, while the embargo still is in play; access by request during embargo; DOI minted on deposit; open landing page in the repository (so people know the data exists, even if they can't access it yet) with end of embargo date on it; and the metadata should be specified on deposit too.

Next the Red group looked at the situation of longitudinal cohort studies which may change and have multi-layered embargoes. Access to variables could be dependent on layers of the dataset, with access to layers potentially increasing in time. The suggestion was to have multiple DOIs for multiple layers, with links between the landing pages to show how the layers fit together.

The Green group also looked at embargoes - specifically the situation where there was retrospective withdrawal of permission for a dataset and the data was embargoed while an investigation took place. (The assumption was that the DOI had already been minted for the dataset.) Suggested action was: retain the same landing page, but add text to it detailing the embargo and the expected date when the investigations would end (compliant with the institution's investigations policy). A user option to register to get notified when the dataset becomes un-embargoed would be a nice thing to have. When the investigation is complete, update the metadata depending on the results. And, at the beginning of the data collection, make sure that the permissions and data policy are set out clearly!

The Blue group were looking at access criteria, in two cases. Firstly was "White rhino numbers and GPS tracking information". The suggestions were: assigning a DOI to the analysed rather than raw data, and apply access conditions to the raw data so as to verify user credentials. The format of the public dataset could be varied, e.g. releasing it as snapshots instead of time series, or delaying the release of the dataset until the death of the tagged rhinos. Some of the rich descriptive data might also be kept back from the DataCite metadata store in order to protect the subjects.

The second scenario the Blue group looked at was animal experiments - medical testing on guinea pigs with photos and survival times. This one was noted as being difficult - though there was agreement that releasing data should be guided by funders and ethics committees. The metadata should not name individuals, and the possibility of embargoing data, or publishing subsets (without photos?) should be investigated.


In the general discussion afterwards it was (quite rightly!) pointed out that it's ok to cite and make available different levels of data (raw/processed) as raw data might well be completely incomprehensible to non-experts. We also had a lot of discussion about those two favourite topics in data citation - granularity and versioning. Happily enough, they'll be the subject of the next workshop, booked for Mon 3rd Dec. 

Friday, 19 October 2012

Why Citation does not equal Open

Open Means Never Having to Say You're Sorry
By cogdogblog. http://www.flickr.com/photos/cogdog/7155294657/in/pool-67039204@N00/ 

Recently I've had a few emails from people expressing concern about data licensing, especially when it comes to assigning DOIs to datasets so they can be formally cited. The assumption seems to be that if a dataset has a DOI assigned to it, the data must therefore be open. This isn't the case.

Citation and open data seem to have got tangled together. Yes, citation is a mechanism for encouraging researchers to make their data open, but it doesn't follow that everything you cite has to be open.

Let's take an example of a journal paper. You can cite a journal paper whether it's open or not, and the citation simply gives information about the paper and where you can find it. The DOI for the paper will take you to a landing page, and the landing page then tells you what restrictions are on the paper (if any). It's commonplace to cite a paper that you have to pay to access - I know I've done it many a time.

Similarly, say you want to cite the Book of Kells (Trinity College Dublin MS 58). That's easy - in fact I've just done it. But for the casual reader to access it, you'd need to travel to Dublin, go to Trinity College Library, pay €9 and look at whatever page happens to be open on display at that particular time. (I'm sure there are more stringent restrictions on researchers who actually want to be able to flick through the pages!)

So, there's plenty of precedent for researchers citing things that aren't open, or are restricted in some way. Data will be no different.

DataCite themselves have accounted for some situations where access to data might be restricted (because of confidentiality issues, embargo periods, etc.) in the publication year in the mandatory properties and also in the date element in the optional properties of the DOI metadata schema.

Publication Year “If an embargo period has been in effect, use the date when the embargo period ends. “

The landing page for a DOI-ed dataset needs to be completely open with the relevant information about why there is restricted access and/or what to do to get full access.

There's a planned JISC-British Library DataCite Workshops, focusing on managing and citing sensitive data, taking place on Monday 29th October in the British Library Conference Centre, which will look in greater detail at exactly these sorts of issues. Registration is still open!

For me, I want to spread the word that you can cite data without having to make it open. Open data is, of course, something to be encouraged wherever possible. But scientists are nervous enough about open data and the possibility of getting scooped, or having legal or IPR issues causing problems. Going for the softly, softly approach of citing data whether it's open or not will allow researchers to get used to the idea of data citation. Once they get credit for their work in creating the datasets, that's when we can show them how much more credit they can get for making them open.

And in a lot of cases, data needs to be restricted for very good reasons (for example protecting patient confidentiality). Penalising the researchers who created those datasets by not allowing them citations because their data can't be make open just seems unfair.

Wednesday, 27 June 2012

DataCite Summer Meeting June 2012


Sand castles by experts in Copenhagen


This is the last one of these posts - as it's the end of my notes from the Talinn/Copenhagen trip. Unfortunately it wasn't the last of the meetings I had to go to; the final one was a CODATA working group on data citation report drafting meeting, which doesn't have any presentation notes, but meant I missed the second half of the DataCite meeting.

Anyway, notes from the DataCite Summer Meeting presentations I did get to see are below:

NordBib conference notes, Copenhagen, June 2012


Inside the lecture theatre at the Black Diamond

The NordBib conference was all about Structural Frameworks for Open Digital Research
- Strategy, Policy & Infrastructure. I kind of fell into attending it by accident, as I was in Copenhagen for the OpenAIREplus workshop before it, and the DataCite meeting after it, so it seemed sensible to go to this one too.

It was an interesting conference, on one hand very high level, EU strategising, while on the other, the audience seemed to mainly consist of librarians and people interested in data without that much by way of actual concrete experience in data management. So I wound up having lots of conversations with lots of people, all interested in finding out what we in the UK and the NERC data centres have been up to.

All the presentations from the conference are available here (which kind of make my notes redundant, but nevermind!)

Monday, 25 June 2012

OpenAIREplus workshop - notes from the breakout session


One of Copenhagen's bridges being opened, so we can sail through!

1. Funders and data policy
 * Lots of interest in the data value checklist - compare UK and Australian data value checklist
 * It's cheaper to keep data rather than recreate it
 * Can you require open availability of data brought into a project? Case by case negotiation
 * Multiple funders - which data policy will be applied?
 * CODATA preparing a toolkit for funders about open data policy
 * Role of institutional repositories? Data centres are good places to handle data pools
 * Need clear metadata!
 * How to handle data management plans once the project is over? Fund data management post-project. Should remain institutional responsibility.
 * Identifiers - need researcher identifiers, funder acknowledgements, DOIs - all to pull together project information and data
 * Are there international approached in data management plans?

2. Institutional policy
 * Most institutions don't have a policy yet because they're not easy to create
 * What other steps need to be done before policy?
 * Hierarchy - who to get involved - academic champions
 * Broad overview - what are the needs of researchers -  don't want extra admin
 * Don't contradict other policies or legislation
 * Smaller institutions don't have monet or effort to get into big data infrastructures
 * policy can guide researchers on what to do with their data
 * What should be deposited, what should be kept
 * How can we help insitutions develop data management plans?
 * Guidelines on developing data management policies
 * What kind of questions do we need to know before drafting policy?

3. Researchers and publishers
 * Current examples are life and environmental sciences
 * we need other examples in other fiels
 * Researchers need acknowledgement for their work on data - not having it stops them shring
 * Quality issues are important - need principles for peer review of research data
 * Users of data are candidates to review it
 * there are varying degrees of openness in peer-review - which will be appropriate for data?
 * What stopes researchers sharing data? Quality, promotion, confidence in the value of the data
 * We can give researchers more confidence in their data by promoting community standards
 * Change beahviour so that data management is done every day, instead of just at the beginning and end of the project.
 * Publishers can influence researchers when it comes to data management.
 * Metrics are needed, data citation, but also alt.metrics
 * Need for good examples of data management to educate researchers
 * Need a list of trusted databases/repositories
 * URLs aren't trusted, because they break!

4. Technical
 * Finland is constructing a national data catalogue, containing a mix of metadata records and data
 * OpenAIRE data model and services are using trust levels for entities and (automatic and man-made) relations
 * Need to guarantee long term data availability for enhanced publications to be trustworthy, or at least know what bits will last for how long
 * Level of trust needed to develop services to show levels of preservation
 * Services should still exist for low trust objects - e.g.g use a robot to check if the object is still there, and if not, drop the connection.

5. OpenAIREplus
 * Are there licensing restrictions for metadata?
 * Case studies of scientific communities should be published as soon as possible
 * Credit for researchers is important
 * Libraries have a role too - even if there is a fear of data management
 * Universities are very disparate - makes it hard politically to agree on data policy.