Showing posts with label cloud computing. Show all posts
Showing posts with label cloud computing. Show all posts

Wednesday, 25 August 2010

Is the Cloud aware that it has 'the future of digital archiving in its hands'?

As anyone in the audience at the ECA conference in Geneva earlier this year will be aware, one of issues which I’ve been mulling over in recent months relates to roles and responsibilities in ‘the cloud’. The question I was asked to address in Switzerland was ‘in whose hands does the future of digital preservation lie?’ and my succinct response was: ‘Google's’. This was (for reasons evident in the paper I gave) meant both literally – given their increasing dominance of the cloud space but also metaphorically, as an encapsulation of all cloud service providers.

And certainly when my colleague, Doug Belshaw, pointed me in the direction of this post regarding Facebook’s archiving policy it became clear that I’m not the only one thinking about the (unintended?) consequences for all parties of where this might lead us.

Its tempting to see things only from our (by that I mean the archival community) side of the fence – to lament the inevitable decline in our future professional role that the handing over of content to commercial external service providers for its long term preservation will entail and to worry about what it may mean for the archives (and their users) of the future.

But maybe we should also pause to reflect on what it may mean for these service providers themselves and whether they actually have as much concern about the implications of this new found responsibility on their side as we do on ours.

For as I concluded my paper in Geneva:

"Perhaps we should actually stop to ask Google and their peers whether they are indeed aware of the fact that the future of digital preservation lies in their hands and the responsibilities which comes with it and whether this is a role they are happy to fulfil. For perhaps just as we are in danger of sleepwalking our way into a situation where we have let this responsibility slip through our fingers, so they might be equally guilty of unwittingly finding it has landed in theirs.

If so, might this provide the opportunity for dialogue between the archival professions and cloud based service providers and in doing so, the opportunity for us to influence (and perhaps even still directly manage) the preservation of digital archives long into the future".


To again quote from the conclusion of my paper:

"Maybe the interconnection of content creation and use and its long term preservation need not be as indivisible within the cloud as it might first appear. Yes Google’s appetite for content might appear insatiable, but that does not necessarily mean that they wish to hold it all themselves – after all, their core business of search does not require them to hold themselves every web page they index, merely to have the means to crawl it and to return the results to the user. Might we be able to persuade them that the same logic should also apply to the contents of Google Apps, Blogger, YouTube and the like? If so, might the door be open for us, the archival community through the publicly funded purse to create and maintain our own meta-repository within which online content can be transferred, or just copied, for controlled, managed long term storage whilst continuing to provide access to it to the services and companies from which it originated?

That way they get to continue to accrue the benefit of allowing their users to access and manipulate digital content in ways which benefit their bottom line, the user continues to enjoy the services they have grown accustomed to and the archival community can sleep soundly, safe in the knowledge that whilst service providers are free to do what they want with live content, its long term preservation and safety continues to lie in our own experienced and trusted hands".


I wonder if such dialogue is already occurring between Google, Facebook et al and the likes of NARA, NAA and TNA. Lets hope so…

Sunday, 2 May 2010

Is digital preservation now routine?

It’s been a while since I attended a conference specifically themed around digital preservation / electronic archiving and having spent a few days last week in Geneva at the excellent European Conference on Archiving I was struck by the change. Not many years ago such conferences were dominated by debate about the technical complexities it posed, about the relative merits of competing theoretical approaches such as emulation and migration and the risks we faced if and when we got it wrong. The fragility of digital media was stressed and compared unfavourably to the durability of their traditional counterparts (encapsulated by seemingly endless comparisons between the original Domesday Book and its 1980s electronic equivalent).

I heard none of this at ECA, at least not in the sessions I attended or the conversations I was party to. Instead there were plenty of case studies from around Europe of organisations who are quietly and successfully getting on with it. On the evidence of the past few days we seem to have found ourselves in the situation where our ability to actually preserve this stuff indefinitely and to continue to provide access to it seems, without much triumph or fanfare to now be taken as read. This is not, of course, the same as saying that no more problems or challenges exist, but they seem to be of a more prosaic, ‘routine’ nature revolving around the need to secure budgets and improve the user experience etc.

More interestingly still if a single concern dominating the conference can be identified it seemed to be one related to the volume of information being created and stored today and estimated to be created tomorrow. I lost count of the number of presentations which contained jaw-dropping predictions of the amount of data soon to be at our fingertips and the challenges this will pose in terms of resource discovery, legal discovery and overall management. But interestingly virtually never its preservation. So, based on the evidence of this conference alone, it seems as though within a few short years we have jumped from a situation where we used to worry obsessively that we were in danger of losing everything to one where we now stress about how to manage a world where we will lose nothing, with barely a pause for reflection on this change.

My other observation was that (my own meagre contribution to one side) there was little or no discussion about the growing impact of the web as a storage ‘repository’ as heralded by the rise of ‘the cloud’. The unspoken assumption behind most of the debate and the projects and initiatives they represented seemed to be that these organisations will always have physical control of the electronic information they wish to preserve both prior to and after its ingest into their electronic archive, but as I tried to stress in my paper, I wonder how safe an assumption that will prove to be?

Wednesday, 24 February 2010

Information Management: the forgotten issue of the cloud

There was an interesting supplement on Cloud Computing from MediaPlanet within this Saturday’s Daily Telegraph (Ok I know it’s now Wednesday, but it takes me most of the week to wade through the weekend papers!).
The supplement - of which I can sadly find no English language online version - appears to be aimed at a senior management audience and is deliberately light on the technical detail, choosing to focus more on the benefits to the organisation which moving towards cloud-based computing can bring (institutional agility, flexibility and cost saving seem to be the main arguments in favour). It also includes ‘5 steps to making the most of cloud’ which are:

1. See the possibilities
2. Consider security
3. Use it to your advantage
4. Push the boundaries
5. Consider logistics

It would be hard to disagree with any of these, but its steps 2 and 5 which interested me most. For whilst steps 1, 3 and 4 (and, indeed, the rest of the content in the supplement) is designed to articulate the advantages and to push the potential it is these two steps which are designed to sound a note of caution and to instil the need for a cautious, managed approach to the management of the risks involved.

But if you were to rely solely on this supplement for guidance you’d be mistaken for assuming that data security should be your only concern when adopting a cloud-based computing environment (especially as the ‘logistics’ which Step 5 encourages you to consider relate to issues of security and mobile devices so is, in effect, just an extension of Step 2: Consider Security).

Aside from a passing mention of data protection and the potential need for some organisations to keep certain data within ‘certain geographic boundaries’ (which I’m assuming is again essentially related to the requirements of the Data Protection Act) what is entirely missing is an appreciation of the information management implications of moving data to the cloud. There is no acknowledgement of the need to ensure that current levels of record and information management control, say in relation to resource discovery or retention, must be continued into the cloud; nor any recognition of the potential problems of ensuring that this is so.

Interestingly, some of the issues which may come to the surface if these concerns are ignored are obliquely and inadvertently acknowledged – for example the point is made that in the cloud you pay as you consume, but the point is not expanded to its logical conclusion that it therefore pays to know exactly what information you still need to store (and pay for) and what can safely be destroyed. Likewise, the point is made that one of the biggest advantages foreseen for the ‘G Cloud’ (the UK Government Application Store which is currently being trialled) “could be allowing departments to share non-sensitive data so paper work is reduced and processes sped up” but no consideration is given as to how ‘sensitive’ and ‘non-sensitive’ data might be appropriately identified and controlled within the cloud.

On a more positive note Mark Taylor from Microsoft draws attention to the need for increasing standardisation so that the cloud ‘runs along the same principles and business models no matter who is managing the hosting’. Might the development of such standards and interoperability offer a potential means by which a single management layer can be placed on top of the cloud to allow organisations to consistently manage their information wherever it happens to reside in the cloud? And in doing so might it help address some of the management information issues which this supplement failed to acknowledge?