This week’s Economist has an excellent special report on managing information entitled ‘Data, Data everywhere’. It looks at the changes, opportunities and challenges posed by our new found ability to create and manipulate vast quantities of data – big data. There are lots of impressive/daunting (depending on your point of view) statistics about just how much data we are now talking about (40 billion photos on Facebook for example) during this “industrial revolution of data”. It also explores the concept of ‘data exhaust’ the trail of clicks which users leave behind them and which Google and others have been able to put to such incredible use: from search to speech recognition and from spell checking to language translation. All made possible not by attempting to training computers the rules which determine how these concepts work, but instead by tracking the activities of billions of user transactions which do the work of refining, correcting and adding relative value to words. Those who have heard my ‘Meet the future of Records Management: Amazon.com’ conference paper will know that I have long suspected that we could and should be making use of this exact same ‘exhaust’ to help us manage information, as well as profit from it - what I describe as 'Automated Records Management' (See also Records Management Journal Vol 19 No.2 2009 for a paper I wrote on this entitled 'Forget Electronic Records Management its Automated Records management that we desperately need')
There’s also interesting stuff in the Economist supplement on the problems of how to make sense of all this data, including new ways of visualising it and the prediction that statistics will soon be one of the coolest jobs around(!). It also makes some interesting points about the need for management to be trained in how to make sense of all this data. This chimes with a conversation I had with a Chief Exec a few weeks ago who also made the case for ensuring that senior management were aware of good old fashioned archival concepts such as provenance and context to give them a better appreciation of what the data they are looking at is actually telling them or how much it can be relied upon (rather than what they wish it was telling them and how much faith they may wish to place in it).
To give the Economist its due it does also look beyond the potential and address some of the challenges (and not just in relation to security – see previous post). Admittedly it does appear a little confused about the subject of data retention stating that ‘current roles on digital records state that data should never be stored for longer than necessary because they might be misused or inadvertently released’. It then goes on to state that ‘in future it is more likely that companies will be required retain all digital files, and ensure their accuracy, rather than to delete them’ – a vision of the future likely to strike fear into every records managers heart. There are some immediate flaws obvious in this logic (in the EU at least) where current data protection laws prevent this in relation to personal data, and elsewhere the Economist itself draws attention to the problems that storing such massive amounts of data is causing to the existing technical and resource infra-structure that Google et al rely on which would seem to favour a more selective approach to data retention on pragmatic grounds if nothing else. But whether such concerns are considered enough to stop the ‘lets keep and exploit everything’ bandwagon which lies behind much of this report is at best debatable and at worst, I suspect, distinctly unlikely.
The world is changing fast. Changes in technology are having a profound effect on the role of records management. The purpose of this blog is to give records managers and others interested in this area a 'heads up' as to what these changes might mean and how the profession needs to adapt to keep pace and maintain its relevance in the years ahead.
Showing posts with label retention. Show all posts
Showing posts with label retention. Show all posts
Friday, 5 March 2010
Wednesday, 30 January 2008
How to keep those servers cool!
There’s an interesting article in the latest edition of Government Computing magazine (unfortunately there is no e-version of the article to link to). The article in question is entitled ‘Greening the data centre’ and refers to the problems that organisations are encountering in terms of rising energy costs resulting from the ‘hot and hungry’ new blade servers that organisations are cramming into their data centres. According to the piece it is predicted that between 2000 and 2010 we will have "installed six times the amount of servers in our data centres and 69 times the amount of storage".
The problem is apparently that our data centre buildings are not designed to cope with the power required and heat generated by such machines, plus of course energy-consumption is now a political and ethical hot potato.
Now the interesting thing is the range of possible solutions outlined in the article. These vary from ‘better power management in data centres’, through to ‘taking your servers… and virtualising them’ or simply replacing old technology with new.
Nowhere does the rather obvious suggestion of ‘keeping less information’ get a look in. It would be interesting to know what percentage of the content of these steaming servers is actually still useful and still required? Of course the volume of information organisations create and need to retain is always increasing – but I bet there is still a huge percentage that could safely be destroyed if only anyone knew what it was, and whether it was still actually required…
But given that the main contributor to the piece is a Vice President at IBM perhaps its not that that surprising that the suggestion is to buy more kit, rather than to make better use of what exists already…
The problem is apparently that our data centre buildings are not designed to cope with the power required and heat generated by such machines, plus of course energy-consumption is now a political and ethical hot potato.
Now the interesting thing is the range of possible solutions outlined in the article. These vary from ‘better power management in data centres’, through to ‘taking your servers… and virtualising them’ or simply replacing old technology with new.
Nowhere does the rather obvious suggestion of ‘keeping less information’ get a look in. It would be interesting to know what percentage of the content of these steaming servers is actually still useful and still required? Of course the volume of information organisations create and need to retain is always increasing – but I bet there is still a huge percentage that could safely be destroyed if only anyone knew what it was, and whether it was still actually required…
But given that the main contributor to the piece is a Vice President at IBM perhaps its not that that surprising that the suggestion is to buy more kit, rather than to make better use of what exists already…
Thursday, 4 October 2007
Retention by format, not content?
It is one of the basic truisms of records management theory: that decisions regarding retention requirements are made based on their content and 'regardless of format'. This was an especially useful mantra a decade or so ago when some users would otherwise see the electronic version of a document as some alien construct, completely divorced from its paper counterpart.
Its still a mantra we cling to today and trot out as our first instinctive response, but its limitations are becoming more and more obvious, for example when it comes to email. Yes, its easy for us to say to our users that they must manage the contents of their inbox not as emails, but according to the content they contain but this is seldom reflected in reality. The simple fact is that the sheer volume of emails faced by users makes this virtually impossible to achieve and the majority of decisions taken regarding the fate of email are either taken on an individual ad hoc basis ("I don't think I need this any more") or en masse ("I've run out of space allocation so lets delete all last year's emails/all emails with large attachments etc").
The news that UK phone companies are now bound by law to retain information about all telephone calls and text messages for one year sounds a further death knell in the practicality of the 'regardless of format' concept. Even though this data might be used for one of three levels of enquiry the decision has been made that all such information must be retained for the same period: regardless of content, subject or any other criteria. If its information about a phone call it is kept for 1 year - and that is retention based purely on format and 'regardless of content'.
Its still a mantra we cling to today and trot out as our first instinctive response, but its limitations are becoming more and more obvious, for example when it comes to email. Yes, its easy for us to say to our users that they must manage the contents of their inbox not as emails, but according to the content they contain but this is seldom reflected in reality. The simple fact is that the sheer volume of emails faced by users makes this virtually impossible to achieve and the majority of decisions taken regarding the fate of email are either taken on an individual ad hoc basis ("I don't think I need this any more") or en masse ("I've run out of space allocation so lets delete all last year's emails/all emails with large attachments etc").
The news that UK phone companies are now bound by law to retain information about all telephone calls and text messages for one year sounds a further death knell in the practicality of the 'regardless of format' concept. Even though this data might be used for one of three levels of enquiry the decision has been made that all such information must be retained for the same period: regardless of content, subject or any other criteria. If its information about a phone call it is kept for 1 year - and that is retention based purely on format and 'regardless of content'.
Subscribe to:
Posts (Atom)