About Me

My photo
I am Professor of Digital Humanities at the University of Glasgow and Theme Leader Fellow for the 'Digital Transformations' strategic theme of the Arts and Humanities Research Council. I tweet as @ajprescott.

This blog is a riff on digital humanities. A riff is a repeated phrase in music, used by analogy to describe a improvisation or commentary. In the 16th century, the word 'riff' meant a rift; Speed describes riffs in the earth shooting out flames. The poet Jeffrey Robinson points out that riff perhaps derives from riffle, to make rough.

Maybe we need to explore these other meanings of riff in thinking about digital humanities, and seek out rough and broken ground in the digital terrain.
Showing posts with label archives. Show all posts
Showing posts with label archives. Show all posts

6 September 2015

Big Data: Some Historical Perspectives



This was a contribution to a plenary panel at the European Policy on Intellectual Property conference organised by CREATe at the University of Glasgow in September 2015. In the nature of a short contribution to a panel on a wider theme, it barely scratches the surface of the possibilities implied i the title, but here it is for the record. 

It is already very evident from our discussions here that a distinctive feature of research into intellectual property is the emphasis on historical understanding. Petra Moser’s keynote yesterday was a wonderful illustration of how intellectual property researchers find historical data which has a wider cultural significance and is more than simply a lab for exploring different models of access. It may seem that, since big data is meant to present issues of scale and potential that we haven’t encountered before, historical perspectives won’t be particularly helpful. What I want to suggest here that we perhaps need to widen our historical terms of reference, and not restrict ourselves to precise historical precedents and analogies. 

Sometimes, big data requires big history (although we should also be aware of the caveat of Tim Hitchcock about the dangers of thinking exclusively on a large scale — we need both microscope and macroscope).

Let’s go back a long way, to 1086, when William the Conqueror, who had won the English crown at the Battle of Hastings twenty years previously, gave orders that a detailed survey should be undertaken of his English dominions. The Anglo-Saxon chronicle described how William sent his men to every county to enquire who held what land. The chronicler was horrified by the amount of information William collected: ‘So very thoroughly did he have the inquiry carried out that there was not a single hide, not one virgate of land, not even — it is shameful to record it, but it did not seem shameful for him to do — not even one ox, nor one cow, nor one pig which escaped notice in his survey’. Collating all this information required William’s clerks to develop innovative data processing techniques, as they prepared a series of summaries of the data, eventually reducing it to two stout volumes. The motives of William in collecting this information are still debated by historians, but the data was immediately put to use in royal courts and tax collection. Within a short period of time, this eleventh-century experiment in big data had become known as Domesday Book — the book of the day of judgement, from which there is no appeal — just as there can today be no appeal from the algorithms that might be used to set our insurance policy or credit rating.

Domesday Book was the first English public record and is a forcible reminder that anxieties about government data collection are nothing new. 2015 marks the anniversary of the grant of another celebrated English public document, Magna Carta. King John is remembered as the tyrant forced to grant Magna Carta at Runnymede, but his reign was also important because many of the major series of records recording government business began in his reign. John’s reign saw an upsurge in the use of technologies of writing and mathematics in the business of government. One important thread in the Magna Carta story is that it was both a reaction to, and at the same time an expression of, this growth in new technologies of government. It’s intriguing that Tim Berners-Lee and others have called for a Magna Carta for the world wide web to address issues of privacy and openness. There are a number of problems with this. One is, of course, that Magna Carta is linked to a common law system — it hasn’t even been adopted by the whole of Britain, as Scotland with its roman law system has always had a semi-detached relationship with Magna Carta. The other is that granting that the embedding of Magna Carta in English political life was a complex process, spread over several centuries and involving two civil wars.

In considering the issues of governance, ethics and identity posed by big data, this kind of longue durée approach can be very helpful. Jon Agar’s wonderful book, The Government Machine: A Revolutionary History of the Computer describes how the conception of the modern computer was influenced by the type of administrative processes developed by government bureaucracies in the nineteenth century which sought to distinguish between high level analytical policy work and routine mechanical clerical labour. Charles Babbage’s work was a sophisticated expressions of this nineteenth-century urge to identify and mechanise the routine. Closely linked to this urge to mechanise government was a concern, in the wake of the industrial revolution and the growth of population, to gather as much statistical information as possible about the enormous changes taking place. In a way, data can be seen as an expression of modernity. Another key big data moment was the 1890 United States census when the huge quantity of data necessitated the use of automatically sorted punch cards to analyse the information. Jon Agar vividly describes the achievements of this analogue computing and the rise of IBM. His account of the debates surrounding the national registration schemes introduced in wartime and the anxieties about linking these to for example employment or health records illustrate how our current concerns have long antecedents.

However, I think looking at big data concerns in this way does more than simply remind us that there is nothing new under the sun. It is also helpful in clarifying what is distinctive about recent developments and in identifying areas which should be policy priorities. First is the ubiquity of data. 
For governments from the eleventh to the twentieth century, data was something gathered with enormous clerical and administrative effort which had to be carefully curated and safeguarded. Data like that recorded in Domesday Book or records of land grants was one of the primary assets of pre-modern governments. Only large organisations such as governments or railroad companies had the resources to process this precious data — indeed one of the changes that is very evident is the shift in processing power, and perhaps we should be talking more about big processing rather than big data. Data was used in order to govern and was integral to the political compact. Now data is ubiquitous and comparatively cheap to acquire and process, this framework of trust no longer applies. Moreover, the types of organisations deploying data have changed. In particular, it is noticeable that the driving forces behind the development of big data methods have frequently been commercial and retail organisations: not only Google and Amazon, but also large insurance, financial and healthcare corporations. This is a contrast to earlier developments, both analogue and digital, where governments have been prominent and private sector involvement more limited.

The Oxford English Dictionary draws a distinction between the term big data as applied to the size of datasets and big data referring to particular computational methods, most notably predictive analytics. Predictive analytics poses very powerful social and cultural challenges, especially as more and more personal data such as whole genome sequences becomes cheaper and more widely available. How far can your body be covered by existing concepts of privacy? And is the likely future path of your health, career and life a matter of purely personal concern? In many ways, it is this idea of prediction which most forcibly challenges many of our most cherished social and cultural assumptions. Predicitive policing — an early contact by the police with people considered likely to commit crimes — is already being tested in some American cities. Predictivity almost dissolves privacy because it shifts the way in which we look at freedom of choice. It starts to become irrelevant as to what my reading or music choices are if they can be readily predicted from publicly available data. How we cope with a society in which many of our actions can be predicted is one of the chief challenges posed by big data. As my colleague Barry Smith, from the AHRC’s Science in Culture theme, has emphasised, the neuroscience surrounding predictivity — the way in which the brain copes with this predictivity — will become a fundamental area of research. As predictive analytics shades into machine learning, these questions will become even more complex, since we will start to see the distinction described by Agar between analytic work and routine labour breaking down in large organisations, posing major social and cultural challenges.

Finally, it is worth noting that generally the most important large data sets (censuses, tax records) have been about people, but increasingly big data will become about things. For example, machine tools frequently have sensors attached to them which enable the state of the tools to be monitored remotely by the manufacturer. This might encourage the manufacturer to monitor use of their products by clients in ways that could have commercial implications. The monitoring of medical implants will raise even more complex issues. A hint of the kind of complications that these developments might raise was given in the concurrence of Justice Alitto in the US Supreme court judgement in US v Jones 2012, which concerned the use of GPS tracking devices by police. Struggling to imagine how the framers of the US constitution would have viewed such devices, and imagined the analogy of ‘a case in which a constable secreted himself somewhere in a coach and remained there for a period of time in order to monitor the movements of the coach’s owner’.
For the Anglo-Saxon chronicle complaining about Domesday Book, the objects of the king’s greed were evident: land and animals; our future anxieties may be very different because our chief anxiety may be about objects linked to us in much more distant and complex ways.

Read more »

2 February 2014

Dennis the Paywall Menace Stalks the Archives


A story that hit the news this week was a report that the Prime Minister David Cameron is distantly related to the comedian Al Murray, and that they both had ancestors who worked for the East India Company. This news item was part of the publicity for the release online of 2.5 million genealogical records which are part of the India Office Records in the British Library. The digitisation of this huge historical archive is at first sight exciting news, but there’s a catch. The digitisation was undertaken by the family history company findmypast and you can only access these records via a subscription. A full World subscription to findmypast costs over £150, more than a television licence or a premium subscription to Spotify, although admittedly Pay As You Go credits are also available. Fortunately, many public and other libraries offer access to family history services such as findmypast, but this doesn’t fully address the profound issues of ethics, access and public ownership of archives posed by the activities of findmypast and other similar firms.

Among other new records recently made available by findmypast are the Rate Books for Westminster and Southwark. I am currently undertaking some research relating to houses in Westminster, so I immediately went to the National Library of Wales in Aberystwyth to use their findmypast subscription to check the new records. Rate books are one of the fundamental sources for the study of all aspects of the history of localities in England, and are not just of use for genealogy. One of the annoying things about packages like findmypast is the way in which they assume I’m only interested in my relatives (a matter of surpassing little interest to me), so that accessing other information in the documents available via these packages can be very awkward. As I’m working on some streets in Westminster, what I ideally would like to do is browse images of the relevant sections of the rate books. Although the presentation in findmypast is dreadful for this kind of wider research, I nevertheless found the images are there and that I could browse the sections I want. Except that the NLW library subscription provided by findmypast does not offer access to images or transcripts. The NLW subscription to findmypast, in a kind of digital dance of the seven veils, gave me lists of the names of people who owned property in the streets I was interested in, but when I went to check the images, said I would need to give money to findmypast to learn more. All I was presented with at the National Library of Wales was an extended advert which inevitably resulted in the request that I take out a subscription.

The subscription structure of findmypast is quite obscure, offering basic levels of subscription, then requiring the purchase of additional credits to undertake such everyday research tasks as viewing an image of the document.  I'm not sure at the moment whether the library subscriptions offered by libraries such as The National Archives at Kew or The British Library in London offer users free access to the images, but the description of the 'findmypast.co.uk Community Edition(TM)' on the corporate website doesn't encourage optimism.

Findmypast is a subsidiary of the Dundee-based firm D. C. Thomson, which I had hitherto thought of merely as the benign publisher of children’s comics such as the Dandy and the Beano and home of comic creations such as Dennis the Menace and Desperate Dan (although the founder of the firm, David Coupar Thomson, was notorious for his refusal to employ trade unionists or Roman Catholics). Findmypast is part of Brightsolid, the IT division of Thomson. Thomson has great hopes that its family history activities will offset the steep decline in profits from its newspaper and other conventional publications, and has recently reorganised Brightsolid, establishing D. C. Thomson Family History, in order to build its presence in this sector. This seems a reasonable business strategy – a report by Global Industry Analysts states that ‘genealogical enthusiasts are spending between US$1000 to US$18000 a year to discover his or her roots. The growth of the genealogy research market is being spurred by the spending of over 84 million genealogists’. (I would have linked to the original report but it costs $1450). It is estimated that the family history sector overall as a business is worth $84 billion dollars. According to Business Week, ‘genealogy ranks second only to porn as the most searched topic online’. Family history is big business, and the new CEO of D. C. Thomson Family History, Annelies van den Belt, a former TV executive, has declared her intention of ensuring that D. C. Thomson becomes a ‘truly global digital family history business’.

I suppose I would wish D. C. Thomson well in moving on from Dennis the Menace to history, if it wasn’t for the fact that it involves the theft of public cultural property. D. C. Thomson see partnerships with organisations like the Imperial War Museum, The National Archives, the British Library and The Scottish National Archives as their strong suit in the battle with American behemoths such as Ancestry.com. That means that it is our access to our archives that is being traded to help shore up Thompson’s profits. The argument in favour of a commercial approach to the digitisation of the India Office Records is that bodies like the National Archives and the British Library can’t afford to undertake digitisation on this scale themelves. But if digitisation is locked up behind high paywalls, then it is not a very useful activity. Instead of increasing access, subscription services limit access to those social categories (white, retired, middle class) who can afford comparatively expensive leisure activities. The justification offered by the British Library Press Office that the site can be accessed freely in the British Library Reading Rooms seems to miss a lot of the point of digitisation. To make matters worse, the determination of companies like D.C. Thomson to milk the genealogical market for all it is worth restricts the research use that can be made of the online records by locking them into the narrow types of search required by family historians.

The problem is not simply paying for access to this material, but also the enormous damage that is being done to public and scholarly understanding of history and culture by the resulting digital divides. In Britain, university access to digital resources depends on licensing deals secured by the excellent work of JISC Collections which allow university libraries to acquire packages like Early English Books Online, Eighteenth Century Collections Online and the Burney Newspaper Collections at very reasonable levels which ensure that most universities can afford them. As a result, an accepted canon of scholarly electronic resources has developed, supplemented by major resources available as open access, such the Proceedings of the Old Bailey. However, online publishers specialising in family history, being in a highly competitive and profitable market, are apparently unwilling to strike such deals. A subscription to Ancestry is for most university libraries prohibitively expensive. As a result, university-based researchers give priority to the JISC-licensed resources over the records available via family history firms. The bizarre results that this can cause are apparent from the current situation with British nineteenth-century newspapers. While one tranche of nineteenth-century newspapers in the British Library is available via JISC Collections, the bulk of the historic newspapers have been digitised by Brightsolid and are only available via a subscription service. This means that scholars will inevitably (and for no good scholarly reason) privilege the material to which they have free access, thereby creating profound and unnecessary distortions and biases. In this way, paywalls are shaping and distorting scholarship by creating hierarchies in the availability of material and imposing new and unlooked for canonicities.

The British Library recently (and rightly) got a great deal of praise for making available as open access on Flickr one million images from its nineteenth-century books. But one has to question how seriously an institution is committed to open access when, just a month later, it releases such an important part of the national heritage as the registration records associated with British rule in India on a subscription-only basis, and in a form that is really only useful for genealogical research. It is difficult to overstate the devastating implications for future scholarship of the depredations of firms such as D. C. Thomson. Archival records such as rate books are the backbone of the study of English local history, but in the form in which they are presented online, it is very difficult to use them other than for the study of individual family members. It would be wonderful to see the Westminster rate books linked to the London Lives resource, to help further the potential of linked data to trace the lives of everyday eighteenth-century Londoners, but I fear that is unlikely to happen. (Some of the later Westminster Rate Books are linked to London Lives, but the coverage is less comprehensive than findmypast, illustrating once again the confusing and fragmented landscape that is being created by these commercial partnerships). A horrible vision of the future is the Scotland's People resource, which is run by D. C. Thomson Family History in partnership with the National Archives of Scotland. This offers free surname search of over 90 million records from such key series as births, marriage and death registers, wills and probate records, and valuation records (containing details of properties), but access to images of the record themselves is largely pay-as-you-go. The business model is presumably one here that was ultimately determined by the National Archives of Scotland (and thus the Scottish Government). Presumably the justification is that it would have been impossible to undertake such large-scale digitisation otherwise, but is digitisation in this way worthwhile? What is the point of digitising and then being able to undertake only the most basic research because of the cost? It seems as if archivists have been gripped by a mania to digitise as quickly as possibly, regardless of the implications for future scholarship of how this is done.    

It is this kind of development that makes me worry as to whether digital technologies will turn out to be a boon or disaster for scholarship. If we end up with the bulk of our archival records only available via the expensive and cumbersome route offered by firms like findmypast, digitisation might prove to be the greatest disaster for scholarship of recent times. Melissa Terras in an excellent post has recently protested against the insistence by publishers on extracting processing charges for publishing books and articles on an open access basis. However, since we as scholars are the producers of those books and articles, the power to remedy this situation lies in our own hands. The decisions about the use of rapacious family history firms to digitise archives are more difficult for us to influence. Bodies like the British Library are funded separately from universities and are subject to different policy pressures. In the face of the enormous comercial possibilities of family history, the requirements of university researchers look puny. Yet surely we must protest against this enclosure of our cultural commons. We should also congratulate cultural institutions when they do make digital resources available on an Open Access basis. Although I couldn’t get very far with findmypast at the National Library of Wales, NLW has been a staunch standardbearer for the cause of Open Access. The excellent Welsh Journals and Welsh Newspapers projects are fully open access. Because of the NLW’s enlightened approach, Scottish students in Glasgow now study Welsh wills (freely available) rather than Scottish wills (locked behind a brightsolid paywell) – a lesson for the Scottish government to ponder there, surely.

In the meantime, I’m nevertheless pondering whether I need a subscription to findyourpast. Except of course that since I work in London, there is an alternative – I can just go to the very pleasant searchroom of Westminster Archives and consult the original Rate Books (or more likely microfilms) there. And, as the depredations of companies like D. C. Thomson continue, I think this is an alternative that many of us might be taking more and more in the future.

Read more »