Showing posts with label ANDS. Show all posts
Showing posts with label ANDS. Show all posts

Wednesday, November 21, 2012

Is the Thomson Reuters Data Citation Index Worth Paying For?

Greetings from the Hancock Learning Commons at the Australian National University in Canberra, where I am attending a presentation on the Thomson Reuters Data Citation Index. This attempts to add data to the traditional papers and books which academics get credit for producing. Thomson Reuters' offering was released a few weeks ago and ANU has been trying it out. Initially there were 2 million records in 69 repositories, including the Australian Data Archive (ADA). About half the current contents are from the USA, 42% from Europe. About half is from the life sciences.

What was not clear to me is what benefit the academic and research community gain by using Thomson Reuters' service. Presumably Thomson Reuters will charge money for use of their service. The information being indexed is almost all free open access material paid for by the public. It is not clear why academics should then pay Thomson Reuters for accessing free information.

The Australian Government funded Australian National Data Service the Australian Research Data Commons and similar free open access repositories are being linked up around the world.

If Thomson Reuters can add value to this, then the benefits they are offering need to be compared with what they propose to charge and a decision made if this investment is in the public interest.  It may not be a good use of public money for each university in Australia individually buy a subscription from Thomson Reuters.

Thomson Reuters collect up the metadata provided by repositories around the world and provide a global search facility to subscribers. The same service could be provided by others by harvesting this metadata.

The other service which is likely to be of more interest to academics, is that they harvest the citations of the datasets. Academics get hired and promoted partly on how many times their work is mentioned (cited) in published work. It is now possible to cite a dataset in the same way as publications, such as using APA. This will then increase the academics' citation ranking. If a Digital Object Identifier (DOI) is used, this makes the collection of the citations relatively easy.

Thursday, October 20, 2011

Benefits of Open Access to Data

The report "Costs and Benefits of Data Provision" (by John Houghton, September 2011) has been released. This is a 66 page (1.6Mbyte PDF) document, sponsored by the Australian National Data Service (ANDS). In one case, for the Australian Bureau of Statistics, the benefit of open access are estimated to be at least five times the cost.

As the reports was released in PDF, it is difficult to read on-line. Here is the table of contents and Summary of findings, converted to web format (diagrams omitted):
Contents

SUMMARY OF FINDINGS
  1. BACKGROUND AND CONTEXT
    1. ECONOMIC PRINCIPLES
    2. ECONOMIC EVIDENCE
    3. POLICY RESPONSES
    4. MEASURING VALUE, COSTS AND BENEFITS
  2. A FRAMEWORK FOR ESTIMATING COST-BENEFIT
    1. AGENCY COSTS AND COST SAVINGS
    2. USER COSTS AND COST SAVINGS
    3. EFFICIENCY AND PRODUCTIVITY IMPACTS
    4. WIDER ECONOMIC IMPACTS AND BENEFITS
    5. A COST-BENEFIT MODEL 12
    6. GUIDE TO DATA REQUIREMENTS 13
  3. PUBLIC SECTOR INFORMATION CASE STUDIES
    1. NATIONAL STATISTICS (AUSTRALIAN BUREAU OF STATISTICS)
      1. Agency costs and benefits
      2. User costs and benefits
      3. Indicators of use (web statistics)
      4. Wider impacts of use
      5. Summary of impacts
    2. SPATIAL DATA (OFFICE OF SPATIAL DATA MANAGEMENT & GEOSCIENCE AUSTRALIA)
    3. HYDROLOGICAL DATA (NATIONAL WATER COMMISSION & BUREAU OF METEOROLOGY)
  4. WIDER IMPACTS OF OPEN ACCESS TO PSI
    1. REPORTED IMPACTS OF SPATIAL DATA IN AUSTRALIA
    2. IMPACTS OF PSI MORE BROADLY
  5. LESSONS FOR THE RESEARCH SECTOR
    1. COST-BENEFIT ANALYSES OF RESEARCH DATA CURATION AND SHARING
    2. LESSONS FOR THE RESEARCH SECTOR AND NEXT STEPS
ENDNOTES AND REFERENCES
Figures and Tables

Summary of findings

Over the last decade there has been increasing awareness of the potential benefits of more open access to Public Sector Information (PSI) and the findings of publicly funded research. That awareness is based on economic principles and evidence, and it finds expression in policy at institutional, national and international levels.

Public Sector Information (PSI) policies seek to optimise innovation by making data available for use and re-use with minimal barriers in the form of cost or inconvenience. They place three responsibilities on publicly funded agencies: (i) to arrange stewardship and curation of their data; (ii) to make their data readily discoverable and available for use and re-use with minimal restrictions; and (iii) to forgo fees wherever practical.

This report presents case studies exploring the costs and benefits that PSI producing agencies and their users experience in making information freely available, and preliminary estimates of the wider economic impacts of open access to PSI. In doing so, it outlines a possibly method for cost-benefit analysis at the agency level and explores the data requirements for such an analysis – recognising that few agencies will have all of the data required.

There are many ways in which the provision of more open access to PSI can impact upon the costs faced by the government agency producers and the many existing and potential users of the information. This study focuses on three main elements:

  • The costs and cost savings experienced by PSI producing agencies involved in the provision of free and open access to information;
  • The costs and cost savings experienced by the users of PSI in accessing, using and re- using the information; and
  • The potential wider economic and social impacts of freely accessible PSI.
It is always more difficult to identify benefits than costs. Benefits may accrue in a variety of ways, including cost savings, efficiency gains, and new opportunities to create value through doing things in new ways and doing new things. These are, successively, more difficult to quantify: not least because they often emerge over time and can only be realised in the future.

An obvious approach is to begin with the most direct and directly measurable benefits, namely agency and user cost savings. Wider benefits are more difficult, and in some cases impossible, to measure. In this study, we explore impacts on consumer welfare and attempt to estimate the impacts of increased access and use, as measured by increased downloads, on returns to expenditure on data production.

While there are some one-off costs involved in the change to open access, most are recurring annual costs (e.g. agency IT and hosting costs, revenues foregone, etc.). Hence, both the agency and user costs that are modelled are annual costs, and the cost savings annual savings. In terms of the wider benefits of open access to PSI, returns to investment in data production are recurring annual returns, lagged and discounted over the useful life of the data – using a perpetual inventory method. Consequently, the cost-benefit comparisons presented in this study include annual agency and user costs and cost savings as well as the wider benefits arising from increased returns to annual expenditure on data production (Figure 1). They compare the costs and benefits at the time of the transition to open access (i.e. at the prices and levels of activity of
the time).

[Figure 1 omitted]

It is clear from the case studies presented that even the subset of benefits that can be measured outweigh the costs of making PSI more freely and openly available. It is also clear that it is not simply about access prices, but also about the transaction costs involved. Standardised and unrestrictive licensing, such as Creative Commons, and data standards are crucial in enabling access that is truly open (i.e. free, immediate and unrestricted).

For example, we find that the net cost to the Australian Bureau of Statistics (ABS) of making publications and statistics freely available online and adopting Creative Commons licensing was likely to have been around $3.5 million per annum at 2005-06 prices and levels of activity, but the immediate cost savings for users were likely to have been around $5 million per annum. The
wider impacts in terms of additional use and uses bring substantial additional returns, with our estimates suggesting overall costs associated with free online access to ABS publications and data online and unrestrictive standard licensing of around $4.6 million per annum and measurable annualised benefits of perhaps $25 million (i.e. more than five times the costs).

While data are more limited, there appears to have been an even more compelling case for making fundamental geospatial data freely available. Of course, the relative cost-benefits apply to the form of PSI involved and do not reflect in any way on the performance of the producing agencies. Some forms of PSI underpin major industries and contribute to their growth and prosperity. Other forms of PSI may have an important influence on policy decisions, but the economic impacts may be more limited and difficult to trace.

The publications and data arising from publicly funded research differ somewhat from other forms of PSI. Consequently, it is difficult to draw direct lessons for the research sector from the case studies explored in this report. Nevertheless, it is clear that many of the same issues arise when attempting to measure the value of the information and/or the costs and benefits associated with providing open access to it.

The evidence from previous studies suggests that individual cases vary greatly, making generalisation extremely difficult. Perhaps, what could more usefully be generalised are the methods of analysis. For example, it would be useful to combine the frameworks and models into a tool that could be applied in assessing the costs and benefits of research data curation and sharing, and to further develop the framework for estimating cost-benefits outlined in this study to produce a tool tailored to the analysis of the costs and benefits of providing open access to PSI. These tools might consist of a template for data collection, a draft questionnaire outlining the questions needed to elicit the necessary information, and a simple spreadsheet-based online model that people could use to perform a cost-benefit analysis. The models should include all possible quantifiable costs and benefits, but must also include qualitative issues to help to prioritise data preservation, access and curation projects (e.g. incorporate a balanced scorecard approach to weighing the more intangible benefits).

What this study demonstrates is that the direct and measurable benefits of making PSI available freely and without restrictions on use typically outweigh the costs. When one adds the longer-term benefits that we cannot fully measure, and may not even foresee, the case for open access appears to be strong. ...

From: "Costs and Benefits of Data Provision", John Houghton, prepared for the Australian National Data Service (ANDS), September 2011

Tuesday, March 30, 2010

Responsible Conduct of Research in Australia

Greetings from the Great Hall of the Austrlaian National University in Canberra, where a Research Data Workshop on the Australian Code for the Responsible Conduct of Research by the Australian National Data Service (ANDS). There has been recent controversy over the distribution of climate change data. ANDS has been set up to help Australian researchers collect even larger collections of data online, so it is timely to have a look at the ethics of this. There will be a second workshop tomorrow on the services which ANDS provides.

ANDS has produced a short guide "Research data policy and the Australian Code for the Responsible Conduct of Research". This suggests institutions review policies on: Intellectual property (covering copyright, moral rights, patent), Data management (Storage, Retention, Disposal, Access), Conflict of interest, Collaboration and contractual agreements, Ethics and privacy and Compliance. Many of these issues are covered in my lecture notes on Metadata and Electronic Data Management.

At question time I asked if ANDS would require organisations contributing data to indicate if they comply with the code. The reason for this is that ANDS, by referring people to data sources take on an ethical and legal responsibility for what is done with that data. Even if there is no black letter law requiring the use of the code, the fact that it exists is likely to be taken into account by a court or other body assign the actions of researchers. Given that ANDS has endorsed the code, it would be difficult for ANDS to claim that the code does not apply to them. It would not be possible to say that the data ANDS refers people to is not ANDS data and they have no responsibility for it: by referring people to data ANDS takes on obligations. One way to discharge those obligations might be to record if the organisation providing the data complies to the code or another code. Data uses could then make an informed decision as to if they should use the data.

The code itself (Reference No: R39 508kbytes PDF, 41 pages) was published in 2007 by the National Health and Medical Research Council (NHMRC), with help from the Australian Research Council and Universities Australia. There is also a Summary

Synopsis of publication:

The Australian Code for the Responsible Conduct of Research guides institutions and researchers in responsible research practices and promotes integrity in research for researchers. The Code shows how to manage breaches of the Code and allegations of research misconduct, how to manage research data and materials, how to publish and disseminate research findings, including proper attribution of authorship, how to conduct effective peer review and how to manage conflicts of interest. It also explains the responsibilities and rights of researchers if they witness research misconduct.

Developed jointly by the National Health and Medical Research Council, the Australian Research Council and Universities Australia, the Code has broad relevance across all research disciplines. It replaces the Joint NHMRC/AVCC Statement and Guidelines on Research Practice (1997).

Compliance with the Code is a prerequisite for receipt of National Health and Medical Research Council funding. ...

From: Australian Code for the Responsible Conduct of Research, NHMRC, 2007

Wednesday, August 26, 2009

Australian National Data Service

Greetings from the Australian National University in Canberra where Ian Barnes is giving a presentation on the Australian National Data Service to my e-commerce students. ANDS aims to make the data behind research in Australia accessible online. The services provided are both for human readers to search for data and then machine to machine web services to access the data.

One interesting question is why scientists would share their data. One reason would be that this will tend to result in the scientist's work being more widely cited and thus promoting their career. A less obvious reason is that by making the data available will help ensure the data is preserved for long term use, including by the original creator. Another reason is that this may be required by funding bodies to prevent academic fraud.

Interestingly some of the initial data in the ANDS system is research data about humpback whales near the location for the proposed north west multi-billion dollar gas platform.

The service provides a google map interface using the geotagging information in the metadata. The service uses a different metadata standard to the ISO19115/19139 standard which is used for some repositories, but there is provision for conversion of the data. The service uses the same OAI interface as used for electronic document repositories. The service also provides persistent identifiers.

The service has a harvesting process to search registered data sources to find new collections of data to index. The service uses RIF-CS based on ISO2146 to represent data in XML format.