OPEN DATA TO IMPACT

By Cianan Clancy

From Open Data to Real Impact: What I Built at the Build for Ireland Hackathon

Ireland has more than 22,000 open datasets published by over 140 public authorities and bodies. They cover everything from company records and transport to the environment and public services.

That represents a substantial public resource. But publishing data is only the beginning.

How do we know which datasets are useful? Which need improvement? Which are helping people, supporting businesses or creating opportunities for new products and services?

For the Build for Ireland hackathon at Dogpatch Labs, I built a prototype to help answer those questions: an Open Data Impact Framework, applied across Ireland’s national open data catalogue.

The problem: we can count datasets, but understanding their value is harder

Open data teams need to decide where to focus limited time and resources. Public bodies need feedback on the data they publish. Developers, researchers and startups need to identify reliable data worth building on.

Yet the evidence needed to make those decisions is often incomplete or scattered.

A dataset might be available in a useful format and updated regularly, but say very little about who uses it or what difference it makes. Another might have significant potential, without any documented examples of reuse.

Without a consistent assessment method, it is difficult to compare datasets or decide which to improve, promote or investigate further.

The challenge is to connect publication with evidence of value.

The solution: a transparent framework for assessing every dataset

I used ChatGPT to help design a framework that assesses datasets against eight criteria:

  • Data quality and usability
  • Public value
  • Strategic alignment
  • Usage and reach
  • Economic and innovation potential
  • Operational value
  • Risk and ethics
  • Demonstrated social and environmental impact

Each criterion uses a one-to-five scale, with published weights that add up to 100%. Together, they produce an overall Open Data Impact Score.

The framework was informed by EU Open Data Maturity work and Spain’s approach to assessing open data impact.

Transparency matters here. A score should have an explanation: what evidence contributed to it, what is missing and what a publisher could improve.

What I built: applying the framework across Ireland’s catalogue

I then used Codex to build a prototype that reads data.gov.ie catalogue records, applies the scoring rules and creates a ranked dataset impact register.

The prototype processed 22,608 catalogue records in the snapshot used for the project.

It assessed information including titles, descriptions, publishers, licences, formats, update dates, themes and available usage signals. The scoring follows written rules rather than asking a language model to invent a rating.

This is an assessment of catalogue metadata. It does not inspect every underlying file or establish the full real-world impact of each dataset. That distinction is essential: a weak description can mean weak evidence, rather than weak value.

One example: Ireland’s company records

The prototype scored the Companies Registration Office’s Company Records as Ireland’s dataset.

The CRO catalogue record describes several useful characteristics: a CSV format, daily updates, a CC BY 4.0 licence and an EU High-Value Dataset designation.

The prototype gave it an overall score of 3.192 out of five.

The interesting part was the explanation behind that number. The dataset scored strongly on quality and strategic alignment, while its public-value score was constrained by limited information about who benefits and how it is used.

That gives the publisher an actionable next step: document reuse and outcomes alongside the technical information.

What the prototype revealed

The catalogue run identified 175 potential reuse candidates for further validation. These are leads to investigate, rather than confirmed examples of impact.

It also exposed a measurement issue: the CKAN API returned zero views for every dataset, while dataset pages displayed nonzero counts. Those sources need to be reconciled before views can reliably inform scoring or reporting.

Both findings reinforce the same point. Better assessment depends on better evidence, including evidence of what happens after data is published.

What comes next

My next step is a human review of a 100-record sample to calibrate the framework and check whether its scores reflect informed judgement.

I also want to develop a standard way to record reuse: the applications, services, research and public benefits supported by a dataset. A further extension would assess economic and monetisation potential, helping identify opportunities for startups, AI applications and better government services.

The longer-term ambition is a European open data intelligence layer, with Ireland as the testbed.

The hackathon prototype is a first step towards that ambition. It offers a repeatable way to turn a large catalogue into questions, priorities and opportunities that people can act on.

We have invested in making public data available. Now we need better ways to understand what it enables—and where its next opportunity lies.

Spread the love

Leave a Reply

Your email address will not be published. Required fields are marked *