What is metadata?

Metadata is ‘data about data’ or ‘cataloguing information’ that enables data users to find or use a dataset.  It makes data discoverable, understandable, reusable and reproducible.  Metadata is essential infrastructure for  supporting integrity and rigour in research.  Data that is well described with metadata can be replicated and/or combined in different settings.  You should produce metadata as your project moves forward as creating high quality documentation and metadata at the end of the project is very difficult and time consuming.

Ulster University image

There are three types of metadata you should maintain for your research data

  1. Administrative metadata
  2. Descriptive metadata
  3. Structural metadata
Metadata type explanation
Metadata TypeExplanation
Administrative metadata Data about a project or resource that are relevant for managing it; for example, project/ resource owner, principal investigator, project collaborators, funder, project period. They are usually assigned to the data, before you collect or create them.
Descriptive metadata Data about a dataset or resource that allow people to discover and identify it; for example, authors, title, abstract, keywords, persistent identifier, related publications.
Structural metadata Data about how a dataset or resource came about, but also how it is internally structured.  This metadata focuses on the technical processes used to produce, or required to use, the data. Structural metadata should be gathered by the researchers according to best practice in their research community and be published together with the data.

Examples of structural metadata

  • Fieldwork activities: participant code; date/place of activity (e.g. interview); logged confirmation of consent; filename of research schedule used; filename(s) of interview record(s) obtained (audio record, scanned version of handwritten interview notes); filename of transcribed record; filename of anonymised transcribed record.
  • The location of key data sources: field notes, consent forms, transcripts etc.
  • Experimental conditions, sample preparation and sample processing, on the performed measurements
  • Information about the clinical samples, biological reagents (e.g., cell lines, antibodies, siRNAs), chemical reagents (e.g., drugs), etc. used to generate the data.
  • Coding frames for qualitative interviews and focus groups.
  • Standard Operating Procedures: Clinical researchers often apply institution and domain specific SOPs that outline specific processes for data management.

Once you have established which metadata you will produce, you should consider how you will record it.

Documentation on metadata can be maintained in a variety of forms, including, but not limited to:

  • README: A README File is a text file located in a project-related folder that describes the contents and structure of the folder and/or a dataset so that a researcher can locate the information they need.
  • Data Dictionary: Also known as a codebook, a data dictionary defines and describes the elements of a dataset so that it can be understood and used at a later date.
  • Protocol: A protocol describes the procedure(s) or method(s) used in the implementation of a research project or experiment.
  • Lab Notebook: For research groups that use them, lab notebooks are often the primary record of the research process. They are used to document hypotheses, experiments, analyses, and interpretations of experiments.
  • Textual data file:  Background and contextual information, key biographical characteristics and thematic features of participants. Audiovisual files: creator, date, location, subject, content, copyright, keywords, equipment used.
Ulster University image

Read more about looking after your data during your project

Ulster University image