The quality of your research data depends on good organisation
A crucial part of making data user-friendly, shareable and with long-lasting usability, is to ensure they can be understood and correctly used by other researchers and by your future self.
Efficiently organising your research data also saves times, prevents data loss and makes collaboration easier.
Naming, versioning and formatting your data files
Data files should be well named and version-controlled throughout the research lifecycle. Using standard and interchangeable or open lossless data formats ensures longer-term usability of your data.
Good file names provide useful cues to the content, status and version of a file. They should uniquely identify a file and help in classifying and sorting files.
File names that reflect the file content also facilitate searching and discovering files. In collaborative research, it is essential to keep track of changes and edits to files via the file name. File names should be independent of the location of the file on a computer.
Best practice is to:
- Create meaningful but brief names.
- Order the elements in a file name in the most appropriate way to retrieve the record.
- Use dates in the format YYYY-MM-DD.
- Avoid using spaces, dots and special characters (& or ? or !).
- Use capital letters to delimit words, not spaces or underscores.
- Avoid very long file names.
- When using a personal name in a file give the family name first followed by the initials.
- Avoid using words such as ‘draft’ a the start of the file name, unless doing so will make it easier to retrieve the record.
- Include versioning within file names where appropriate.
There are important things to consider when choosing a file format for digital data, and the choice should be planned early in the research cycle to ensure that the format suits all purposes that might be necessary.
Formats for long-term accessibility
When thinking about long-term accessibility and usability of research data, sustainable digital file formats and software are needed. For many formats, there is a danger that they will become obsolete in the future, which would make the data impossible to read and interpret.
Despite the backward compatibility of many software packages to import data created in previous software versions and the interoperability between competing popular software programmes, the safest option to guarantee long-term data access is to convert data to standard or open formats.
Not only can most software packages interpret these, but they are also suitable for data interchange and transformation, and are likely to stand a better chance of being reused well into the future.
Information on the file formats recommended by the UK Data Archive for long-term preservation
File formats can be proprietary or open
- Proprietary formats are owned by a company that claims intellectual property rights for the use of the software by granting licenses. Standard formats include the widely used proprietary Microsoft Office software products, (MS Word, Rich Text Format and MS Excel), or the popular SPSS format. These are likely to have long-term sustainability as they are so widely used.
- Examples of open file formats are PDF/A, CSV, TIFF, OpenDocument Format (ODF), ASCII, tab-delimited format, comma-separated values and XML.
- File formats can also be lossy or lossless. Lossy formats save space by removing detailed information that is assumed to be unimportant. For example, the lossy format JPEG removes fine detail in images, whilst the lossless format TIFF keeps all the detail. Also, repeatedly editing and saving files in lossy format results in a greater loss of information.
Is the format suitable for conversion?
While researchers will use the most suitable data formats and software according to planned analyses during their research, once data analysis is completed and data are to be prepared for long-term storing, data conversion must be considered. Using open, standard, interchangeable and longer-lasting formats, avoids being unable to use the data in the future. This is also recommended for any backups. For long-term digital preservation, data centres and archives hold data in open and standard formats.
It can be difficult to locate a correct version or to know how versions differ after some time has elapsed. A suitable version control strategy depends on whether files are used by single or multiple users, in one or multiple locations, and whether versions across users or locations need to be synchronised or not, so that if information in one location is altered, the related information in other locations is also updated.
Version control can be done through the following:
- The date recorded in the file name or within the file, for example, HealthTest-2008-04-06.
- Version numbering in the file name, for example, HealthTest-00-02 or HealthTest_v2.
- A file history, version control table or notes included within a file, where versions, dates, authors and details of changes to the file are recorded.
- Version control facilities within the software used.
- Using versioning software, e.g. Subversion.
- Using file-sharing services, such as Dropbox or Google Docs.
- Controlling rights to file-editing.
- Manual merging of entries or edits by multiple users.
Metadata plays the key role in providing the context of how your data was created, analysed and stored.
Metadata is descriptive information about your data. Good metadata is not a side-issue, it is fundamental to making your research data accessible, understandable and usable.
Metadata are valuable in and of themselves, when planning research, especially replication studies. Even if the original data are missing, tracking down people, institutions or publications associated with the original research can be extremely useful.
You should produce metadata as your project moves forward as creating high quality documentation and metadata at the end of the project is very difficult and time consuming.
Read more about metadata
Read more about the types of metadata you should produce and how to record the specified metadata.
Preparing for storing active data during your research
What is Active Data?
Active research data is data that is currently being used, or is planned to be used in the near future.
Correctly storing your active research data is essential to maintaining the authenticity, reliability and integrity of data and it will support data accessibility long after publication
Following good practice in the planning stages of your research, you will have asked yourself the following questions:
- Will I have enough university networked storage space available for all of the data will be created or generated?
- If not, how will I ensure that data is regularly synchronised between devices and backed-up according to university requirements?
- If I need to collaborate with non-Ulster University partners, how will this be done in a secure and organised manner?
- Are any my data sensitive? If so, what appropriate security measures should be in place?
Thinking through these questions in the planning stages of your research gives you a roadmap for storing your research data.
The university requires all researchers to ensure that all active research data in digital and computer-readable form is stored securely in a durable format appropriate for the type of research data in question and is backed up regularly in accordance with best practice in the relevant field of research.
Critical parameters to adhere to when backing up your active research data.
- All original, irreplaceable electronic project data and electronic data from which individuals might be identified must be stored on University-supported media, that are managed, secured, encrypted, supported, operationally controlled, and securely maintained, where access is managed through strong and appropriate levels of secure authentication, good cyber hygiene and identity access management. Such data must never be stored on portable devices or temporary storage media.
- Personal Data must not be kept in an identifiable form for longer than is necessary for the purposes for which the Personal Data was processed or as may be required by law. Further support can be found at the Data Protection at Ulster University webpage.
- All other electronic project data must be held on appropriate centrally-allocated secure server space which is accessible to members of the project team; such data must not be held on personal or portable devices unless these are encrypted in line with University requirements and except when this is necessary for the purposes of working off-site; amended documents must be returned to the appropriate University-maintained shared space when the work has been completed.
- Under no circumstances should original, irreplaceable data or sensitive personal data be stored using cloud storage services as this can place data outside UK and EU legal control.
| Suitable for storing active research data | Not suitable for storing active research data |
|---|---|
| Faculty/School storage systems | External hard drives and USB sticks |
| Network drives | Local storage that is not backed-up |
| SharePoint | Third party cloud storage |
Data Security
Ensuring the security of data requires paying attention to physical security, network security, plus the security of computer systems and files to prevent unauthorised access or unwanted changes to data, disclosure or destruction of data. Researchers should think carefully about their research processes to ensure data is stored and transferred securely throughout the research project.
This might involve:
- Ensuring physical data is stored in secure locations and digitised where appropriate.
- Logging the removal of, and access to, media or hardcopy material in storerooms.
- Not storing confidential data, such as those containing personal information on servers or computers connected to an external network, particularly servers that host internet services.
- Firewall protection, security-related upgrades and patches to operating systems to avoid viruses, trojans and malicious codes.
- Anonymising data, or pseudo-anonymising data and storing the key code in a separate location to the data.
- Implementing password protection of, and controlled access to, individual data files, for example, allocating ‘no access’, ‘read only’, ‘read and write’ or ‘administrator only permissions.
- Ensure that you lock your computer if you leave it temporarily unattended by pressing Ctrl-Alt-Del and clicking on 'Lock Computer'.
- Not sending personal or confidential data via email. This should be encrypted and sent via a secure means, not email.
- Destroying data in a consistent and robust manner when needed.
You should think about data curation and storage from a FAIR perspective
Some thoughts to start....
Making the data Findable and Accessible:
- Folders relating to each work package provide a secure shared workspace (e.g. University OneDrive or SharePoint folder) where the project team can contribute to the data collected; team members can be given editing or full admin rights as appropriate.
- University OneDrive and SharePoint systems have appropriate infrastructure and support in place to ensure data is protected and preserved.
- A data index maintained within the site ensures data consistency/quality control and enable specific data to be easily located.
- Well thought out access procedures to the data including authentication and authorisation steps provide safeguarded access to sensitive data.
- Handwritten notes transcribed into Word files on the day the notes are taken and saved on the SharePoint password-protected server ensures integrity and accessiblility.
Making the data Interoperable:
- Naming conventions and formatting templates standardise the information acquired.
- Metadata follows standards relevant to the discipline.
- Open and standard file formats facilitate interoperability.
Making the data Reusable:
- Clear and detailed provenance for the data: how, why, and by whom is the data created and/or processed.
- Comprehensive metadata allows the research to be replicated.
Further ulster policies, standards and guidelines information
Ulster University Information Services Directorate maintains a range of policies, standards and guidelines for the secure and reliable delivery of services across the University.
A simplified and practical overview for staff on how these policies can help to protect University information and handling electronic data is available in a staff information handbook.
Read more about Research Data Management



