Guidelines for data preparation
Appropriate use of data is only possible if methodology and assumptions behind original datasets are known, as well as weaknesses and limitations of the data. Thus the data archive collects not just row data, but also all documentation relevant to data use including instruments used for data collection.
Data files must be clean and clear, i.e., interpretable and searchable by ways of a codebook by other researchers. Supplementary documentation should give as much material as possible in order to make the study as suitable as possible for secondary use. This kind of accompanying material could include methodological reports, codebooks, questionnaires, coding instructions, interviewer guides, database dictionaries, bibliographies of publications related to the data, and links to online tools.
Data and documentation preparation checklist
- Remove all direct identifiers from your dataset (names, addresses, telephone number and any other variables that would allow easy identification of individuals).
- Name your variables clearly and logically, explain shortly their meaning and link them with corresponding questions in the questionnaire. Give short description also to response options.
- Replace missing values with explicit code (e.g. 88 for “don’t know”). Missing values should not appear as an empty case or as a missing value attributed by default by the statistical program you are using.
- Check frequencies and detect and repair or delete any inconsistencies or abnormalities in your data.
- Prepare instruments (questionnaires, interview guides) used for the collection of data in separate file, as well as
- any materials sent in advance to respondents (e.g., advance letters or postcards),
- any materials presented to respondents during the interviews (e.g., showcards), and
- any instructions or materials for use by interviewers (e.g., question by question explanations, frequently asked questions).
- Prepare relevant documentation which should include the following type of information:
- the context of the data: project history, objectives, research design, and hypotheses
- population, sampling design, and sample size
- unit of analysis
- data collection mode (CATI, CAPI, mail, web, etc.)
- response rate
- temporal and geographic coverage
- structure of data files, cases, relationships between files (if applicable)
- data validation, checking, proofing, cleaning and other quality assurance procedures carried out
- information on data confidentiality, access and use conditions
- weighting
- recoded and derived variables created after collection, with code, algorithm or command file used to create them
Data formats
Regarding the data and documentation file formats that are accepted by the data archive, we prefer formats that are most likely to be accessible in the future. In other words, formats that are non-proprietary, openly documented, unencrypted, and uncompressed that are commonly used by the research community.
We therefore consider the following formats as appropriate:
- Tabular data: SPSS portable format (.por), SPSS (.sav), Stata (.dta), Excel or other spreadsheet format files, which can be converted to tab- or comma-delimited text, ), R (.txt);
- Text: Adobe Portable Document Format (PDF/A, PDF) (.pdf), plain text data, ASCII (.txt);
- Audio: Waveform Audio Format (WAV) (.wav) from Microsoft, Audio Interchange File Format (AIFF) (.aif) from Apple, FLAC (.flac);
- Images: TIFF (.tif) ideally version 6 uncompressed, JPEG (.jpeg, .jpg) only when created in this format, Adobe Portable Document Format (PDF/A, PDF) (.pdf), RAW image format (.raw), Photoshop files (.psd);
- Video: MPEG-4 (.mpg4), motion JPEG 2000 (.mj2);
Compressed files: are accepted as long as they can be uncompressed by using open and freely available software, such as 7-Zip or Winzip.
Data deposit agreement
The data deposit agreement describes the legal framework under which material is deposited and states the rights and responsibilities of both parties – the depositor and the archive – regarding copyright and ownership of the data and the access conditions. It is a legal agreement between the depositor and the archive that covers arrangements regarding usage rights, authenticity, data protection responsibilities, and disposal.
The data producer can during the negotiation phase place the data under an embargo or special conditions, which means that data are not accessible to data users for a pre-defined period or only with certain restrictions. However, in the spirit of Open Access we highly recommended that data are made available to users as soon as possible after a research project has come to an end. The data depositor can also impose certain conditions on the use of the data (e.g., to be informed beforehand about who is going to use the data and for what purposes).
Data management planning
The quality of documentation can be significantly improved if its creation and collation is planned at the beginning of the data life cycle, during the project conception phase. If needed, archive staff will offer assistance in this respect.