This text document contains a description of the files uploaded to SND. The data files contain the raw files and scripts for processing as used in the manuscript "Coverage of Web Accessibility Guidelines Provided by Automated Checking Tools" authored by Thomas Fischer, Björn Lundell, and Jonas Gamalielsson in 2024.

Primary author of data files and scripts is Thomas Fischer <thomas.fischer@his.se>. All data unless noted otherwise is released under the Creative Commons Attribution-Share Alike (BY-SA) 4.0 license.

Personal information in this data collection includes among others:
- Names and contact details of contributing authors.
- In one archive file (details below): Textual data collected from publicly available web pages of Swedish public sector organizations (PSOs), which may include names, contact details, or other personal or biographical information. Due to the directory structure, for every file the origin of the data is determined, so any further questions about the handling of personal data shall be directed to the respective PSO.


20240226-scan-webpages-run-experiment-myndigheter-dom-files.tar.xz
20240226-scan-webpages-run-experiment-myndigheter-without-dom-files.tar.xz

Both tar archives contains a single subdirectory (scan-webpages-run-experiment-myndigheter) which in its turn contains about 550 subdirectories and three text files.

The three text files (installed-packages*txt) contain information which npm packages were installed in the containers that ran the three automated accessibility checkers during the experiments.

There is one subdirectory for each webpage visited for the automated accessibility analysis. In most cases, the subdirectories' names match the organizations' domain names (for example, "rymdstyrelsen.se") as only landing pages were visited. In few cases, an organization did not have its own top-level domain, but was hosted under another organization's domain, which is then reflected in the subdirectory's name. One example is "regeringen.se-myndigheter-med-flera-oljekrisnamnden".

The two tar archives differ in which files are contained in each subdirectory. Within each tar archive, the files per subdirectory have the same names. The "without-dom-files" tar archive contains only the results of the six engines' analysis in JSON format, but no potentially personal data retrieved from webpages, this this tar archive can be freely shared. The other tar archive (without "without" in its filename) contains the content of the analyzed webpage, potentially including personal data which may not be freely shared.

Per-domain subdirectories in 20240226-scan-webpages-run-experiment-myndigheter-without-dom-files.tar.xz contain the following files: "a11ywatch-axe.json", "a11ywatch-htmlcs.json", "pa11y-axe.json", "pa11y-htmlcs.json", "qualweb-act.json", "qualweb-wcag.json" are JSON log files as generated by the tools and engines as described in the manuscript. The JSON files were modified to remove longer fragments of HTML source code recorded by the analysis tools.

Per-domain subdirectories in 20240226-scan-webpages-run-experiment-myndigheter-dom-files.tar.xz contain the following files:
- dom.html: Mimicking what the analysis tools would do, a headless Chromium browser was used to dump the DOM (document object tree) of the visited webpages. Notably, this document does not necessarily match the HTML code that can be downloaded from the web servers, but matches the internal representation after the execution of the webpages' initial JavaScript and similar page-modifying activities. The accessibility checking tools did not use this file, but directly retrieved their version of the webpage from the respective web servers.
- dom.txt: For reference only, a text dump (without HTML) of dom.html.


installed-packages.zip

This zip file contains only the three text files (installed-packages*txt) included in 20240226-scan-webpages-run-experiment-myndigheter*.tar.xz which document npm packages were installed.
This allows to easily see the used software packages without handing an archive of 18MB.


scan-webpages-run-experiment-myndigheter---file-list-md5sums.txt.xz

This file is a compressed text file, which contains a combined list of files included 20240226-scan-webpages-run-experiment-myndigheter*.tar.xz as well as each files MD5 checksum.


coverage.json.xz

This JSON file aggregates information about WCAG 2.x success criteria and various checkers' related information. It can be updated with the help of refresh-coverage-json.py (see below) and other helper scripts.
The JSON file contains, among others, the following information:
- Links to W3C's webpage for WCAG principles, guidelines, and success criteria for each WCAG 2.0, 2.1, and 2.2
- For each of the used automated checkers links to webpages and source code files that are in relation to the given success criterion.
This JSON file can be queried to check, for example, which Axe-core checks cover success criterion 1.3.1 and where the corresponding source code files for Axe-core 4.7 can be found.

This file is XZ-compressed only for archival purposed. For practical use, such as for the Python scripts below, expand this file manually first, for example by using the tool "unxz".


snd-python.zip

This archive contains a collection of Python script files that are useful to process the research data from JSON documents as describe above into LaTeX code (tables, TikZ diagrams, link rewriting, ...).
The scripts were run with Python 3 interpreters and modules supplied with Fedora Linux versions current during the time of the study, e.g. Fedora Linux 40.
Thus, Python version 3.12 as well as slightly older or newer Python version should work be able to run the scripts.
Only Python packages shipped with Fedora Linux were used; no packages got installed in virtual environments or via pip, for example. Installed and used Python packages include: python3-binaryornot, python3-beautifulsoup4

Some of those scripts may have hard-coded paths:
- It may be expected that archive "20240226-scan-webpages-run-experiment-myndigheter-without-dom-files.tar.xz" got extracted in /tmp/, i.e. it is expected that there is a file "/tmp/scan-webpages-run-experiment-myndigheter/rymdstyrelsen.se/a11ywatch-axe.json".
- Python file "disagreement-between-checkers.py" expects that archive "20240226-scan-webpages-run-experiment-myndigheter-dom-files.tar.xz" got extracted in /tmp/, i.e. it is expected that there is a file "/tmp/scan-webpages-run-experiment-myndigheter/rymdstyrelsen.se/dom.html".
- Some Python scripts expect that file "coverage.json" is located in directory "../WebAccessibility/shared/coverage.json" or "../shared/coverage.json", respectively.

The paths are as they were on the author's computer, and for use on other systems, the paths (typically specified at the beginning of a Python file) may need to be adjusted.

Selected Python scripts will be described in the remainder of this text document.


coverage-map.py

This file generates TikZ code for embedding in LaTeX documents that became Figure 3 ("Coverage of WCAG success criteria by accessibility checkers.") in the manuscript. Data for this diagram is loaded from "coverage.json", but also from file "disagreement-between-checkers.txt" which has to be pre-generated using "disagreement-between-checkers.py" (which in its turn makes use of "coverage.json" and files from "20240226-scan-webpages-run-experiment-myndigheter-without-dom-files.tar.xz" and "20240226-scan-webpages-run-experiment-myndigheter-dom-files.tar.xz").


refresh-coverage-json.py

Loads, updates, and writes back the JSON data of file coverage.json. Updated information is retrieved from a number of other files (see script's source code for details), such as "qualweb-wcag-techniques.json", "qualweb-act-rules.json", various JSON files on Axe-core code, by parsing webpages of W3C on WCAG (may break if they ever redesign their webpages), and by fetching tools' source code from GitHub (for example, to check which source code files can raise errors).

The updated JSON code is printed to stdout, but on a Unix/Linux shell can be redirected into a JSON file. It is recommended to carefully compare the old and the new coverage.json before overwriting the old file, in case "refresh-coverage-json.py" failed and erased data for any reason.


link-to-external.py

This is a Python script that can rewrite LaTeX code to insert URLS (\href{..}{..}) to external sources for text strings that match a number of pre-defined patterns. For example, if a LaTeX file contains the string "1.3.1", it is assumed that this is a success criterion identifier and the text will be replaced by the LaTeX code
\href{https://www.w3.org/WAI/WCAG22/Understanding/info-and-relationships.html}{1.3.1}
