← Research using Delve

arXiv · June 2023

Quality Issues in Machine Learning Software Systems

Pierre-Olivier Côté, Amin Nikanjam, Rached Bouchoucha, Ilan Basta, Mouna Abidi, Foutse Khomh

SWAT Lab, Polytechnique Montréal, Québec

Grounded theory HCI & computer science

How they used Delve

Software engineering researchers at Polytechnique Montréal used Delve to code 37 practitioner interviews about quality issues in machine learning systems, double-coding every transcript with a third researcher moderating disagreements, and publishing every coder's and moderator's coded transcripts in the replication package.

“To help us code the transcripts, we used Delve qualitative analysis tool, because the researchers are familiar with it and it is easy to use. Delve is a computer-assisted qualitative data analysis software (CADQAS) that provides simple interfaces to code and analyze data. In order to ensure the quality of the analysis, each document is coded by two researchers. In case of a disagreement in codes, a third researcher plays the role of moderator and selects the final code for a text segment.”

Field
Empirical software engineering; a practitioner-derived catalog of quality issues in machine learning software systems and the strategies used to mitigate them
Data
37 interviews with machine learning practitioners and experts about the quality issues they hit in real ML software systems, yielding 18 recurring issues and 21 mitigation strategies, subsequently validated by a practitioner survey
Approach
Grounded theory following Strauss and Corbin: open coding of interview transcripts, memos capturing interesting information and preliminary category ideas, memo sorting, and axial coding to group codes into categories so the issues could be reasoned over. Every document was coded by two researchers; where their codes disagreed, a third researcher acted as moderator and chose the final code for the segment. The coded transcripts of both coders and the moderators are published in the replication package. Findings were then validated through a survey of ML practitioners.
Data types
Interviews

Abstract

Context: An increasing demand is observed in various domains to employ Machine Learning (ML) for solving complex problems. ML models are implemented as software components and deployed in Machine Learning Software Systems (MLSSs). Problem: There is a strong need for ensuring the serving quality of MLSSs. False or poor decisions of such systems can lead to malfunction of other systems, significant financial losses, or even threats to human life. The quality assurance of MLSSs is considered a challenging task and currently is a hot research topic. Objective: This paper aims to investigate the characteristics of real quality issues in MLSSs from the viewpoint of practitioners. This empirical study aims to identify a catalog of quality issues in MLSSs. Method: We conduct a set of interviews with practitioners/experts, to gather insights about their experience and practices when dealing with quality issues. We validate the identified quality issues via a survey with ML practitioners. Results: Based on the content of 37 interviews, we identified 18 recurring quality issues and 21 strategies to mitigate them. For each identified issue, we describe the causes and consequences according to the practitioners’ experience. Conclusion: We believe the catalog of issues developed in this study will allow the community to develop efficient quality assurance tools for ML models and MLSSs. A replication package of our study is available on our public GitHub repository

Citation

Pierre-Olivier Côté, Amin Nikanjam, Rached Bouchoucha, Ilan Basta, Mouna Abidi, Foutse Khomh (2023). Quality Issues in Machine Learning Software Systems. arXiv. https://doi.org/10.48550/arxiv.2306.15007

Resources

Start your 14 day free trial of Delve