Bridging the Gap: How Practicing Data Scientists Use LLMs in the Wild and What It Means for Data Science Education
Since the widespread availability of generative artificial intelligence (GenAI), particularly Large Language Models (LLMs), fundamental questions have emerged about the future of coding in data science. Some predict that data scientists will no longer need traditional coding skills, while others question whether LLMs might replace data scientists entirely. However, these discussions have largely proceeded without empirical evidence of how practicing data scientists actually use these tools.
This study addresses this gap by surveying trained, practicing data scientists to understand if and how they integrate LLMs into their workflows, particularly for writing and editing code and performing other data science tasks. Building on our recent investigation of data science educators' perspectives on LLMs, this research examines real-world usage patterns among practitioners to bridge the gap between current practice and educational preparation.
Our findings will contribute to the data science community in two critical ways. First, by documenting how data scientists are actually working with LLMs four years after their initial release, we provide actionable insights that allow practitioners to learn and adopt effective strategies for integrating these tools into their work. Second, we inform data science education by evaluating whether current pedagogies adequately prepare students for this evolving landscape.
This research will help answer key pedagogical questions: Should coding education emphasize code reading, tracing, and editing over writing large amounts of de novo code? Should greater focus be placed on writing high-quality documentation and specifications, given their value as prompt context for LLMs? Should testing receive increased emphasis to enable verification of LLM-generated code? By grounding these questions in empirical evidence of practitioner behaviour, we aim to provide data-driven guidance for evolving data science curricula.
To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.
This talk is one of the Teaching and Learning in Statistics and Data Science Seminar Series.
