Module 1: Text preprocessing and linguistic signals#
Theme#
Text preprocessing and linguistic signals
Essential Question#
What is lost and gained when language becomes data?
Module Components#
Book prose: conceptual framing, domain scenario, methods, and failure modesAssignment: evidence-backed production of a specific artifactSlides: presentation sequence for seminar or lecture deliveryNarration: spoken version of the slide flowRubric: criteria for evaluating the module artifactNotebook: executable lab aligned with the module theme using synthetic support messages, retrieval snippets, intent labels, and factuality checks
Module Artifact#
NLP evaluation packet with task framing, retrieval/evaluation design, and deployment guardrails focused on text preprocessing and linguistic signals: Compare tokenization choices on a small corpus.
Professional Setting#
Students work as if advising a product team evaluating an NLP workflow before using it in customer-facing communication. Their work must be intelligible to product manager, support lead, privacy reviewer, and model evaluator.
Use This Module in Order#
Review the slide deck with the matching narration.
In Populi, open the private student-repository link for this course and enter
modules/module-1.Clone the repository once or open its Codespace/Colab copy; run
lab.ipynband completeexercise.ipynbthere.Self-check with the rubric, commit and push the work, then submit exactly what Populi requests.