Andrés Felipe Puerta Velez

Andrés Felipe Puerta Velez

Asistente de investigación y estudiante de maestría en matemáticas aplicadas. @ Universidad EAFIT

About

He is a master's student in applied mathematics and a research assistant, with experience in natural language processing (NLP), integration of heterogeneous databases, and data analysis. He currently works on using the Bayesian g-formula as a solution to the multiplicity of causal effect estimators and on using machine learning models to estimate pollutant gas emissions and fuel consumption for low-displacement motorcycles in the Colombian context. As a research assistant and scholarship holder on the Comparative Analysis of Perceptions on Protective Behaviors against COVID-19 in Colombia project, he has worked on teams made up of students and researchers from diverse disciplines to study how public discourse circulates on social networks and how it behaved regarding care habits during the pandemic. At PyCon Colombia 2026 he will participate as moderator of a workshop on natural language processing for corpus analysis in Digital Humanities, sharing experiences on how Python can become a meeting tool between technology and the humanities.

Workshop

Artificial IntelligenceData Science

NLP in Practice: From Corpus Linguistics to RAG with Python

FORMAT: WorkshopLEVEL: IntermediateLANGUAGE: Spanish

Natural language processing today offers a mature set of tools for analyzing textual corpora systematically and reproducibly, but the path between having the documents and obtaining results is not always clear. This workshop covers that path from start to finish. In two hours, participants will build an understanding of the NLP ecosystem: its history, logic, and methods. The session opens with a timeline from the earliest rule-based models to transformers, followed by a map of techniques organized by problem type (classification, entity extraction, semantic search, generation) so each participant can identify which method they need for a specific textual problem. The second part covers two implementations with Python. First, topic modeling with BERTopic, reviewing the internal pipeline of embeddings, UMAP, and HDBSCAN. Second, a conversational assistant with RAG: corpus indexing, semantic retrieval, and connection with a language model to answer queries about the documents. Upon completion, each participant will have a functional notebook with both pipelines and a clear map of the ecosystem to guide their own textual analysis projects.

Andrés Felipe Puerta Velez

Andrés Felipe Puerta Velez

Asistente de investigación y estudiante de maestría en matemáticas aplicadas. @ Universidad EAFIT

Biviana Marcela Suárez Sierra

Biviana Marcela Suárez Sierra

Profesora vinculada al área de Computación y analística @ Universidad EAFIT

Dora Cecilia Alzate Gallo

Dora Cecilia Alzate Gallo

Estudiante de la Maestría en Estudios Humanísticos @ EAFIT

Karen Melissa Gomez Montoya

Karen Melissa Gomez Montoya

Ingeniera matemática - Asistente en investigación @ Universidad EAFIT

View talk

Want to know more?

Join PyCon Colombia newsletter and get a complete overview of our events, speakers and community participation.