Introduction to Data Science and Artificial Intelligence with Python
Course Description
Characteristics
A four-day workshop course that guides participants from data analysis and preparation in Python (Pandas), through building and evaluating machine learning models (regression and classification), to a practical introduction to large language models (LLM) and their applications in analytical and product work. The sessions combine short theoretical blocks with intensive hands-on exercises with data, so that participants complete the training with a set of ready-to-use techniques and best practices for implementation in daily tasks.
Learning Objectives
After the course, participants will be able to:
- prepare data for analysis (cleaning, filtering, aggregations, feature engineering) using Pandas
- build, evaluate, and compare regression and classification models in scikit-learn
- apply techniques to improve model quality: normalization, encoding categorical variables, cross-validation, and hyperparameter tuning
- recognize common risks (e.g., overfitting) and select appropriate evaluation metrics
- understand the basics of working with LLMs (API, embeddings, multimodal models) and safely prototype simple solutions
Book the Training
- Format: Remote
- Language: English
- Type: Open Enrollment
- Date: 24-27 Feb 2026
- Duration: 4 days (7h/day)
- Trainer: Patryk Palej
- Validator: Bartosz Wójcik
Net price per participant.
Target Audience
The course is designed for individuals who want to develop data/AI skills in practice, particularly:
- data analysts, business analysts, and reporting specialists
- professionals from finance, sales, operations, and marketing departments working with data
- developers and engineers who wish to solidify their foundations in ML/LLM using Python
- individuals preparing for a Data Scientist / ML Engineer role at a junior to regular level
Prerequisites
- basic knowledge of Python (variables, loops, functions, working in a notebook)
- basics of data handling (tables, data types, simple calculations)
- readiness to work on your own computer in a remote environment
Training Methods
Practical workshops dominate (approx. 75% of the time). During the sessions, we use:
- individual and team exercises in notebooks (Python)
- short theoretical introductions preceding labs
- case studies based on data resembling real-world business applications
- mini-projects summarizing each day and work with pipelines
- Q&A sessions and consultations regarding participants' work issues (if data/cases are provided)
Training Materials
Participants receive a complete set of materials used during the training, including:
- Jupyter notebooks and exercise files
- dataset packs for labs
- cheat sheets with key Pandas functions and scikit-learn elements
- links to recommended resources and documentation
Technical Conditions for Remote Training
Participation requires:
- computer with Windows/macOS/Linux OS (min. 8 GB RAM, 16 GB recommended)
- stable internet connection (min. 10 Mb/s)
- headset with microphone and camera (recommended for workshop work)
- access to MS Teams or Zoom communicator
- optional: access to the DaDesktop environment provided by NobleProg (if launched for the group)
Validation and Certificates
Validation of learning outcomes is based on practical tasks (notebooks), mini-projects, and quality checklists. Upon completion of the training, participants receive a NobleProg certificate in electronic format.
Information Document
The program can be tailored to the group's needs after an requirements analysis (PCQ).
Learning Outcomes and Verification Criteria
Data Analysis with Pandas
- Creates and modifies Series/DataFrame objects, performs operations on tables and columns
- Applies filtering, sorting, grouping, and aggregations to solve analytical problems
Verification criteria: Verification: practical tasks in a notebook (calculation results and correctness of data transformations)
Regression Models
- Selects and trains a regression model for a given problem
- Evaluates the model using appropriate metrics and interprets results
- Applies normalization and data preparation, as well as prevents overfitting
Verification criteria: Verification: regression mini-project + discussion of metrics and insights.
Classification Models
- Builds and compares classification models, selecting metrics (e.g., accuracy, precision/recall, F1, ROC-AUC)
- Applies ensemble techniques to improve prediction quality
Verification criteria: Verification: classification task + comparison of at least 2 models with justification for the choice.
Working with LLMs
- Understands key concepts: prompts, context, embeddings, tokenization, model limitations
- Uses LLM APIs for text generation and processing and creates embeddings for semantic search
- Knows basic security and quality principles (data protection, testing, limitations)
Verification criteria: Verification: integration exercise (simple prototype) + data security checklist.
Training Program
Day 1 - Data Analysis with Pandas
- Basic data types: Series and DataFrame (creation, indexing, data types).
- Table operations: loading data, merging/joining, concatenation.
- Filtering, sorting, grouping: groupby, aggregations, pivot tables.
- Value modification: mapping, replace, missing data (NaN) and imputation strategies.
- Column operations: feature creation and transformation, apply/assign functions.
Day 2 - Machine Learning: Regression Algorithms
- Introduction to regression: problem, data, ML pipeline.
- Model evaluation: metrics (MAE, MSE/RMSE, R2), train/test split.
- Data normalization and standardization: when and why.
- Handling categorical variables: One-Hot Encoding and Label Encoding.
- Overfitting: diagnosis and reduction methods.
- Cross-validation: selection of validation strategy.
- Grid Search and hyperparameter optimization (including pipeline).
Day 3 - Machine Learning: Classification Algorithms
- Introduction to classification: binary and multi-class.
- Classification evaluation: confusion matrix, accuracy, precision/recall, F1, ROC-AUC.
- Overview of classification algorithms (e.g., logistic regression, trees, SVM, kNN).
- Ensemble: combining classifiers (bagging, boosting, stacking - overview and practice).
Day 4 - Large Language Models (LLM)
- Introduction to LLMs: capabilities, limitations, typical business use cases.
- OpenAI API and other models: integration basics, costs, limits, best practices.
- Multimodal models: working with text and images - scenarios and limitations.
- Embeddings: semantic search, clustering, simple recommendations, RAG basics.
Lack of funds for training? Get funding!
A program that allows you to easily and quickly obtain funding for training for individual individuals.
Why guaranteed training?
- Guarantee of delivery. The training will take place regardless of the number of participants.
- Knowledge and experience exchange with specialists from other industries.
- Interactive, live-led sessions. Not just theory, but also practical exercises and discussions.
- Flexible remote format. Join from anywhere.
Need Help?
Reach out to learn more about our team and the kinds of tailored solutions we can offer your organization.
Get in Touchwroclaw@nobleprog.pl or +48 (22) 103 3718