Get in Touch

Introduction to Data Science and Artificial Intelligence with Python

 

Course Description

Characteristics

A four-day workshop course that guides participants from data analysis and preparation in Python (Pandas), through building and evaluating machine learning models (regression and classification), to a practical introduction to large language models (LLM) and their applications in analytical and product work. The sessions combine short theoretical blocks with intensive hands-on exercises with data, so that participants complete the training with a set of ready-to-use techniques and best practices for implementation in daily tasks.

Learning Objectives

After the course, participants will be able to:

  • prepare data for analysis (cleaning, filtering, aggregations, feature engineering) using Pandas
  • build, evaluate, and compare regression and classification models in scikit-learn
  • apply techniques to improve model quality: normalization, encoding categorical variables, cross-validation, and hyperparameter tuning
  • recognize common risks (e.g., overfitting) and select appropriate evaluation metrics
  • understand the basics of working with LLMs (API, embeddings, multimodal models) and safely prototype simple solutions

Book the Training

  • Format: Remote
  • Language: English
  • Type: Open Enrollment
  • Date: 24-27 Feb 2026
  • Duration: 4 days (7h/day)
  • Trainer: Patryk Palej
  • Validator: Bartosz Wójcik

BOOK NOW - 5160 PLN 

Net price per participant.

Target Audience

The course is designed for individuals who want to develop data/AI skills in practice, particularly:

  • data analysts, business analysts, and reporting specialists
  • professionals from finance, sales, operations, and marketing departments working with data
  • developers and engineers who wish to solidify their foundations in ML/LLM using Python
  • individuals preparing for a Data Scientist / ML Engineer role at a junior to regular level

Prerequisites

  • basic knowledge of Python (variables, loops, functions, working in a notebook)
  • basics of data handling (tables, data types, simple calculations)
  • readiness to work on your own computer in a remote environment

Training Methods

Practical workshops dominate (approx. 75% of the time). During the sessions, we use:

  • individual and team exercises in notebooks (Python)
  • short theoretical introductions preceding labs
  • case studies based on data resembling real-world business applications
  • mini-projects summarizing each day and work with pipelines
  • Q&A sessions and consultations regarding participants' work issues (if data/cases are provided)

Training Materials

Participants receive a complete set of materials used during the training, including:

  • Jupyter notebooks and exercise files
  • dataset packs for labs
  • cheat sheets with key Pandas functions and scikit-learn elements
  • links to recommended resources and documentation

Technical Conditions for Remote Training

Participation requires:

  • computer with Windows/macOS/Linux OS (min. 8 GB RAM, 16 GB recommended)
  • stable internet connection (min. 10 Mb/s)
  • headset with microphone and camera (recommended for workshop work)
  • access to MS Teams or Zoom communicator
  • optional: access to the DaDesktop environment provided by NobleProg (if launched for the group)

Validation and Certificates

Validation of learning outcomes is based on practical tasks (notebooks), mini-projects, and quality checklists. Upon completion of the training, participants receive a NobleProg certificate in electronic format.

Information Document

The program can be tailored to the group's needs after an requirements analysis (PCQ).

Learning Outcomes and Verification Criteria

Data Analysis with Pandas

  • Creates and modifies Series/DataFrame objects, performs operations on tables and columns
  • Applies filtering, sorting, grouping, and aggregations to solve analytical problems

Verification criteria: Verification: practical tasks in a notebook (calculation results and correctness of data transformations)

Regression Models

  • Selects and trains a regression model for a given problem
  • Evaluates the model using appropriate metrics and interprets results
  • Applies normalization and data preparation, as well as prevents overfitting

Verification criteria: Verification: regression mini-project + discussion of metrics and insights.

Classification Models

  • Builds and compares classification models, selecting metrics (e.g., accuracy, precision/recall, F1, ROC-AUC)
  • Applies ensemble techniques to improve prediction quality

Verification criteria: Verification: classification task + comparison of at least 2 models with justification for the choice.

Working with LLMs

  • Understands key concepts: prompts, context, embeddings, tokenization, model limitations
  • Uses LLM APIs for text generation and processing and creates embeddings for semantic search
  • Knows basic security and quality principles (data protection, testing, limitations)

Verification criteria: Verification: integration exercise (simple prototype) + data security checklist.

Training Program

Day 1 - Data Analysis with Pandas

  • Basic data types: Series and DataFrame (creation, indexing, data types).
  • Table operations: loading data, merging/joining, concatenation.
  • Filtering, sorting, grouping: groupby, aggregations, pivot tables.
  • Value modification: mapping, replace, missing data (NaN) and imputation strategies.
  • Column operations: feature creation and transformation, apply/assign functions.

Day 2 - Machine Learning: Regression Algorithms

  • Introduction to regression: problem, data, ML pipeline.
  • Model evaluation: metrics (MAE, MSE/RMSE, R2), train/test split.
  • Data normalization and standardization: when and why.
  • Handling categorical variables: One-Hot Encoding and Label Encoding.
  • Overfitting: diagnosis and reduction methods.
  • Cross-validation: selection of validation strategy.
  • Grid Search and hyperparameter optimization (including pipeline).

Day 3 - Machine Learning: Classification Algorithms

  • Introduction to classification: binary and multi-class.
  • Classification evaluation: confusion matrix, accuracy, precision/recall, F1, ROC-AUC.
  • Overview of classification algorithms (e.g., logistic regression, trees, SVM, kNN).
  • Ensemble: combining classifiers (bagging, boosting, stacking - overview and practice).

Day 4 - Large Language Models (LLM)

  • Introduction to LLMs: capabilities, limitations, typical business use cases.
  • OpenAI API and other models: integration basics, costs, limits, best practices.
  • Multimodal models: working with text and images - scenarios and limitations.
  • Embeddings: semantic search, clustering, simple recommendations, RAG basics.
People flying on paper airplanes at training.

Lack of funds for training? Get funding!

Entity Financing System for Adults logo

A program that allows you to easily and quickly obtain funding for training for individual individuals.

View More

Open training offer with guaranteed dates shown in the form of pictograms in screws.

Why guaranteed training?

  • Guarantee of delivery. The training will take place regardless of the number of participants.
  • Knowledge and experience exchange with specialists from other industries.
  • Interactive, live-led sessions. Not just theory, but also practical exercises and discussions.
  • Flexible remote format. Join from anywhere.

View More

Two persons looking at a tablet

Need Help?

Reach out to learn more about our team and the kinds of tailored solutions we can offer your organization.

Get in Touch

wroclaw@nobleprog.pl or +48 (22) 103 3718