Posts

Showing posts from September, 2026

Fine-Tuning Mistral 7B using QLoRA with PyTorch pt. 2: K8s & GKE | ML Engineering & MLOps

Image
  Hi All Continuing from Part 1, this post details the Kubernetes and Observability configs of the project. The full source is available here . Let's break the code down shall we.  1. K3's Server Config ( infra/server-config.yaml )  write-kubeconfig-mode: "0644"  *      Sets file permissions for kubeconfig file (readable by all users in the group) *      0644 means owner can read/write, group and others can only read disable: - traefik - servicelb - local-storage - metrics-server *      Disables default k3s components that we'll replace with better alternatives. Components Disabled: *      _traefik: Replaced with ingress-nginx for better control *      _servicelb: Replaced with MetalLB or cloud load balancer *      _local-storage: Replaced with Longhon for dynamic provisioning *      _metrics-server: Replaced with Prometheus for better monit...

ML Observability with eBpf and OTel Pt. 1: Basics | ML Engineering and MLOps

Image
  Hi All  In this first part of a series on eBPF and Open Telemetry . We detail a couple of useful eBPF scripts tailored for an MLE with no background of the project(s). They focus on observability, performance monitoring and data collection - key areas where eBPF shines. I tried to make these as practical as possible, let's get to it. 1. Tracing Python Function Calls in ML Pipelines Monitor which Python functions are being called in an ML training script (eg. PyTorch/Tensorflow) and their execution time. This helps identity bottlenecks in data loading, preprocessing or model training. Tools: *    bpftrace (for high level scripting) *    libbpf (for custom C-based eBPF programs) *    Run the below script while your ML script is executing. sudo bpftrace -e 'uprobe:python3:PyEval_CallObject { printf("Function called: %s\n", str(arg1)); }' *    Trace all torch.Tensor method calls in a PyTorch script *    The output will show ...

Deploying LoRA Optimised BERT as a FastApi service on GKE | ML Engineering & MLOps

Image
  Hi All  Alright today on "Bored MLE" we're doing some inference, using a LoRA optimized BERT model on GKE . Now there's a lot more to MLOps than just this (like cluster level observability for example), but for this example I'm keeping it brief.  The full code is available on GitHub with added minikube deployment instructions (I only cover the Cloud deployment here).    I assume familiarity with Python, PyTorch, K8s, FastApi, Minikube and GKE. Let's get on with it.  View the full FastApi service source below, also available here : Let's breakdown the above code, block by block. 0. Install Dependencies :  pip install fastapi uvicorn torch transformers peft accelerate  1. Imports from fastapi import FastAPI, HTTPException from pydantic import BaseModel from typing import List import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer from peft import PeftModel, PeftConfig  *      fastapi : F...

Fine-Tuning Mistral 7B using QLoRA with PyTorch pt. 1: The Model | ML Engineering

Image
     Hi All Today we're working with a popular and slightly bigger model than our previous example. Mistral 7B is capable of chat and light coding tasks, for older hardware it's a winner for sure.  Here's a complete, runnable example of fine-tuning Mistral 7B using QLoRA with the peft , transformers , and bitsandbytes libraries. This example assumes you're working with a single GPU (eg. an A100 or similar). First install the required packages: pip install -q bitsandbytes datasets accelerate peft transformers trl View full script below, also available here :   Full breakdown of the script above, block-by-block. 1.      Dataset Loading dataset = load_dataset("timdettmers/openassistant-guanaco", split="train") *      Loads a preprocessed instruction-following dataset (Guanco, derived from OpenAssistant). *      split="train" selects the training portion *      The dataset is in a conversational ...