Posts

Showing posts with the label FastApi

Deploying LoRA Optimised BERT as a FastApi service on GKE | ML Engineering & MLOps

Image
  Hi All  Alright today on "Bored MLE" we're doing some inference, using a LoRA optimized BERT model on GKE . Now there's a lot more to MLOps than just this (like cluster level observability for example), but for this example I'm keeping it brief.  The full code is available on GitHub with added minikube deployment instructions (I only cover the Cloud deployment here).    I assume familiarity with Python, PyTorch, K8s, FastApi, Minikube and GKE. Let's get on with it.  View the full FastApi service source below, also available here : Let's breakdown the above code, block by block. 0. Install Dependencies :  pip install fastapi uvicorn torch transformers peft accelerate  1. Imports from fastapi import FastAPI, HTTPException from pydantic import BaseModel from typing import List import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer from peft import PeftModel, PeftConfig  *      fastapi : F...