BACK TO PORTFOLIO
AI INFRASTRUCTURE2026

GEMMA LOCAL API SERVER

A specialized side project built as internal developer tooling to support RMIT Capstone offline AI testbenches. Packages google/gemma-4-E2B-it locally with FastAPI, Hugging Face Transformers, and CUDA acceleration.

USERS

Capstone Project Team

LAUNCH TIME

1 week

IMPACT

100% Offline Capstone Testing

PERFORMANCE

Quantised 4-bit CUDA Runtime

Mini Case Study

PROBLEM

Testing AI features for the Capstone project directly against cloud APIs was costly, internet-dependent, and slowed down development velocity.

BUILD

Engineered a local microservice with FastAPI and Hugging Face Transformers to run a quantized Gemma-4 model on local CUDA GPUs with standard OpenAI-compatible endpoints.

RESULT

Provided a dedicated offline development testbench for the capstone team, eliminating cloud API costs and speeding up feature verification.

Tech Stack

PythonFastAPIGemma-4CUDAPyTorchHugging Face