← back to case filesCASE FILE · 08 / 09
StyleNova product screenshot
product snapshot ↗
VISION-LANGUAGE · RECOMMENDERS

StyleNova

StyleNova fuses what an outfit looks like, what a shopper has liked before, and what similar shoppers bought into a single hybrid recommendation score. CLIP vision-language embeddings, TF-IDF content signals, and collaborative filtering feed the base model, while an online preference-learning loop keeps adjusting to each shopper in real time.

Stack

AI + ML
PyTorchCLIPscikit-learn
Backend + Data
PythonFastAPIPrisma
Frontend
Next.jsTypeScriptZustandFramer Motion

The Problem

Vision-language fashion recommender that learns a shopper's taste from just a handful of interactions.

What I Built

StyleNova fuses what an outfit looks like, what a shopper has liked before, and what similar shoppers bought into a single hybrid recommendation score. CLIP vision-language embeddings,…

The Difference

Hybrid scorer (CLIP ViT-B/32 + TF-IDF + collaborative filtering) reaches Precision@10 of 0.73

Numbers That Matter

32primary scale
10measured result
0.73system signal
32%quality outcome

Engineering Notes ✎

01

Engineering Note 1

Hybrid scorer (CLIP ViT-B/32 + TF-IDF + collaborative filtering) reaches Precision@10 of 0.73

02

Engineering Note 2

Online preference learning improves recommendation relevance 32% after just 5 user interactions

03

Engineering Note 3

Cached CLIP embeddings deliver sub-200ms p95 API response for 10K+ concurrent users with zero cold-start

Architecture Overview

Taste signalsHybrid Recommender
CLIP EmbeddingsTF-IDF ScorerPreference Loop
Prisma StorePersonalized looks