Based in Belgrade · Serbian and Croatian/EU citizen

Uroš Savurdić

M.Sc. EECS, University of Belgrade · previously Applied Science Intern at Microsoft

I work on machine learning systems from both ends — building them, and finding out what they actually do. Retrieval pipelines, evaluation harnesses and containerized services on one side; controlled experiments, causal interventions and preregistered analysis on the other.

The through-line is measurement: whether a change actually did anything, and how you would know. That question is the same whether the answer ships or gets written up, which is why the work runs across both — LLM evaluation at Microsoft, a controlled study of what preference optimization does to a model’s safety representation, and performance measurement across CPU, GPU and TPU.

Latest

  1. Finished Internal Representations of Safety Across Post-Training — nine checkpoints, a frozen 654-prompt benchmark, preregistered analysis.

  2. Finished the Applied Science internship at Microsoft — LLM evaluation, knowledge graphs and multi-hop reasoning.

  3. Gave a talk on attention across computational paradigms at SICAAI, the 5th Serbian International Conference on Applied Artificial Intelligence, in Kragujevac.

  4. Started as Research Collaborator at IPSI Belgrade, with Prof. Veljko Milutinović.

Experience

Experience

Applied science at Microsoft, and an ongoing research collaboration at IPSI Belgrade.

Microsoft — Applied Science Intern

Mar 2026 – Jul 2026
LLMsevaluationknowledge graphsmulti-hop reasoning

ResultA consensus of graph-based and surface-level features agreed with human complexity labels 81.2% of the time, against 64.6% for a standard hop-count baseline.

  • Built a multi-view evaluation framework combining graph-based and surface-level features; its consensus score reached 81.2% agreement with human-annotated complexity labels, against 64.6% for a standard hop-count baseline — a 16.6-point gain.
  • Developed an end-to-end pipeline generating evidence-grounded, controllable multi-hop QA datasets from enterprise documents using knowledge graphs and LLMs.
  • Ran ablations across extraction, retrieval, canonicalization and graph-construction strategies to find what actually drives QA generation quality.
A consensus of graph and surface features tracks human judgement far better than hop count. Multi-hop question difficulty, scored against human-annotated complexity labels.
0%25%50%75%100%Multi-view consensusMulti-view consensus: 81.2% agreement with human labels81.2%Hop-count baselineHop-count baseline: 64.6% agreement with human labels64.6%+16.6 pts

Applied Science internship, Microsoft. Hop count is the standard proxy for multi-hop difficulty; it is a real signal but an incomplete one.

Table view
ScorerAgreement
Multi-view consensus81.2%
Hop-count baseline64.6%

IPSI Belgrade — Research Collaborator

Feb 2026 – present
with Prof. Veljko Milutinović
  • Working with Prof. Veljko Milutinović on a comparative survey of emerging computing paradigms for AI workloads.
  • Running the literature review and contributing technical review for the survey and an accompanying book.
Projects

Projects

Six projects. Filter by the kind of work you care about.

Optimization Methods

Nov – Dec 2025

Backpropagation written from scratch in NumPy and raced against a genetic algorithm, plus 2-opt, simulated annealing and differential evolution.

Research

Research and engineering, in detail

Three projects with the measurements behind them. Every figure has a table view underneath with the exact numbers.

Independent researchPyTorch, LoRA, TRL, interpretability

FindingAblating the refusal direction changed behavior in all four training branches, but the size of that change varied by roughly 13x depending on the path taken — and the branch with the largest behavioral shift carried the smallest direction-specific effect.

QuestionDoes preference optimization create a new internal safety representation, or modify one that is already there?
SetupNine Qwen2.5-1.5B checkpoints across instruction tuning, safety SFT and DPO, crossed over two instruction corpora with a direct-DPO control.
ResultIn this setup the refusal direction was largely established before DPO; DPO mainly changed its magnitude, and how much it mattered depended on the training path.
CaveatsOne model, one scale (1.5B), one seed per cell. The transfer study is post hoc and outside the preregistration. No multiplicity correction.
  • Trained a nine-checkpoint chain on Qwen2.5-1.5B — base → helpful SFT → safety SFT → DPO — crossed over two instruction corpora, with a direct-DPO arm that skips safety SFT as the control. Nine LoRA checkpoints, matched preference data, one factor changed at a time.
  • Fixed the benchmark, the direction-estimation procedure and two confirmatory behavioral endpoints in a preregistered analysis plan before looking at any outcome, then recorded every deviation from it.
  • Found the harmful/benign contrast forms during instruction tuning rather than DPO (adjacent-stage cosine 0.65 → 0.96 → 0.93), and that DPO amplifies it instead of adding structure: residualize the contrast away and the remaining space carries no linearly decodable harm category, +0.004 with a 95% CI of [-0.018, +0.027].
  • Measured the direction's causal weight by ablation against a magnitude-matched random control, cross-fitted out-of-fold, n=120 per branch. Every interval excludes zero, so the direction is load-bearing in all four — but the effects span roughly thirteen-fold, and the corpus × training-history interaction is +0.097 [+0.040, +0.150].
  • Ran a cross-branch transfer study over all 12 ordered pairs: injecting one branch's DPO activation delta into another's pre-DPO model moves behavior toward that branch's post-DPO profile, class-conditioned rather than prompt-specific, surviving per-row norm matching, and not reducible to a single refusal direction.
  • Built the whole pipeline: training, generation, a refusal classifier combining a frozen regex scorer with two LLM judges, activation capture, intervention tooling, cross-fitted statistics, CI, and a provenance sidecar binding every committed number to the benchmark hash it came from.
Sharpening an existing axis drags the ambiguous cases with it. Harmful prompts with the wrongdoing vocabulary removed, measured against each stage’s own refusal direction. Skip safety SFT and preference optimization overshoots — those prompts end up past the overtly harmful cluster.
instruction tuning → safety SFT → DPODPO straight from instruction tuning
benign clusterovertly harmful cluster0.00.51.0BaseInstructiontuningSafetySFTDPOBase: 0.33Instruction tuning: 0.72Safety SFT: 0.65DPO: 0.90Direct DPO: 1.040.901.04

Layer 24. 0 = benign cluster, 1 = overtly harmful cluster.

Table view
StagePosition
Base0.33
Instruction tuning0.72
Safety SFT0.65
DPO0.90
DPO direct from instruction tuning1.04
The mechanism is load-bearing everywhere — but thirteen times more in one branch than another. Ablating the refusal direction, cross-fitted out-of-fold. Every interval excludes zero, so the direction matters in all four. How much depends entirely on the training path.
via safety SFTDPO direct
0.000.050.100.150.20effect vs magnitude-matched random controlSafety SFT then DPO, AlpacaSafety SFT then DPO, Alpaca: +0.154, 95% CI [+0.105, +0.203]+0.154Safety SFT then DPO, DollySafety SFT then DPO, Dolly: +0.044, 95% CI [+0.005, +0.085]+0.044DPO direct from M1, AlpacaDPO direct from M1, Alpaca: +0.025, 95% CI [+0.013, +0.039]+0.025DPO direct from M1, DollyDPO direct from M1, Dolly: +0.012, 95% CI [+0.003, +0.023]+0.012

n=120 per branch, 95% confidence intervals. Post hoc, one seed per cell, not multiplicity-corrected.

Table view
BranchEffect95% CI
Safety SFT then DPO, Alpaca+0.154[+0.105, +0.203]
Safety SFT then DPO, Dolly+0.044[+0.005, +0.085]
DPO direct from M1, Alpaca+0.025[+0.013, +0.039]
DPO direct from M1, Dolly+0.012[+0.003, +0.023]
Benchmarking and profilingPyTorch, JAX, NumPy

FindingOne PyTorch call, 92x apart. scaled_dot_product_attention silently dispatches to the math fallback in FP32 and the memory-efficient kernel in FP16 — 104 against 9,633 GFLOP/s on the same T4, with the chosen backend logged on every run to prove it.

  • Benchmarked attention across 12 computational paradigms and 17 implementations, decomposed into five stages and measured for FLOPs, memory traffic and arithmetic intensity. Eight implementations ran on hardware I had access to — CPU, an NVIDIA T4 and a TPU v5e; the rest are simulators or analytical cost models.
  • Traced a 92x throughput gap in a single PyTorch call to silent backend dispatch between FP32 and FP16, confirmed by logging the selected kernel on every run.
  • Built the harness as an extensible package: an abstract paradigm interface, a registry decorator, dataclass metrics and a structured experiment logger.

Presented at the 5th Serbian International Conference on Applied Artificial Intelligence (SICAAI), Kragujevac, May 2026.

Common-scale comparison: 14 implementations at n=2048, d=256. Roughly seven orders of magnitude separate a TPU from a Turing machine. The FLOP count is the same; the execution cost is not.
measured on hardwaresimulator or cost model
10⁻⁴10⁻³10⁻²10⁻¹10⁰10¹10²10³10⁴GFLOP/sTPU, fused (JAX/XLA)TPU, fused (JAX/XLA): 15,772 GFLOP/s (measured)15,772GPU FP16, decomposedGPU FP16, decomposed: 11,327 GFLOP/s (measured)11,327GPU FP16, fused SDPAGPU FP16, fused SDPA: 9,633 GFLOP/s (measured)9,633TPU, decomposed (per-op JIT)TPU, decomposed (per-op JIT): 4,908 GFLOP/s (measured)4,908Dataflow (Maxeler-style)Dataflow (Maxeler-style): 977 GFLOP/s (modeled)977Optical (photonic MZI)Optical (photonic MZI): 510 GFLOP/s (modeled)510GPU FP32, decomposedGPU FP32, decomposed: 199 GFLOP/s (measured)199GPU FP32, fused SDPAGPU FP32, fused SDPA: 104 GFLOP/s (measured)104CPU, decomposed (AVX2)CPU, decomposed (AVX2): 18 GFLOP/s (measured)18CPU, fused SDPACPU, fused SDPA: 18 GFLOP/s (measured)18GaAs RISC-VGaAs RISC-V: 0.15 GFLOP/s (modeled)0.15IoT edge (ARM M4F)IoT edge (ARM M4F): 0.11 GFLOP/s (modeled)0.11WSN, 64 nodesWSN, 64 nodes: 0.01 GFLOP/s (modeled)0.01CdTe Turing machineCdTe Turing machine: 0.0007 GFLOP/s (modeled)0.0007

Blue rows ran on hardware I had access to — CPU, an NVIDIA T4 and a TPU v5e. Orange rows are simulators or analytical cost models. Where the roofline model can be checked against a measurement of the same thing it does badly: it put decomposed FP32 GPU at 6,592 GFLOP/s where the T4 delivers 199.09, a 33× overestimate. So the models are order-of-magnitude arguments about architecture, not predictions of what hardware would do. The full benchmark covers 12 paradigms and 17 implementations; a paradigm is the technology, and CPU, GPU and TPU each carry more than one implementation. This chart shows the 14 that share n=2048, d=256 — the chemical, biological and quantum simulators only run at n=32, d=16 and are reported separately.

Table view
ImplementationGFLOP/sSource
TPU, fused (JAX/XLA)15,772measured
GPU FP16, decomposed11,327measured
GPU FP16, fused SDPA9,633measured
TPU, decomposed (per-op JIT)4,908measured
Dataflow (Maxeler-style)977modeled
Optical (photonic MZI)510modeled
GPU FP32, decomposed199measured
GPU FP32, fused SDPA104measured
CPU, decomposed (AVX2)18measured
CPU, fused SDPA18measured
GaAs RISC-V0.15modeled
IoT edge (ARM M4F)0.11modeled
WSN, 64 nodes0.01modeled
CdTe Turing machine0.0007modeled
PyTorch Lightning, FAISS, Sentence-Transformers

FindingContrastive fine-tuning moved MRR@10 from 0.786 to 0.911 on CoSQA — a 16% relative gain, with NDCG@10 and Recall@10 up alongside it.

  • Fine-tuned a bi-encoder (all-MiniLM-L6-v2) with InfoNCE contrastive loss on 20k CoSQA query-code pairs, improving MRR@10 from 0.786 to 0.911 (+16.0%), with NDCG@10 +12.1% and Recall@10 +2.8% over the pretrained baseline.
  • Enabled semantic search across 550+ indexed code snippets with FAISS and cosine similarity.
  • Set up a reproducible training and evaluation pipeline with PyTorch Lightning, AdamW with warmup, Weights & Biases tracking, and a torchmetrics evaluation suite.
Contrastive fine-tuning, measured three ways. The same bi-encoder before and after InfoNCE training on 20k CoSQA query-code pairs: +0.125 MRR@10, a 15.9% relative gain.
pretrained baselineafter fine-tuning
0.70.80.91.0MRR@10MRR@10 before: 0.786MRR@10 after: 0.9110.911NDCG@10NDCG@10 before: 0.831NDCG@10 after: 0.9320.932Recall@10Recall@10 before: 0.969Recall@10 after: 0.9960.996

Evaluated with torchmetrics on the held-out split. Recall@10 was already high, so the gain lands in ranking quality rather than coverage.

Table view
MetricBeforeAfter
MRR@100.7860.911
NDCG@100.8310.932
Recall@100.9690.996
Background

Background

Where I studied, what I volunteer at, the tooling I actually reach for, and what has been recognized.

Volunteering

Volunteer with EESTEC, the European EECS student association — 34 universities and 4,000+ members — since February 2019. It is where I learned to run teams and teach.

EESTEC — Soft Skills Trainer

Aug 2021 – present
  • Delivered 100+ hours of training across 15+ topics to 400+ participants from across Europe — leadership, emotional intelligence, team dynamics, nonviolent communication.
  • Qualified through a 12-day intensive Training for Trainers, then two years delivering under mentorship before being certified by the EESTEC Training Team in 2023.

EESTEC — Training Team Coordinator, LC Belgrade

Feb 2022 – May 2023
  • Owned the soft skills training program for the Belgrade branch — what got taught, by whom, and to whom — and helped organize the educational program for Soft Skills Academy Belgrade, a self-development project open to any student.
  • Mentored around 30 members across two generations; several went on to Chairperson, fundraising coordinator, and international project roles. Supported the branch's HR processes and worked with the board on team development and conflict resolution.

EESTEC — Regionalization Coordinator, then Project Leader

Aug 2019 – Aug 2021
  • Led a 30-member international team through the COVID-19 transition, moving the program's events online, and restructured the project — merging two international teams into one and redefining how they operated.
  • Ran the Balkan regionalization program: online regional meetings and an in-person one hosted with LC Skopje, on how committees actually get run — applying for grants, fundraising, recruitment and HR practice.

EESTEC — PR Team Leader and Event Organizer

2019
  • Led PR for Job Fair, one of Serbia's largest student career events — media outreach and coverage, reporting to the event coordinator.
  • Main organizer of a local student exchange from September to December 2019, and ran the Belgrade branch's social media as PR assistant to the board.
First page of a recommendation letter from the EESTEC international board, on headed paper.Recommendation letterEESTEC International Board · Zürich, August 2025Christa Ward, Chairperson · Atina Velinovska, Vice Chairperson for Internal AffairsRead it — PDF, 2 pages

Topics I trainFacilitation and meeting design · feedback and nonviolent communication · emotional intelligence · team dynamics · presentation · strategic and design thinking · problem solving and creativity · training design · learning from experience

Education, skills and recognition

Education

expected Sep 2027M.Sc. Electrical Engineering & Computer ScienceUniversity of Belgrade — School of Electrical Engineering (ETF) · GPA 10/10 · evaluation, system security, optimization
graduated Sep 2024B.Sc. Electrical Engineering & Computer ScienceUniversity of Belgrade — School of Electrical Engineering (ETF) · signal processing, information theory · thesis: optimizing urban traffic congestion using data-based crowdsourcing systems
Online courseworkFull Stack Deep Learning UC Berkeley · CME295: Transformers & Large Language Models Stanford · An Introduction to Statistical Learning Stanford · Machine Learning Workshop Polytechnic University of Madrid, 2019

Technical skills

LanguagesPython, C/C++, SQL
MLPyTorch, Transformers, Hugging Face, TRL, LoRA/PEFT, JAX
Interpretabilityactivation capture and steering, directional ablation, linear probing, representation geometry
Statisticspreregistration, cross-fitted out-of-fold estimation, bootstrap CIs, McNemar tests
ToolsDocker, AWS, Git, CI/CD, Weights & Biases
SpokenSerbian/Croatian — native · Macedonian — native · English — fluent · Russian and Bulgarian — understood · German — basic

Awards & talks

AwardEESTECer of the Year — Board of the Association, one recipient per year across a 4,000+ member network (2021)
Award3rd place, National Mathematics Competition — Mathematical Society of Serbia (2017)
AwardBest Student Start-Up Award — Science-Technology Park Belgrade (2022)
TalkAnalysis of Attention Mechanisms in the Context of Computational Paradigms — presented at the 5th Serbian International Conference on Applied Artificial Intelligence (SICAAI), Kragujevac, May 2026.
Additional recognition (3)
  • Certified Soft Skills Trainer — EESTEC Training Team (2023)
  • 3rd place, Case Study Competition — BEST Belgrade (2018)
  • National silver medal, judo — junior level