Representation Engineering · AI Monitoring

Maciej Chrabąszcz

Struggling with the Polish spelling?

Put together these common English sounds and you have it:

Maciek (Ma-ciek)
“Ma” (grandma) + “ch” (cheese) + “y” (yes) + “eck” (neck) → ma-ch-y-eck
Chrabąszcz (Chra-bąszcz)
“h” (hat) + “rah” (hurrah) + “bon” (James Bond) + “sh” (shoe) + “ch” (church) → h-rah-bon-sh-ch

Fun fact: Chrabąszcz means “May Beetle” in Polish. 🪲

I study the hidden representations of generative models: how to read them, steer them and monitor them, and how that can help make these systems safer.

PhD candidate, Warsaw University of Technology
AI Safety Researcher, NASK National Research Institute

Maciej Chrabąszcz

About

I am a PhD candidate at Warsaw University of Technology and an AI Safety researcher at NASK National Research Institute. My thesis, supervised by Prof. Tomasz Trzciński and Dr. Sebastian Cygert, is on efficient machine-learning methods for the safety of AI models.

My work centres on the hidden representations of generative models. A model's internal activations often carry useful information about what it is doing, so I look for ways to read them, steer them and monitor them. In practice that covers representation engineering, probing, activation steering and AI monitoring, approaches that treat the latent space as a useful place to work on safety.

Part of what interests me here is cost. Guardrails built from activations the forward pass has already computed can be cheaper than running a separate moderation model, and monitoring a reasoning trace as it unfolds may surface problems earlier than filtering the final output.

At NASK I co-led the safety and evaluation work for PLLuM, the Polish government's family of Polish large language models. I hold an MSc in Mathematical Statistics and Data Analysis from WUT.

Research

Representation Engineering

Internal activations often carry signal about what a model is doing. I work on reading those representations, and on writing to them: steering behaviour by transporting activations rather than retraining the model or filtering its outputs.

AI Monitoring

Observing a model while it works rather than after. Probes applied across layers and along a reasoning trace turn hidden state into a signal that can be followed as the model generates.

Safety from Activations

Guardrails built from the model's own latent space. Multi-layer latent prototypes estimate whether an input is safe using activations the forward pass has already computed, which avoids running a separate moderation model.

Publications

Google Scholar

* denotes equal contribution.

2026

2025

2024

2023

Get in touch

Happy to talk about AI safety, representation engineering or possible collaborations. Email is the best way to reach me.