AI research & development labAI R&D labEst. 2020

We break AI models
to learn how to
defend them.

trifero is an independent lab working inside the weights of language, vision, speech and agent models. We build them, fine-tune them, strip out their guardrails through abliteration, and heal what that breaks. Underneath it all is philosophy: epistemology and logic, turned on the assumptions everyone else takes for granted.

  • Development
  • Fine-tuning
  • Abliteration
  • Healing
Fig. 01 trifero.ai Est. 2020

01 · Philosophy

Every model is an argument.
We check the premises.

Philosophy drives the research as much as engineering does. Epistemology asks how we know what we claim to know. Logic asks whether the conclusion actually follows. Put to AI, those two questions expose the hidden assumptions that most failures grow from.

Epistemology

How do we know?

Every claim a model makes, and every claim we make about a model, needs a warrant. We ask what the evidence is, how it was gathered, and what result would change our minds.

  • Calibration & uncertainty
  • Evaluation design
  • Data provenance

Logic

Does it follow?

We test whether conclusions follow from their premises: in a model's reasoning, in a benchmark's design, and in our own write-ups. Validity comes before persuasion.

  • Reasoning faithfulness
  • Consistency checks
  • Counterexample search
Hidden assumptions we challengeThe question we ask instead
  1. 01A refusal means the model is safe.Is the safety in the weights, or only in the wording?
  2. 02A high benchmark score means the model is capable.Does the test measure the skill, or familiarity with the test?
  3. 03Fine-tuning only adds what you train.What did it quietly take away?
  4. 04The model knows when it does not know.Is its confidence calibrated to the evidence?
  5. 05If it reasons step by step, the steps are the reason.Do the stated steps actually cause the answer?

02 · Research

Four disciplines.
Every kind of model.

The disciplines are how we work on a model. Language models are where they were sharpened, and the same methods now reach vision, speech, agents and systems that learn by evolution.

01

Development

Building models and the tooling around them: architectures, training runs from scratch, data pipelines, and the evaluation harnesses that show what a model can and cannot do.

  • Pretraining from scratch
  • Evaluation harnesses
  • Vision, speech, evolved nets
02

Fine-tuning

Adapting open-weight models to a domain, a task or a hardware budget, then proving the result against the base model on the same tests.

  • Supervised & preference tuning
  • Adapters & distillation
  • Quantization for the edge
03 · Offense

Abliteration

Finding the directions in a model's activations that mediate refusal, then editing the weights to remove them. The most direct way to learn how thin a safety layer really is.

  • Refusal-direction analysis
  • Weight orthogonalization
  • Guardrail robustness testing
04 · Defense

Healing

Recovering the capability that abliteration, quantization or pruning takes from a model. Targeted retraining, measured against the original so nothing is claimed that was not tested.

  • Post-edit capability recovery
  • Regression benchmarks
  • Drift & side-effect checks

Where we apply them

A · 01

Speech & audio

Text-to-speech, voice cloning and source separation, built and run on the lab's own GPUs.

We probeCloned voices and spoofed speakers.

A · 02

Vision

Image understanding, detection and motion analysis, from production vision services to models we train ourselves.

We probeAdversarial inputs and confident misreads.

A · 03

On-device models

Quantized models that run offline on phones and edge hardware, with no cloud in the loop.

We probeWhat compression quietly breaks.

A · 04

Retrieval & knowledge

Systems that answer from curated corpora and show where every answer came from.

We probePoisoned and misleading sources.

A · 05

Agents & tool use

Models that act through tools, including autonomous red-team agents that chain an attack from first foothold onward.

We probePrompt injection and over-broad permissions.

A · 06

Evolutionary & game AI

Neuroevolution, self-play and search: the line of work the lab started with, still running.

We probeReward hacking and brittle strategies.

03 · Method

Offense informs
defense.

Safety behavior in a model lives in its weights, and weights can be edited. We run the attacks ourselves, in a controlled lab, so the people deploying these systems learn what fails, what holds, and what it costs to fix.

  1. Step 01

    Question

    Name the assumption the system depends on, and decide in advance what result would prove it wrong.

  2. Step 02

    Break

    Abliterate, jailbreak and tamper with the model under controlled conditions, the way a real attacker would.

  3. Step 03

    Measure

    Quantify what was exposed and what was lost: refusal behavior, capability benchmarks, side effects.

  4. Step 04

    Harden

    Turn each finding into a defense: guardrails, monitoring, and guidance for deploying the model safely.

  5. Step 05

    Heal

    Restore the capability the attack cost, then rerun the original attack to verify the fix holds.

Repeat until the fix holds

Ground rules

Controlled environments

Offensive work runs on the lab's own hardware, against models the lab controls.

Findings, not exploits

What leaves the lab is understanding: failure modes, measurements, and the mitigations that answer them.

Measured, not asserted

Every claim about a model is backed by a before-and-after evaluation on the same tests.

04 · Since 2020

It started with
evolution.

The lab's research began in 2020 with networks that grow their own structure and with production computer vision. The models have grown since. The method has not: take the system apart to understand it.

  1. 2020

    inininout +node

    NeuroEvolution of Augmenting Topologies

    Evolving the structure and the weights of neural networks together, starting from minimal topologies and adding nodes and connections only when they make the network better.

  2. 2020

    mountainclock?

    Microsoft Computer Vision

    Working with Microsoft's Computer Vision service to understand how a production vision model sees and labels an image, and where it gets it wrong.

  3. Now

    Language models and beyond

    Development, fine-tuning, abliteration and healing across language, vision, speech and agent models. Philosophy sets the questions. Offensive research tests the answers.

Find out how your model fails before someone else does.

For research collaboration or questions about the lab's work, get in touch.

Contact the lab [email protected]