Avalia

Independent evaluations of AI models in Brazilian Portuguese. We compare limits and risks where they matter. In real Brazil.

Loading results…

How we evaluate

IDJÉ Avalia rankings run on Régua: our proprietary multiple-choice evaluation harness. Same answer reading, same scoring rule, with care for Brazilian Portuguese.

Question

Each test item is a multiple-choice question with two to ten options.

Answer reading

Régua identifies which option the model chose, the same way every time.

↓ offline scoring, no LLM judge ↓

What counts

Splits valid answers, refusals, invalid replies and errors. Clear what enters each rate.

Test metric

Accuracy, prudence, bias: each evaluation defines what to measure. Counting follows Régua.

This site only shows the ranking. Scoring runs offline and does not call models here. Each test defines what to measure; counting follows Régua.