How good is Clef on OrganAMNIST?

Using Cloudflare Clef to classify 17,778 CT crops from OrganAMNIST without any training, and how it compares to other models.

· 4 min read

A CT crop with the eleven possible answers around it
A kidney crop from OrganAMNIST with the eleven labels that Clef can choose from.

In the body part post, Clef worked well for finding the body part of a slice but I wanted a real number for how good it is on medical image classification on public datasets and benchmarks. I could not find any image benchmark from Cloudflare so I ran one myself. I used OrganAMNIST from MedMNIST with both Clef models and the same kind of one-line prompt.

The dataset

OrganAMNIST is one of the twelve 2D datasets in MedMNIST. MedMNIST is a collection of small and standardized medical images in the spirit of the MNIST digits. OrganAMNIST is made from about 200 abdominal CT scans from the LiTS liver tumor challenge where eleven organs are boxed in 3D. Each image is a slice through the middle of a box. The slice is cropped to the box and windowed for the abdomen. Then it is shrunk to a small square and labeled with the organ.

Boxing the organs in 3D then slicing through the middle of the liver box and cropping it into one labeled sample.

This gives 58,830 images in total. 34,561 are for training and 6,491 for validation. The other 17,778 are for testing and come from 70 scans that are not in the training set. I used the 128 × 128 version.

A grid of real OrganAMNIST crops
Real crops from OrganAMNIST at 128 × 128 pixels.

The task

Each image has to be classified as one of eleven organs. Five of them are the bladder, heart, liver, pancreas and spleen. The other six are the left and right femur, kidney and lung.

The eleven classes
The eleven classes with the easy ones in green and the left and right kidney in yellow.

Some of them are easy like the heart and the liver. Others are hard because the left and right kidney look almost the same and a femur crop is only a few blurry pixels. Random guessing gets 9.1%. Liver is the most common organ so always answering liver gets 18.5%.

How I ran it

  1. Images from the test split are sent as 128 × 128 PNGs without any change. There is one request for each image and each model.
  2. Classification is done via Clef on Workers AI in a Cloudflare worker which receives the image and is asked one multiple choice question. The instruction is "Classify this image." and the choices are the eleven label names. There is no mention of CT and no example or description.
  3. Result is the label that Clef chooses together with the probability of each label. It is saved to a CSV file for every image.

This is the whole request:

{
  "model": "clef-flash",
  "state": "",
  "questions": {
    "label": {
      "type": "choice",
      "instructions": "Classify this image.",
      "criteria": { "bladder": null, "femur-left": null, "…": null, "spleen": null }
    }
  },
  "images": ["data:image/png;base64,…"]
}
Clef returns a probability for each label and picks the most likely one. For this crop it is heart.

I ran both clef and the smaller and cheaper clef-flash on all 17,778 test images. That is 35,556 requests in about an hour and $1.83 in total.

Results

clef-flash got 45.0% of the test images right and clef got 34.1%.
clef-flash clef
Accuracy, 11 classes 45.0% 34.1%
Accuracy with left and right merged 53.4% 42.2%
Macro AUC 0.883 0.833
Median latency 210 ms 323 ms
Cost of the whole test set $0.50 $1.33

The smaller model was better by eleven points. It was also faster and 2.7 times cheaper. The heart was right 86% of the time. The liver was right 71% of the time and the pancreas 61%. The bladder and the femurs were almost never right.

Accuracy per organ for both models
Accuracy for each organ with clef-flash in blue and clef in yellow.

Left and right look more like a habit than something read from the image. When clef-flash says kidney it says left three times out of four. Clef says right four times out of five. The real split is close to half and half.

Share of left and right answers on the kidney images
The share of left and right answers when the models say kidney next to the real split.

How it compares

Clef next to published zero-shot and supervised results
Clef compared with chance, medical CLIP models, the best zero-shot results and supervised models.
Model Setup Accuracy
ResNet-18, 224 px trained on the 34,561 training images 95.1%
MGLL (ICLR 2026) zero-shot, pretrained on biomedical figures 52.7%
FG-CLIP zero-shot, pretrained on biomedical figures 47.9%
clef-flash zero-shot, one-line prompt 45.0%
clef zero-shot, one-line prompt 34.1%
BiomedCLIP zero-shot 25.5%
Always "liver" 18.5%
Random guess 9.1%

The clef-flash model is within eight points of the best zero-shot result I could find with only a one-line prompt and no training on this dataset. It is also well above BiomedCLIP. A model trained on the data still roughly doubles the accuracy. The papers use different image sizes, prompts and preprocessing so this is more of a rough ranking than a leaderboard.

The main numbers from the run
The main numbers from the run.

Takeaways

  • Zero-shot with a bare prompt is competitive with the medical CLIP models which were built for this.
  • Speed and cost are good at about 0.2 s for each image and $0.50 for the whole test set.
  • Left and right are a weak spot because each model has a side that it prefers.
  • Training still wins if you have labels. A ResNet is about 50 points better.
  • The smaller model was better than the bigger one.

One dataset is done. You can also watch the narrated 4-minute version.


Data is from MedMNIST v2 by Yang et al. in Scientific Data 2023 under CC BY 4.0. The supervised results are from the MedMNIST v2 benchmark. FG-CLIP and MGLL are from Table 24 of the MGLL paper at ICLR 2026. BiomedCLIP is from Table 1 of "Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs" from 2026. The latency was measured from my machine through a local worker. The figures are from the explainer video which is narrated with Kokoro-82M running locally.