hjLabs AI Playground
Starting the on-device runtime…
ViT-GPT2 gives a short factual caption in a 250 MB download; Mozilla's DistilViT is roughly a quarter the size for the same one-line style, and a flower-tuned variant handles botanical subjects. The obvious use is accessibility — writing alt text for a large image library without paying per call or exposing unpublished assets — and the second is bulk cataloguing. Everything runs on your GPU, so a folder of a thousand product shots costs nothing but time.
image caption generatorai alt text generatordescribe image ai freevit-gpt2 demodistilvit browserlightweight captioning modelautomatic alt textimage to text description
ViT-GPT2 is the most reliable general choice; DistilViT is about a quarter of the size if you are captioning in bulk and can accept slightly plainer wording.
Not here. These captioners describe an image but take no text prompt, so they cannot answer questions about it. Vision-language models that can are not yet supported by the browser runtime.