Q4 is smallest. FP16 is sharper but ~2× the download.
Load on start
Off means you tap to load each time.
Replies
Instructions
Sent before every conversation. Short and specific works best on a small model.
Creativity
0 is repeatable. Higher wanders.
Reply length
Maximum tokens per answer.
Memory
Earlier turns kept in context.
Voice
Speak replies
Reads answers aloud as they arrive.
Voice engine
System voices need nothing extra. On-device voices download once and then work offline.
Voice
Voices available in the selected engine.
Speaking rate
Send after speaking
Sends as soon as you stop talking.
Recognition
On-device works without a network or Google services.
Data
Stored on this device
Chats stay in this app. Nothing is uploaded.
About
JARVIS runs a Liquid AI LFM2.5 language model directly on this device with Transformers.js
and WebGPU. Prompts and replies never leave the phone. A small model gets facts wrong —
check anything that matters.
JARVIS
build 1.3
A language model that runs on this phone. The first load downloads about 200 MB, then it works offline.