small language models (slms) on android with llama.cpp
small language models are getting really good, and very tiny. from last year, google already had been using gemma 3 1b as on-device models that even have tool-calling capabilities. apple has also launched openelm, a family of models targeting local inferencing, with models as small as 270 million parameters.
for fun, here are some tokens per second benchmarks for a few slms running on my samsung galaxy a15 (an almost 3 years old phone).