SpeakLocal — an AI language tutor that runs on your phone
The fastest way to learn a language is to speak it badly to somebody patient. This is that, running on your phone after a one-time download, so you can practise on a plane and nobody records the fumbling.
Speaking is the skill nobody practises
You can finish a flashcard app's entire tree and still freeze when someone speaks to you. Reading and listening are practised constantly because they are comfortable; speaking is avoided because it is not, and because it requires a partner willing to sit through your first hundred bad sentences.
An AI tutor solves the partner problem. But a tutor that needs a live server connection stops working the moment you are on a plane, on bad signal, or out of data. SpeakLocal downloads the model once and runs it on the phone instead.
What a session with the offline AI language tutor looks like
You pick a scenario, speak, and get a response plus feedback. Everything below serves that loop.
- Scenario conversations. Practice happens in situations rather than in the abstract: ordering in a coffee shop, asking directions, introducing yourself, and — in the Pro set — a job interview, a doctor's visit, hotel check-in, a first date, a business meeting, filing a complaint, a phone call with no visual cues, an emergency call. Each has its own register and vocabulary.
- Three languages, three tutor personas. English, Mandarin Chinese and Japanese, each with its own prompt template and locale so the tutor behaves appropriately for the language rather than being a translated English bot.
- Feedback on every turn. Each thing you say comes back with one of three responses — Good, Almost, or Try again — plus a short tip naming what to fix. It compares what your phone transcribed against what the scenario expected, which makes it a practice aid rather than a measurement of pronunciation accuracy. That distinction is stated inside the app too.
- A pronunciation drill bank. Eight target sentences per language, played back by the device's text-to-speech so you can hear the target before you attempt it, with the same three-level feedback.
- A vocabulary vault with spaced repetition. Words you save go into a vault scheduled by an SRS algorithm, reviewed with the familiar Again / Hard / Good / Easy grading, with a due-today count on the home screen.
- Progress you can see. A confidence chart, streaks and per-session summaries, plus a daily goal set during onboarding.
- A data inspector. Settings has a panel listing every file the app has written to your device — the model, the encrypted settings store, the progress database — with its size, and how much of the total is encrypted.
One connection, and it is a file download
SpeakLocal connects to GitHub once to download the language model, and after that the tutor runs on the phone — you can finish the download, switch the internet off, and keep practising. That is the whole of the app's own network use. Two SDKs make their own separate requests: Google AdMob on the free plan, and Google Play Billing for purchases. Both are named in the privacy policy, because a privacy claim that quietly omits the ad SDK is not a privacy claim.
Your speech is turned into text by your phone's own speech engine; the app does not record audio files and does not upload audio. Progress and vocabulary are held in an encrypted local store, and there is an Erase all my data action that does exactly that.
Who it is for
Intermediate learners stuck at the point where they understand far more than they can produce. Twenty minutes of daily speaking with immediate correction moves that faster than another month of vocabulary drills.
Anyone who freezes when it is their turn to speak. The value here is volume of attempts without an audience — the tutor will sit through the same scenario twenty times and never sigh.
And anyone preparing for a specific event — a job interview in a second language, a medical appointment abroad, a business meeting — who wants to rehearse that exact conversation a dozen times without an audience.
Privacy and requirements
The language model (about 550 MB) is downloaded once during onboarding from GitHub, after which conversation runs on the device. Speech recognition and text-to-speech use the platform's own engines. Audio is processed locally and not uploaded. Data lives in encrypted local storage and can be erased from Settings. Full details are in the privacy policy, which also carries the Gemma model licence terms.
The app is free, with a banner ad above the navigation bar and an occasional full-screen ad after a session; a Pro subscription removes both and unlocks the wider scenario set. It is a practice tool, not a certified assessment.
Frequently asked questions
Is my voice sent to a server? Your speech is turned into text by your phone's own speech engine and the tutor model runs on the device. The app does not record audio files and does not upload audio. Its only content request over the network is the one-time model download.
Which languages can I practise? English, Mandarin Chinese and Japanese, each with its own tutor persona and locale.
How does the feedback work? Each turn gets Good, Almost, or Try again, plus one short tip. It compares what your phone transcribed against what was expected — a practice aid, not a measurement of pronunciation accuracy. There is deliberately no numeric score, because the on-device model is not calibrated well enough for one to mean anything.
Does it work without an internet connection? Yes, once the model has been downloaded during setup. Conversation, feedback and vocabulary review all run offline after that.
Can I delete everything? Yes. Settings has an Erase all my data action, plus a data inspector listing every file stored on the device and how much of it is encrypted.