Talorys 는 채팅, 기억, 할 일, 노트, 예약 알림이 든 개인 AI 비서를 내 Cloudflare 계정 안에 명령 한 줄로 배포하는 오픈소스다. 고른 이유는 "내 비서는 내 계정에서 돈다"는 요구를 무료 플랜 범위 안에서 끝까지 밀어붙였고, 그 선택이 설계 전체를 어떻게 비틀었는지가 코드에 그대로 남아 있어서다. 돌려보면 문서와 다른 무엇이 나올지도 궁금했다.
자료가 전부 Durable Object(요청이 올 때만 깨어나는, 저장소가 딸린 작은 서버) 하나에 들어간다. SQLite 테이블 열다섯 개와 FTS5 색인이 그 안에 있고, KV, D1, R2, Vectorize 는 쓰지 않는다. 설치 스크립트가 만드는 wrangler 설정에도 바인딩은 Workers AI 와 Durable Object 둘뿐이다. 이 선택의 대가가 기억이다. 매 턴 모델에 들어가는 기억은 사용자의 문장을 두 글자 이상 단어로 쪼개 불용어를 뺀 뒤 접두사 OR 질의로 FTS5 를 치고 BM25 로 순위를 매긴 결과 최대 8개에, 최근 수정된 선호 4개를 얹은 것이다. 임베딩도 벡터 색인도 없다. 그래서 "기억"은 의미 검색이 아니라 키워드 검색이고, 어휘가 다르면 못 찾는다. 대신 외부 서비스가 0 이고, 알림은 같은 객체의 알람으로 돌아가므로 무료 플랜의 Worker 하나로 비서 전체가 닫힌다. 창업자가 가져갈 문장은 이것이다. 기억의 품질을 한 단계 양보하면 개인 비서의 운영 비용이 서버 없이 0 에 가까워진다.
커밋 열두 개가 전부 2026-10-10 하루에 있고, 열두 개 모두 Co-Authored-By 트레일러에 Claude Opus 5.5 가 적혀 있다. Hono 라우터와 Cloudflare Agents SDK 위에 짰고, 모델은 Workers AI 의 glm-4.7-flash 다. README 는 만든 과정을 말하지 않는다.
npm install 에 16초, npm run dev 로 Worker, Pages 함수, Vite 가 한 번에 뜬다. Cloudflare 계정 없이 돌아가는 건 로컬에서 모델을 결정적 모의 공급자로 바꿔 끼우기 때문이다. 그러니 이날 확인한 건 모델의 답이 아니라 그 바깥, 즉 도구 호출과 저장과 스케줄이다. "remember that …" 은 기억 페이지에 선호로 남았고, "add task …" 는 할 일 목록에 떴고, "list tasks" 가 그걸 읽어 왔다. "remind me in 1 minutes" 는 알림 센터에 도착했다. 다만 117초 뒤였다. 코드가 분 단위로 자르면서 59초를 더하는 탓에 1분 알림은 1분 넘게 남은 다음 정분에 걸린다. 설계지 버그는 아니다. README 의 주장 셋(알림은 AI 를 쓰지 않는다, 바깥으로 보내는 요청이 없다, 유료 서비스를 만들지 않는다)은 코드와 실행에서 모두 맞았다. Worker 번들은 gzip 602 KiB 였다.
Total Upload: 3190.12 KiB / gzip: 601.88 KiB
--dry-run: exiting now.
Workers AI 의 실제 답과 하루 할당량이 바닥났을 때의 동작은 이 환경에서 볼 수 없었다. FTS5 의 기본 토크나이저가 조사가 붙은 한국어 기억을 얼마나 되찾는지는 다음에 tal_memories_fts 에 한국어 문장을 넣고 질의를 섞어 재 볼 일이다.
무료 플랜에 비서를 닫아 넣은 값으로 기억이 키워드가 됐고, 그 거래를 코드가 숨기지 않는다.
Talorys is an open-source personal AI agent with chat, memory, tasks, notes, and scheduled reminders, deployed into your own Cloudflare account with one command. I picked it because it pushes the demand "my assistant runs in my account" all the way to the edge of the free plan, and the code shows how that choice bent the whole design. I also wanted to see whether running it would turn up something the docs do not say.
Everything lives in one Durable Object, a small server with its own storage that wakes only when a request arrives. Fifteen SQLite tables and an FTS5 index sit inside it, and KV, D1, R2, and Vectorize are not used. The wrangler config the installer generates has only two bindings, Workers AI and the Durable Object. Memory pays for this choice. The memories that reach the model each turn come from splitting the user's sentence into words of two or more characters, dropping stopwords, running a prefix OR query against FTS5, ranking by BM25, and keeping at most eight hits, plus the four most recently updated preferences. There are no embeddings and no vector index. So "memory" is keyword search, not semantic search, and different wording means no match. In return there are zero external services, reminders run on the same object's alarms, and one Worker on the free plan closes the loop around the entire assistant. The sentence for founders is this: give up one step of memory quality and the running cost of a personal assistant approaches zero, with no server to keep.
All twelve commits landed on 2026-10-10, and every one carries a Co-Authored-By trailer naming Claude Opus 5.5. It is written on the Hono router and the Cloudflare Agents SDK, and the model is glm-4.7-flash on Workers AI. The README says nothing about how it was built.
npm install took 16 seconds, and npm run dev brought up the Worker, the Pages function, and Vite in one go. It runs without a Cloudflare account because the local build swaps the model for a deterministic mock provider. So what I verified was not the model's answers but everything around them: tool calls, storage, and scheduling. "remember that …" landed on the memory page as a preference, "add task …" showed up in the task list, and "list tasks" read it back. "remind me in 1 minutes" reached the notification center, 117 seconds later. The code truncates to the minute after adding 59 seconds, so a one-minute reminder lands on the next whole minute that is more than a minute away. That is a design choice, not a bug. The README's three claims (reminders never use AI, nothing is sent outside, no paid services are provisioned) held in both the code and the run. The Worker bundle was 602 KiB gzipped.
Total Upload: 3190.12 KiB / gzip: 601.88 KiB
--dry-run: exiting now.
Real answers from Workers AI, and what happens when the daily allocation runs out, could not be observed in this environment. How well the default FTS5 tokenizer recalls Korean memories with particles attached is something to measure next time, by putting Korean sentences into tal_memories_fts and mixing up the queries.
Memory became keyword search as the price of sealing an assistant into a free plan, and the code does not hide the trade.