428 research journals · as of Oct 9, 2026

Work less,
keep what we learn

KALTI (칼퇴공학연구소) is an independent research group that experiments firsthand with working alongside AI. We log every experiment as a research journal, and keep only the conclusions worth reusing as shared knowledge.

  • Project
  • Journal
  • Finding
  • Hypothesis
  • Decision
  • Concept
This graph links our actual research records. Projects form the outer ring; at the center are conclusions that several projects reached.
428Research journalsSince October 2025
493Knowledge cardsIncl. 264 findings, 95 hypotheses
37Projects30 active · 5 on hold · 2 done
3ResearchersFounded June 8, 2026

About

We try working with AI
and write it all down

KALTI is a research group formed to study software development and engineering together. At the founding meeting on June 8, 2026, the members adopted a charter and elected a representative. The name comes from "kaltoe" (칼퇴), Korean slang for leaving work right on time.

  • We run the experiments

    Each researcher owns their projects, building and measuring things firsthand. The range is wide, from memory for AI agents to devices you hold in your hand.

  • We write everything down

    Each day's experiments, research and decisions go into a research journal. Journals are evidence, so no one but the author edits them.

  • We keep only the conclusions

    Only conclusions worth reusing become cards. Ideas that proved wrong stay too, with a note on why we dropped them.

Research journals so far

428
As of Oct 9, 2026
42

Each bar is one week: the research journals written each week since March 2026. The busiest week is labeled.

Research

Five areas of research

The projects are independent of each other, but most of them touch on how to work alongside AI.

See all projects

AI agents & memory

We help agents build memory across sessions and bring back only what fits the moment. We also build a runner that wires several agents together on one screen so we can watch them work.

  • neuromem
  • mnemosure
  • numen
  • honcho-selfhost
  • Human-memory-based memory
  • Memory systems integration

6 projects · 114 journals

LLM evaluation

We measure what shape of data helps a model answer well, and whether code structure helps AI coding. First, we doubt whether the measuring tool itself is honest.

3 projects · 33 journals

Developer tools & automation

We build tools that AI uses, or that we use together with AI: browser control, design-to-code verification, consistency measurement.

8 projects · 71 journals

Things we build

A dictation app, a composing environment, a puzzle game, 3D printing, a handheld voice input device: things we actually use, built together with AI.

  • wisp
  • aria
  • locogic
  • 3d-fabrication
  • toython
  • ride-draft
  • sulnote
  • ktheme-works

13 projects · 105 journals

Learning together

We study by building a RAG chatbot that answers from documents, and join hackathons to tackle real company problems with AI.

2 projects · 13 journals

Projects

What we're working on now

The bars under each card count research journals per week since March 2026. As the records pile up, so do the conclusions.

See all 32
AI agents & memoryActive

neuromem

A long-term memory engine that lets an AI agent keep who it is and what it has been through, and recall the memories that fit the moment by association.

MarWeekly journalsOct
Journals42Cards48Last2026.09.26
AI agents & memoryActive

mnemosure

A memory layer that cuts forgotten context and hallucination in AI coding across sessions. It says "I don't know" when it doesn't, and answers with sources when it does.

MarWeekly journalsOct
Journals14Cards30Last2026.09.07
AI agents & memoryActive

numen

A multi-agent runner that launches Claude Code, Codex and Gemini CLI unmodified as boxes on a canvas, wires them together, and shows who started when and what failed.

MarWeekly journalsOct
Journals15Cards22Last2026.08.09
LLM evaluationActive

Serialization formats and token accuracy

Feeds the same data to LLMs in seven formats, including JSON, YAML, CSV and Markdown tables, and traces why accuracy and token counts change.

MarWeekly journalsOct
Journals9Cards11Last2026.08.21
Developer tools & automationActive

agrune

A browser-control layer with no LLM of its own. Any agent connects over MCP and handles web pages by meaning instead of screen coordinates.

MarWeekly journalsOct
Journals21Cards25Last2026.08.16
Developer tools & automationActive

dsforge

Verifies by machine measurement, not human judgment, whether the path from Figma mockups to a design system and code stays faithful.

MarWeekly journalsOct
Journals11Cards14Last2026.09.11

Method

From journal to knowledge,
in one line

What we did is written down that day; what we learned is collected separately. Keeping evidence and knowledge apart lets us trace, months later, why a decision was made.

  1. 01428

    Research journals

    Each researcher writes up the day's experiments, research and decisions in their own folder.

  2. 02493

    Knowledge cards

    Findings, hypotheses and decisions are pulled from journals into cards. Every card points to its source journal.

  3. 0313

    Concepts

    Only conclusions that several projects reached independently are grouped into concepts.

  4. 04Weekly· monthly

    Reports

    A weekly report from each researcher, and a monthly digest for the whole lab.

Five Claude Code skills
run the lab

How to write a journal, refine it into cards and produce reports is packaged as Claude Code skills. Everyone follows the same rules, and when the rules change, only the skills change.

kalti-lab/claude-plugins
# once
❯ /plugin marketplace add kalti-lab/claude-plugins
❯ /plugin install kalti-lab-notes@kalti-lab
❯ /kalti-setup      # vault, own folder, sync

# every day
❯ /kalti-journal    # today's work as a journal
❯ /kalti-ontology   # pull conclusions into cards
❯ /kalti-report     # weekly report · monthly digest
❯ /kalti-context    # what we already know on a topic
❯ 

Findings

Conclusions several projects
reached on their own

A conclusion becomes a concept only when projects reach it separately, without knowing about each other. Each dot is one project; colored dots are the projects that reached that conclusion.

0116 projects · 35 findings

Honest evaluation

Before trusting a conclusion, the tool that measures it has to be honest first.

0216 projects · 26 findings

Silent failure

Failures that look like success, with no error, surface only weeks later.

0313 projects · 22 findings

Limits of verification gates

A green check proves only what that check measures.

0411 projects · 15 findings

Replacement cost and boundaries

Hide what is likely to change behind one layer, and swapping it later is cheap.

057 projects · 15 findings

Long-term memory recall

Build memory across sessions, and bring back only what fits right now.

067 projects · 9 findings

Pinning the runtime

If you don't pin down what actually ran, you blame the subject for what the environment did.

  • Judging by appearance5 projects

    Judge by name or format, and you miss that two things are the same.

  • Axes that don't split5 projects

    If sweeping an axis end to end keeps results in the same range, that axis is not a lever.

  • Trusting LLM judges5 projects

    How far to trust a model when it grades work without an answer key.

  • Failures live in the handoffs4 projects

    Symptoms that look like weak performance often come from how parts pass things to each other.

  • The receiver sets the format3 projects

    The maker decides the content, but the receiving app or device decides the format.

  • Model tier decides the design3 projects

    A conclusion that a method works is tied to the model it ran on.

  • Origin checks for local services3 projects

    Assume a service started for development can be reached from outside, and check where requests come from.

News

Monthly research digest

Once a month, the three researchers' work goes onto two A4 pages. Stories are picked by how much each project changed that month, and every number comes from a single command that counts the repository.

KALTI MonthlyIssue 2 · September 2026

Silent checks
didn't mean all clear

In September we looked hardest at the journal repository itself. Where the checks reported nothing wrong, we kept finding what they had missed, and the cause was always the same: a rule that exists only in writing is kept only if someone remembers it.

  • 56rules a machine could catch that the check code never verified
  • 19banned words left in places the checks never looked
  • 30 / 33journals passed over as "nothing to extract" that held a conclusion on a second read
Journals
31
New findings
51
New decisions
25
Settled hypotheses
2

Published Oct 2, 2026 · Featured: research journal system

KALTI MonthlyIssue 1 · August 2026

We thought making the model reason would erase the format gap

When we had the model write out its reasoning, the accuracy gap between formats didn't shrink; it grew. Of the ten hypotheses settled in August, five turned out to be wrong.

Journals
49
New findings
44
New decisions
1
Settled hypotheses
10

Published Sep 11, 2026 · Featured: serialization formats and token accuracy

46 weekly reports

Every week, each researcher writes up the highlights, what they did in each project, and what carries over to the next week.

History

How we got here

  1. First research journal

    The oldest record we still have.

  2. Founding meeting

    We founded the group, approved its charter and elected a representative.

  3. Domain registered

    kalti.co.kr

  4. GitHub organization

    We set up the research journal repository and the Claude Code skills repository.

  5. Monthly digest launched

    We published the first issue (August).

  6. Monthly digest, issue 2

    We published the September issue.

  7. Website launched

    We passed 400 research journals.

People

Researchers

Three people each run their own projects and look after the conclusions together. Below are the projects each of them has written the most journals for.

aram

Representative · Researcher

  • neuromem
  • honcho-selfhost
  • llm-bench
  • agrune
  • ride-draft
  • numen

jinsik

Researcher

  • locogic
  • mnemosure
  • Research journal system
  • dsforge
  • sulnote
  • dsmonitor

sunghyun

Researcher

  • Serialization formats and token accuracy
  • Research record framework
  • Human-memory-based memory

FAQ

Frequently asked questions

Short answers to what you might want to know first about KALTI.

What is KALTI (칼퇴공학연구소)?

KALTI is an independent research group in South Korea that experiments with working alongside AI. It was founded at a founding meeting on June 8, 2026, where the members adopted a charter and elected a representative. Three researchers each run their own projects and log everything as research journals.

What does the name mean?

It comes from the Korean slang "kaltoe" (칼퇴), which means leaving work right on time. The English name is KALTI, and the website is kalti.co.kr.

What does KALTI research?

32 public projects across five areas: AI agents and memory, measuring and evaluating LLMs, developer tools and automation, things we build, and learning together. Examples include a long-term memory engine for AI agents (neuromem), an experiment measuring how data formats change LLM accuracy, and a browser-control layer that agents use over MCP (agrune).

How is the research recorded?

Each day's experiments, research and decisions go into a research journal. Findings, hypotheses and decisions are pulled from journals into knowledge cards, and only conclusions that several projects reached independently become concepts. Each researcher writes a weekly report, and the lab publishes a monthly digest. As of October 9, 2026, there are 428 journals, 493 knowledge cards and 13 concepts.

What has KALTI found so far?

The three conclusions reached by the most projects: a measuring tool must be honest before you trust its result (honest evaluation, 16 projects); failures that look like success, with no error, surface only weeks later (silent failure, 16 projects); and a green check proves only what that check measures (limits of verification gates, 13 projects).

Can I see the records and tools?

The raw journals live in a private repository shared by the researchers; this website publishes per-project summaries and counts. The Claude Code skills used to write and refine the journals are public at github.com/kalti-lab/claude-plugins.

Where do the numbers on this site come from?

Every number is counted directly from the research journal repository each time the site is built. The current figures are as of October 9, 2026.

How can I contact KALTI?

By email at aram@kalti.co.kr, for research questions, collaboration or requests for materials.

Want to research with us,
or have a question?

Questions about our research, collaboration and requests for materials are all welcome by email.