OSORIO.SYS

Lexical Analysis with Logos

In the Compiler Internals series we’re building a tiny language from scratch.
This part focuses on the lexer: turning raw text into a stream of tokens.

Why start with the lexer?

The lexer is usually the first layer of structure above raw bytes:

  • It strips away whitespace and comments.
  • It normalizes identifiers, keywords, and operators.
  • It gives the parser a clean, typed stream to work with.

A surprisingly large chunk of “compiler weirdness” disappears once your tokens are:

  • Well-defined.
  • Easy to inspect in tests.
  • Easy to dump to logs when something goes wrong.

Defining tokens with Logos

Logos is a derive-based lexer generator for Rust.
We start with an enum describing the tokens of our language:

rust
use logos::Logos;

#[derive(Logos, Debug, Clone, PartialEq)]
pub enum Token {
    // Keywords
    #[token("let")]
    Let,
    #[token("fn")]
    Fn,
    #[token("if")]
    If,
    #[token("else")]
    Else,
    #[token("return")]
    Return,

    // Identifiers and literals
    #[regex(r"[a-zA-Z_][a-zA-Z0-9_]*")]
    Ident,
    #[regex(r"[0-9]+")]
    Int,

    // Punctuation
    #[token("=")]
    Eq,
    #[token("==")]
    EqEq,
    #[token("+")]
    Plus,
    #[token("-")]
    Minus,
    #[token("{")]
    LBrace,
    #[token("}")]
    RBrace,
    #[token("(")]
    LParen,
    #[token(")")]
    RParen,

    // Whitespace & comments (skipped)
    #[regex(r"[ \t\n\f]+", logos::skip)]
    #[regex(r"//[^\n]*", logos::skip)]
    #[error]
    Error,
}

The nice part is that:

  • The enum is just Rust.
  • We can pattern-match on Token in tests and in the parser.
  • Regexes live close to the variant they describe.

Running the lexer

Creating a token stream from a &str is straightforward:

rust
pub fn lex<'a>(input: &'a str) -> impl Iterator<Item = Token> + 'a {
    Token::lexer(input).spanned().map(|(tok, _span)| tok)
}

We can write a tiny snapshot test:

rust
#[test]
fn lexes_simple_function() {
    let src = r#"
        fn add(x: i64, y: i64) {
            return x + y
        }
    "#;

    let tokens: Vec<_> = lex(src).collect();

    insta::assert_debug_snapshot!(tokens);
}

This gives us a human-readable, version-controlled view of whatever the lexer is doing.

Handling errors without losing information

Error tokens

Don’t immediately panic on the first Token::Error.
For a better developer experience, collect them with spans and show all of them in a single diagnostic pass.

A simple strategy:

  1. Keep a Vec<Diagnostic> alongside the token stream.
  2. Whenever you see Token::Error, push a new diagnostic with its span.
  3. Let the parser keep going until it really can't.

Later, the diagnostics renderer can highlight those spans in the source.

Measuring lexer performance

Lexing is often not the bottleneck, but it still pays to keep an eye on it. For a toy language, something like this is enough:

rust
use std::time::Instant;

fn bench_lexer(input: &str, iters: usize) {
    let start = Instant::now();

    for _ in 0..iters {
        let _ = lex(input).count();
    }

    let elapsed = start.elapsed().as_secs_f64();
    println!("lexed {} bytes in {:.3}s", input.len() * iters, elapsed);
}
cargo run --release --bin bench_lexer
lexed 12000000 bytes in 0.042s

This is more than enough signal to see if a “cute refactor” accidentally tanks performance.

What’s next?

In the next article we’ll plug this lexer into a Pratt parser and start building an AST:

  • Expressions with proper precedence.
  • A minimal statement grammar.
  • Enough structure to start generating LLVM IR.

If you’d like to peek ahead, the full series is listed under the Compiler Internals module.