Laya on AWS Lambda - A Curiosity-Driven Experiment with a Decision Model

Note: This is a curiosity-driven experiment, not a production recommendation. Cold starts are long, the ARM64 path doesn't work, and confidence scores aren't calibrated. For real workloads, look at Amazon SageMaker, Amazon Bedrock, or a warm, right-sized container service.
1. The idea
First I ran a 1.58-bit LLM on Lambda. Then an embedding model. Both worked, with the usual serverless caveats. That left one more kind of model I was curious about: one that doesn't generate text or vectors, but makes a decision.
Laya is exactly that. It's an open-weight (Apache-2.0), 421M-parameter model that reads some text, answers a typed question about it, and returns a bounded answer in a single forward pass. Could that run on Lambda, CPU only, with no inference server? Only one way to find out.
