We help teams test and deliver AI products: LLM outputs, agent workflows and the QA and CI/CD process around them. We turn “it seems to work” into measurable quality.
Traditional testing assumes the same input gives the same output. AI doesn’t, so it needs different methods, and we apply them.
AI Application Development
Custom AI-powered applications: LLM integrations, retrieval-based search, copilots and agents, built with evaluation and testing from day one.
LLM and retrieval (RAG) application design
Copilots, agents and AI-driven workflow automation
Guardrails, monitoring and cost controls
AI Quality Assurance & LLM Evaluation
AI features give different answers on every run. We define what good looks like, then test and score it automatically.
Evaluation datasets and scoring criteria
Accuracy, safety and tone checks
Regression gates for model and prompt changes
Agentic Workflow Testing
Multi-step agents fail in ways single-prompt tests miss. We test the whole flow, including tool calls and handoffs.
End-to-end conversation and tool-call tests
Guardrails and failure-path validation
Reproducible traces for every failure
AI-Assisted Test Automation
We use AI tools such as Claude Code and Cursor to write, debug and maintain test suites faster, with human review on every change.
AI-assisted test authoring and debugging
Custom agents and skills for your team's standards
AI-assisted review of test code
AI Readiness & Process Audit
Before adding AI, we check whether your processes, data and team are ready, and where AI will actually save time.
Use cases ranked by value and risk
Process and data readiness review
Prioritized roadmap and pilot plan
Architecture, legal & delivery services
Architecture Consulting
Reviews of system and AI architecture for scalability, security and maintainability, with a clear modernization roadmap.
Software Legal Consultation
Technical and compliance readiness for software contracts, open-source licensing and AI data use, in coordination with licensed counsel.
QA Automation & Strategy
Test strategy and automated suites for web, API, iOS and Android.
DevOps & CI/CD Pipelines
Pipelines with quality gates that give developers fast feedback.
Agile & Dev Process Audit
A review of how your team plans, builds, reviews and releases software.
LegalBytes is not a law firm and does not provide legal advice. Software legal consultation is delivered in coordination with licensed attorneys.
How we work
From AI experiment to dependable product
The goal isn’t more AI. It’s AI your team can stand behind in production. Select a step to see what we do.
Step 1 of 4
Find where it matters
We map the workflow and the AI features involved, then decide what is worth automating or validating first. Some work should not be automated, and we will say so.
Built on experience
Over a decade of quality engineering, now applied to AI
Our work covers fintech, banking and mobile products, and now includes testing LLM features and conversational flows.
12+
Years in QA & test automation
AI
LLM features, chat and agent flows
Web · API · Mobile
iOS and Android coverage included
Daily
AI tools in our own test workflow
Our toolkit
AI-assisted development
AI tools that speed up test authoring, debugging and review, always with human approval.
Claude Code
Cursor
GitHub Copilot
MCP integrations
Custom agents & skills
AI code review
AI apps & evaluation
Building AI applications and testing outputs that change between runs.
OpenAI API
Anthropic API
LangChain
LLM evaluation
Prompt engineering
Conversational flow testing
AI-generated content checks
Test automation
Frameworks for web, API and mobile end-to-end testing.
Playwright
Selenium
Appium
Serenity BDD
Cucumber
RestAssured
Kotlin
Java
Python
TypeScript
Performance, CI/CD & test management
Load testing, pipelines and the tools that track quality.
K6
JMeter
Postman
Jenkins
Azure DevOps
Docker
AWS
Allure
Jira
Xray
TestRail
Quality is what separates an AI demo from an AI product
A demo that works once is easy. A feature that keeps working as models, prompts and users change needs testing built for the job. That is what we do.
LegalBytes was founded in August 2023 by Jhonny Alejandro Pérez Rojas and Luz Adriana Ruiz González.
Hands-on, not just strategy
We write the tests, evaluations and pipeline changes ourselves, then hand them over to your team.
Built for non-deterministic systems
Traditional pass/fail testing is not enough for AI. We design checks that measure quality across varied outputs.
Your team keeps ownership
We work with your developers and QA and leave behind processes they can run without us.
Our team
The people who write the tests and evaluations
Senior engineers who build the checks themselves, then hand them over to your team.
Jhonny Alejandro Pérez Rojas
Co-founder · Senior SDET · AI Quality Engineering
12+ years in QA and test automation across fintech, banking and mobile products. Now focused on AI quality: validating LLM features, conversational flows and AI-generated content.
Senior SDET focused on AI quality. Designs and builds test automation frameworks for AI-driven software, coordinates testing across agile teams, and researches new AI and testing technologies. Hands-on with SerenityBDD, RestAssured, Cucumber, Selenium, Appium, K6 and JMeter, and with CI/CD pipelines on Jenkins and Azure DevOps. Master's degree from Pontificia Universidad Javeriana.
Lawyer with a specialization in computer law and new technologies (Universidad Externado de Colombia). Independent legal consultant since 2021, focused on labor and technology law. Provides legal consultancy on software contracts, with experience reviewing and drafting contracts. At LegalBytes, she leads software legal consultation.
Tell us about an AI feature you don’t fully trust, or a process that is slower than it should be. Book a free 30-minute review and we will tell you honestly where we can help.