Core API¶
The core module provides the main user-facing API for LayoutLens.
LayoutLens Class¶
- class layoutlens.LayoutLens(api_key: str | None = None, model: str = 'gpt-4o-mini', provider: str = 'openai', output_dir: str = 'layoutlens_output', cache_enabled: bool = True, cache_type: str = 'memory', cache_ttl: int = 3600, api_base: str | None = None, temperature: float | None = None)[source]¶
Bases:
objectSimple API for AI-powered UI testing with natural language.
This class provides an intuitive interface for analyzing websites and screenshots using natural language queries, designed for developer workflows and CI/CD integration.
Examples
>>> lens = LayoutLens(api_key="sk-...") >>> result = lens.analyze("https://example.com", "Is the navigation clearly visible?") >>> print(result.answer)
>>> # Compare two designs >>> result = lens.compare( ... ["before.png", "after.png"], ... "Are these layouts consistent?" ... )
- __init__(api_key: str | None = None, model: str = 'gpt-4o-mini', provider: str = 'openai', output_dir: str = 'layoutlens_output', cache_enabled: bool = True, cache_type: str = 'memory', cache_ttl: int = 3600, api_base: str | None = None, temperature: float | None = None)[source]¶
Initialize LayoutLens with AI provider credentials.
- Parameters:
api_key – API key for the provider. If not provided, will try OPENAI_API_KEY environment variable.
model – Model to use for analysis (LiteLLM naming: “gpt-4o”, “anthropic/claude-3-5-sonnet”, “google/gemini-1.5-pro”).
provider – AI provider to use (“openai”, “anthropic”, “google”, “gemini”, “litellm”).
output_dir – Directory for storing screenshots and results.
cache_enabled – Whether to enable result caching for performance.
cache_type – Type of cache backend: “memory” or “file”.
cache_ttl – Cache time-to-live in seconds (1 hour default).
api_base –
Optional base URL for an OpenAI-compatible endpoint. When set, it is passed as
api_baseto every model call, enabling self-hosted backends such as Ollama or vLLM. Example:LayoutLens( provider="litellm", model="ollama/qwen2.5vl", api_base="http://localhost:11434", )
temperature – Optional sampling temperature for the analyze path. When None (default) the analyze path uses 0.1 to preserve historical behavior. Subject to the per-model parameter policy — models that reject non-default sampling params (e.g. Claude Sonnet 5) omit it regardless.
- Raises:
ConfigurationError – If invalid provider or configuration is specified.
Notes
A missing API key does NOT raise here. The requirement is deferred to the first LLM call (see
_ensure_api_key()) so that deterministic, keyless operations such ascheck_accessibility(..., mode="axe")work without any credentials configured.
- async analyze(source: str | Path | list[str | Path], query: str | list[str], viewport: Viewport | str = 'desktop', context: dict[str, Any] | None = None, instructions: Instructions | None = None, max_concurrent: int = 5) AnalysisResult | BatchResult[source]¶
Smart analyze method that handles single or multiple sources and queries.
- Parameters:
source – Single URL/path or list of URLs/paths to analyze.
query – Single question or list of questions about the UI.
viewport – Viewport for capture (Viewport.DESKTOP, “desktop”, etc.).
context – Additional context for analysis (user_type, browser, etc.). Legacy format.
instructions – Rich instruction set with expert personas and structured context. Takes precedence over context if both provided.
max_concurrent – Maximum concurrent operations for batch analysis.
- Returns:
AnalysisResult for single source+query, BatchResult for multiple.
Examples
# Single analysis >>> result = await lens.analyze(”https://github.com”, “Is it accessible?”)
# Multiple queries on one source >>> result = await lens.analyze(”https://github.com”, [“Is it accessible?”, “Mobile-friendly?”])
# Multiple sources, one query >>> result = await lens.analyze([“page1.html”, “page2.html”], “Is it good?”)
# Multiple sources and queries >>> result = await lens.analyze([“page1.html”, “page2.html”], [“Accessible?”, “Mobile?”])
- async compare(sources: list[str | Path], query: str = 'Are these layouts consistent?', viewport: Viewport | str = 'desktop', context: dict[str, Any] | None = None, instructions: Instructions | None = None) ComparisonResult[source]¶
Compare multiple URLs or screenshots.
- Parameters:
sources – List of URLs or screenshot paths to compare.
query – Natural language question for comparison.
viewport – Viewport for captures (Viewport.DESKTOP or string).
context – Additional context for analysis.
instructions – Rich instructions for expert analysis.
- Returns:
Comparison analysis with overall assessment.
Example
>>> result = await lens.compare([ ... "https://mysite.com/before", ... "https://mysite.com/after" ... ], "Did the redesign improve the user experience?")
- async capture(source: str | Path | list[str | Path], viewport: Viewport | str = 'desktop', wait_for_selector: str | None = None, wait_time: int | None = None, max_concurrent: int = 3) str | dict[str, str][source]¶
Smart capture method that handles single or multiple sources uniformly.
- Parameters:
source – Single URL/path or list of URLs/paths to capture.
viewport – Viewport for capture (Viewport.DESKTOP, “desktop”, etc.).
wait_for_selector – CSS selector to wait for before capturing.
wait_time – Additional wait time in milliseconds.
max_concurrent – Maximum concurrent captures for multiple sources.
- Returns:
Returns screenshot path as string. Multiple sources: Returns dict mapping source to screenshot path.
- Return type:
Single source
Examples
# Single URL >>> path = await lens.capture(”https://example.com”) # Returns: “/path/to/screenshot.png”
# Multiple URLs >>> paths = await lens.capture([”https://site1.com”, “https://site2.com”]) # Returns: {”https://site1.com”: “/path1.png”, “https://site2.com”: “/path2.png”}
# HTML files >>> path = await lens.capture(“page.html”) >>> paths = await lens.capture([“page1.html”, “page2.html”])
# Existing images (validation) >>> path = await lens.capture(“screenshot.png”)
- async judge(image_path: str | Path, prompt: str, *, max_tokens: int | _Auto = AUTO, timeout: float = 120.0) JudgeResult[source]¶
Send
promptVERBATIM with an image and return a parsed verdict.This is the faithful judge interface for external evaluation harnesses (e.g. UIJudgeBench). Unlike
analyze(), LayoutLens adds NOTHING to the prompt: no system persona, no scaffolding, no appended JSON-format instruction. The caller owns the entire prompt, including any response contract. The call always hits the model (no caching) and honors the per-model parameter policy (Claude 4.6+/5 omit temperature).- Parameters:
image_path – Path to an existing image file (mime inferred from the extension:
.jpg/.jpeg-> JPEG, otherwise PNG).prompt – The exact text to send as the sole text block.
max_tokens – Maximum tokens to generate. Defaults to
AUTO, which resolves to 8000 for reasoning/thinking models (they spend thinking tokens inside this budget) and 300 otherwise. Pass an explicit integer to override.timeout – Per-call timeout in seconds (default 120 — reasoning models can take well over 30s on a single judgment).
- Returns:
JudgeResult with the parsed answer/confidence/rationale, the raw text, a refusal flag, per-model usage split, and the parse mode.
- Raises:
ValidationError – If
image_pathdoes not exist.AuthenticationError – If no API key is configured for a mapped provider.
- async judge_batch(requests: list[BatchRequest], *, max_tokens: int | _Auto = AUTO, resume: bool = True, manifest_path: str | Path | None = None, poll_interval: float = 10.0, poll_timeout: float = 86400.0) dict[str, JudgeResult][source]¶
Judge many image+prompt requests over a provider batch transport.
The batched counterpart to
judge(): each request sends its prompt VERBATIM with its image (byte-identical tojudge()), honors the reasoning-awaremax_tokensdefault and the per-model parameter policy, and is parsed into aJudgeResult. Batch APIs are ~50% cheaper and the right transport for bulk offline evaluation (e.g. UIJudgeBench). LayoutLens is thus the reference batched judge.The backend is chosen from
self.model:gemini/*(AI Studio) uses the google-genai inline batch (optional extralayoutlens[gemini]); every other model uses the litellm file-based batch (OpenAI, Anthropic, Vertex, Bedrock, …). Both are resumable via a manifest.- Parameters:
requests – The batch items. Each
idkeys its result. A request whose image is missing yields an"unknown"result rather than aborting the batch.max_tokens – Per-request token budget. Defaults to
AUTO(8000 for reasoning models, else 300); an explicit integer overrides.resume – When True (default), collect any prior jobs from the manifest first and submit only uncovered ids.
manifest_path – Where submitted job ids persist for resume. Defaults to a path under
output_dir/batchkeyed by the request-id set and model.poll_interval – Seconds between batch-status polls.
poll_timeout – Max seconds to wait for a single batch job.
- Returns:
{request_id: JudgeResult}for every request.- Raises:
AuthenticationError – If no API key is configured for a mapped provider.
ImportError – If a
gemini/*model is used withoutgoogle-genai.
- async check_accessibility(source: str | Path, viewport: Viewport | str = 'desktop', mode: Literal['hybrid', 'axe', 'llm'] = 'hybrid') AnalysisResult[source]¶
Accessibility check with deterministic axe-core, LLM vision, or both.
- Parameters:
source – URL or file path to analyze.
viewport – Viewport for capture/audit.
mode –
"hybrid"(default) runs deterministic axe-core WCAG A/AA checks and LLM vision analysis, forcing a “no” verdict when axe finds any violation."axe"runs axe-core only (no API key required, no LLM call)."llm"runs the legacy vision-only analysis.
- Returns:
AnalysisResult. In axe/hybrid modes
metadata["a11y"]holds the full axe report,metadata["mode"]the mode, andmetadata["engine"]the axe-core version.
- async check_mobile_friendly(source: str | Path) AnalysisResult[source]¶
Quick mobile responsiveness check.
- async check_conversion_optimization(source: str | Path, viewport: Viewport | str = 'desktop') AnalysisResult[source]¶
Check for conversion-focused design elements.
- async audit_accessibility(source: str | Path, standards: list[str] = None, compliance_level: ComplianceLevel | str = 'AA', viewport: Viewport | str = 'desktop', mode: Literal['hybrid', 'axe', 'llm'] = 'hybrid') AnalysisResult[source]¶
Professional accessibility audit using WCAG expert knowledge.
- Parameters:
source – URL or file path to analyze
standards – Accessibility standards to apply (default: WCAG 2.1, Section 508)
compliance_level – WCAG compliance level (ComplianceLevel.AA or string)
viewport – Viewport for analysis (Viewport.DESKTOP or string)
mode –
"hybrid"(default) combines deterministic axe-core checks with LLM analysis (axe violations force a “no” verdict)."axe"runs axe-core only (no API key required)."llm"runs the legacy vision-only audit. The axe run honorscompliance_level: A ->wcag2a, AA ->wcag2a``+``wcag2aa, AAA additionally includeswcag2aaa.
- Returns:
Detailed accessibility assessment with specific WCAG guidance
- Raises:
ValueError – If compliance_level is not a valid WCAG level
- async optimize_conversions(source: str | Path, business_goals: list[str] = None, industry: str = None, target_audience: str = None, viewport: Viewport | str = 'desktop') AnalysisResult[source]¶
Conversion rate optimization analysis using CRO expert knowledge.
- Parameters:
source – URL or file path to analyze
business_goals – Business objectives (e.g., reduce_cart_abandonment)
industry – Industry context for specialized recommendations
target_audience – Target audience for optimization focus
viewport – Viewport for analysis (Viewport.DESKTOP or string)
- Returns:
Detailed CRO recommendations with A/B testing suggestions
- async analyze_mobile_ux(source: str | Path, device_types: list[str] = None, performance_focus: bool = True) AnalysisResult[source]¶
Mobile UX analysis using mobile expert knowledge.
- Parameters:
source – URL or file path to analyze
device_types – Target devices (smartphone, tablet)
performance_focus – Include performance optimization analysis
- Returns:
Mobile-specific UX recommendations and optimizations
- async audit_ecommerce(source: str | Path, page_type: str = 'product_page', business_model: str = 'b2c', viewport: Viewport | str = 'desktop') AnalysisResult[source]¶
E-commerce UX audit using retail expert knowledge.
- Parameters:
source – URL or file path to analyze
page_type – Type of e-commerce page (product_page, checkout, homepage)
business_model – Business model (b2c, b2b)
viewport – Viewport for analysis (Viewport.DESKTOP or string)
- Returns:
E-commerce specific recommendations for conversion improvement
- async analyze_with_expert(source: str | Path, query: str, expert_persona: Expert | str, focus_areas: list[str] = None, user_context: dict[str, Any] = None, viewport: Viewport | str = 'desktop') AnalysisResult[source]¶
Analyze using a specific domain expert persona.
- Parameters:
source – URL or file path to analyze
query – Question to analyze
expert_persona – Expert to use (Expert.ACCESSIBILITY or string)
focus_areas – Specific areas to focus analysis on
user_context – Rich context about users and requirements
viewport – Viewport for analysis (Viewport.DESKTOP or string)
- Returns:
Expert-level analysis with domain-specific recommendations
- async compare_with_expert(sources: list[str | Path], query: str, expert_persona: Expert | str, focus_areas: list[str] = None, viewport: Viewport | str = 'desktop') ComparisonResult[source]¶
Compare multiple sources using domain expert knowledge.
- Parameters:
sources – List of URLs or file paths to compare
query – Comparison question
expert_persona – Expert to use for comparison (Expert.ACCESSIBILITY or string)
focus_areas – Specific areas to focus comparison on
viewport – Viewport for analysis (Viewport.DESKTOP or string)
- Returns:
Expert comparison with domain-specific insights
- create_test_suite(name: str, description: str, test_cases: list[dict[str, Any]]) UITestSuite¶
Create a test suite from specifications.
Each spec in
test_casesfollows the same shape as a YAML test case, including a requiredexpected_results(see theUITestCasedocstring for schema) — validated identically to, and via the same helper as,UITestSuite.from_dict.- Parameters:
name – Name of the test suite
description – Description of the test suite
test_cases – List of test case specifications, each requiring “name”, “html_path”, “queries”, and “expected_results”.
- Returns:
UITestSuite object
- Raises:
ValidationError – If any spec is missing
expected_results.
- async run_test_suite(suite: UITestSuite, parallel: bool = False, max_workers: int = 4) list[UITestResult]¶
Run a test suite and return results.
- Parameters:
suite – The test suite to run
parallel – Whether to run tests in parallel
max_workers – Maximum number of parallel workers
- Returns:
List of UITestResult objects
Result Classes¶
- class layoutlens.AnalysisResult(source: str, query: str, answer: str, confidence: float, reasoning: str, screenshot_path: str | None = None, viewport: str = 'desktop', timestamp: str = <factory>, execution_time: float = 0.0, metadata: dict[str, ~typing.Any] = <factory>)[source]¶
Bases:
objectResult from analyzing a single URL or screenshot.
- class layoutlens.ComparisonResult(sources: list[str], query: str, answer: str, confidence: float, reasoning: str, individual_analyses: list[~layoutlens.api.core.AnalysisResult] = <factory>, screenshot_paths: list[str] = <factory>, timestamp: str = <factory>, execution_time: float = 0.0, metadata: dict[str, ~typing.Any] = <factory>)[source]¶
Bases:
objectResult from comparing multiple sources.
- individual_analyses: list[AnalysisResult]¶
- __init__(sources: list[str], query: str, answer: str, confidence: float, reasoning: str, individual_analyses: list[~layoutlens.api.core.AnalysisResult] = <factory>, screenshot_paths: list[str] = <factory>, timestamp: str = <factory>, execution_time: float = 0.0, metadata: dict[str, ~typing.Any] = <factory>) None¶
Test Suite Classes¶
- class layoutlens.UITestCase(name: str, html_path: str, queries: list[str], viewports: list[str] = <factory>, metadata: dict[str, ~typing.Any] = <factory>, expected_results: dict[str, ~typing.Any] | None = None, expected_confidence: float = 0.7)[source]¶
Bases:
objectRepresents a single test case for UI testing.
expected_resultsdeclares what the analysis must assert against and is required (seeUITestSuite.from_dict). Schema:expected_results: answer: "yes" # or "no" — compared against the # parsed yes/no of the analysis answer contains: ["nav", "contrast"] # optional; each term must appear # (case-insensitively) in answer + reasoning
Both keys are individually optional, but at least one of them must be present — an empty or missing
expected_resultsis a load-time error.expected_confidencesets the minimumresult.confidencerequired in addition to any content assertions.
- class layoutlens.UITestSuite(name: str, description: str, test_cases: list[~layoutlens.api.test_suite.UITestCase], metadata: dict[str, ~typing.Any] = <factory>)[source]¶
Bases:
objectRepresents a collection of test cases.
- test_cases: list[UITestCase]¶
- classmethod from_dict(data: dict[str, Any]) UITestSuite[source]¶
Create test suite from dictionary.
- Raises:
ValidationError – If any test case is missing
expected_results(or declares neither “answer” nor “contains”). Assertions are required per case — there is no confidence-only fallback.
- classmethod load(filepath: Path) UITestSuite[source]¶
Load test suite from JSON file.
- class layoutlens.UITestResult(suite_name: str, test_case_name: str, total_tests: int, passed_tests: int, failed_tests: int, results: list[~layoutlens.api.core.AnalysisResult], duration_seconds: float, metadata: dict[str, ~typing.Any] = <factory>)[source]¶
Bases:
objectResults from running a test suite.
- results: list[AnalysisResult]¶