Ghillie emits structured log events for ingestion health monitoring. All events
use femtologging with consistent field schemas suitable for parsing by log
aggregators (Datadog, Loki, CloudWatch Logs Insights, etc.). Log levels accept
TRACE, DEBUG, INFO, WARN/WARNING, ERROR, and CRITICAL.
Log events
The ingestion worker emits the following structured events:
| Event Type | Level | Description |
|---|---|---|
ingestion.run.started |
INFO | Ingestion run begins for a repository |
ingestion.run.completed |
INFO | Ingestion run finished successfully |
ingestion.run.failed |
ERROR | Ingestion run failed with error |
ingestion.stream.completed |
INFO | Stream (commit/PR/issue/doc) ingested |
ingestion.stream.truncated |
WARNING | Stream hit max_events limit, backlog exists |
Each event includes repo_slug and estate_id for filtering. Completion
events include duration and event counts. Failure events categorize errors for
alert routing.
Example log output:
ghillie.github.observability [INFO] [ingestion.run.completed]
repo_slug=octo/reef estate_id=wildside duration_seconds=45.200
commits_ingested=12 pull_requests_ingested=3 issues_ingested=5
doc_changes_ingested=2 total_events=22
Error categories
Failed ingestion runs are classified into categories for alerting:
| Category | Description | Typical Action |
|---|---|---|
transient |
GitHub 5xx errors | Retry automatically |
client_error |
GitHub 4xx errors | Check token/permissions |
schema_drift |
Unexpected API response shape | Investigate API changes |
configuration |
Missing or invalid config | Fix environment/config |
database_connectivity |
DB connection issues | Check infrastructure |
data_integrity |
Constraint violations | Investigate data issues |
database_error |
Other DB errors | Check database health |
unknown |
Unclassified errors | Investigate logs |
Querying ingestion lag
Use IngestionHealthService to query per-repository lag and identify stalled
ingestion:
import asyncio
import datetime as dt
from sqlalchemy.ext.asyncio import async_sessionmaker, create_async_engine
from ghillie.github import (
IngestionHealthConfig,
IngestionHealthService,
)
async def main() -> None:
engine = create_async_engine("sqlite+aiosqlite:///ghillie.db")
session_factory = async_sessionmaker(engine, expire_on_commit=False)
# Default threshold is 1 hour
service = IngestionHealthService(session_factory)
# Or configure a custom threshold
config = IngestionHealthConfig(stalled_threshold=dt.timedelta(minutes=30))
service = IngestionHealthService(session_factory, config=config)
# Query a single repository
metrics = await service.get_lag_for_repository("octo/reef")
if metrics:
print(f"Lag: {metrics.time_since_last_ingestion_seconds}s")
print(f"Has backlog: {metrics.has_pending_cursors}")
print(f"Stalled: {metrics.is_stalled}")
# Find all stalled repositories
stalled = await service.get_stalled_repositories()
for repo in stalled:
print(f"STALLED: {repo.repo_slug}")
asyncio.run(main())
Lag metrics
The IngestionLagMetrics dataclass provides:
repo_slug: Repository identifier (owner/name)time_since_last_ingestion_seconds: Seconds since newest watermark (None if never ingested)oldest_watermark_age_seconds: Age of oldest stream watermarkhas_pending_cursors: True if any stream has a pagination cursor (backlog)is_stalled: True if lag exceeds threshold or never ingested
Alerting recommendations
Configure alerts based on structured log queries:
-
Transient failures (repeated): Alert if
ingestion.run.failedwitherror_category=transientoccurs 3+ times in 15 minutes for the same repository. -
Configuration errors: Alert immediately on any
ingestion.run.failedwitherror_category=configuration. -
Schema drift: Alert on
error_category=schema_driftfor prompt investigation of GitHub API changes. -
Stalled ingestion: Query
IngestionHealthService.get_stalled_repositories()periodically and alert when the list is non-empty. -
Backlog accumulation: Alert when
ingestion.stream.truncatedevents occur repeatedly for the same repository (indicates sustained high activity or processing issues).