mirror of
https://github.com/ARUP-CAS/aiscr-qgis-amcr-viewer.git
synced 2026-10-08 20:07:36 +02:00
Plánovaný workflow api_monitor.yml (denně 05:17 UTC + ruční spuštění) ověřuje kontrakt API, na kterém plugin závisí: - tests/api_contract.py – stejné dotazy jako plugin, kontrola klíčů, typů a tvarů odpovědí (OK / DRIFT / FAIL / UNAVAILABLE) - tests/api_plugin_live.py – vlastní funkce pluginu (fetch_set, load_amcr_data) proti živému API v qgis/qgis:ltr - tests/api_monitor_report.py – jedno sledovací issue se štítkem api-monitor; čistý běh ho zavře, výpadek ho nemění Neběží na PR, aby výpadek digiarchivu neshodil PR. Popis v AGENTS.md, OpenSpec změna add-daily-api-monitor. Připraveno s pomocí AI (Claude), ověřeno proti produkčnímu API.
10 KiB
10 KiB
Design
Context
- What broke in #67: the facet item shape of
api/search/query(json.nlarrntv→arrarrwith the Solr 10 migration in digiarchiv v4.1.0).fetch_setcaught theTypeError, logged a warning and returned[]; nothing outside QGIS noticed. The plugin now accepts both shapes (amcr_codelists._facet_name) and keeps previous codelist values when a set comes back empty, which makes the next break of this kind even quieter for users – only a check outside the plugin catches it. - What the plugin uses (from the code on
main):GET https://digiarchiv.aiscr.cz/api/assets/i18n/cs.json(amcr_tools.load_translations)GET …/api/search/querywithentity,mapa=true,sort=ident_cely asc,rows=500,page,loc_rpt=minLat,minLon,maxLat,maxLon, filters askey=value:orlists, date ranges; responseresponse.numFound/response.docs[], errors as HTTP 200 withoutresponse(amcr_tools.load_amcr_data,_api_get_json)- PIAN geometry requests in batches (
amcr_tools.load_amcr_data) …/api/search/querywithrows=0&noFacets=false&onlyFacets=true, responsefacet_counts.facet_fields.<f_…>[](amcr_codelists.fetch_setforvedouci,nalezce)GET https://api.aiscr.cz/2.2/oai?verb=ListRecords&metadataPrefix=oai_dc &set=…withresumptionTokenpagination (fetch_set, all other sets inamcr_codelists.slovnicek)POST …/api/user/login,GET …/api/user/islogged,…/logout(amcr_tools.login_to_api, session check)
- Live probe (2026-10-02, anonymous, Praha bbox
49.9,14.3,50.2,14.7, one page of 500):akcenumFound 20 635 (1.0 MB, 0.6 s),lokalita342 (0.8 MB, 0.4 s),pian30 321 (0.8 MB, 0.5 s),samostatny_nalez0 (anonymous SN are sparse – the test bbox must be chosen so that every entity returns records). Bbox chosen: Mikulov48.8,16.6,48.9,16.75– akce 185, lokalita 18, samostatny_nalez 2, pian 294. Only 12 anonymous map-enabled SN exist in the whole CZ, so 2 is accepted: archive records do not disappear (maintainer decision); facet request 54 fields, items are lists. The deployed version read from the web bundle:v4.0.3-237-g96deec70-dirty(footer text is hard-coded and not reliable). - GitHub runs
scheduleonly on the default branch (main) and disables scheduled workflows after 60 days without repository activity.
Goals / Non-Goals
Goals:
- Daily, credential-free check that tells apart three outcomes: the API changed in a way the plugin breaks on (fail), the API changed in a way the plugin survives (drift), the API was not reachable (unavailable).
- Messages specific enough to start a fix without re-probing (field, old shape, new shape, request).
- One tracking issue, no daily noise.
Non-Goals:
- Logged-in access, access levels B–D.
- Full data validation (counts, contents of records).
- Changing plugin code to make it more testable – if a function cannot be
called without UI, the live test provides a fake
iface/ canvas.
Decisions
- Two scripts, two jobs.
tests/api_contract.py(job API contract,ubuntu-latest+actions/setup-python,requestspinned in workflowenv) andtests/api_plugin_live.py(job Plugin against live API,docker run qgis/qgis:ltr, same invocation style as the smoke test). Alternative: one script in the QGIS container – rejected: a QGIS image problem would hide the contract result, and the contract check would pay the image pull every day. Onlyltrfor the live job: the API path does not differ between Qt5 and Qt6, which the PR smoke test already covers on both. - Contract = recorded expectations in the script, not a stored
baseline file. Each check states the expected keys/types/shapes
inline (e.g. facet item is
[str, int],numFoundisint,docs[].ident_celyisstr). A mismatch the plugin tolerates is reported as DRIFT (e.g. facet shape back to{"name":…}, an extra type the parser accepts), a mismatch it does not tolerate as FAIL. Updating an expectation is a reviewed code change. Alternative: snapshot JSON baseline refreshed automatically – rejected: a silent baseline refresh would accept the very change we want to see. - Statuses and exit codes. Each check yields
OK/DRIFT/FAIL/UNAVAILABLE+ detail. Both scripts writeresults-<job>.json(check name, status, detail, request URL without secrets) and a Markdown table to$GITHUB_STEP_SUMMARYwhen set; exit code 1 on anyFAILorDRIFT, 0 otherwise (an all-UNAVAILABLErun is green but visible in the summary). Locally the scripts print the same table. - Retries. Network errors, timeouts and HTTP 5xx: 3 attempts with
backoff (2 s, 8 s). After that the check is
UNAVAILABLE; checks that depend on it are skipped asUNAVAILABLE, notFAIL. HTTP 4xx or a 200 with anerrorbody is a real answer and is judged by the contract. A per-host circuit breaker follows: once a host fails its full retry cycle, later requests to it return "unreachable" without network I/O, so an all-unreachable run ends in ~10 s instead of tens of minutes. - Test inputs from the live API, not from
heslar.csv. Filter values are taken from facets of the same bbox in the same run; the test bbox is a fixed small area where every entity (akce,lokalita,samostatny_nalez,pian) returns a non-zero anonymous count below one page, chosen during implementation by a probe and documented in the script. Pagination is checked separately withrows=100on a larger area (pages must not overlap; downloaded ≥ numFound when it is small enough). - Live plugin test thresholds. For each set in
amcr_codelists.slovnicek:fetch_setmust return ≥ 1 item and at least 50 % of the row count of that category in the bundledcodelists/heslar.csv(a shrunken codelist is the #67 symptom). For each data type:load_amcr_dataon the test bbox (fakeiface, fake canvas in EPSG:5514) must add at least one layer with ≥ 1 feature with a valid geometry and the expected attribute fields. The plugin package is imported as a package (amcr_viewer.amcr_tools), so its relative imports work – a barespec_from_file_locationmakesload_amcr_dataswallow the import error into "0 records". - Deployed version. Fetch
https://digiarchiv.aiscr.cz/home, scan the referenced*.jsbundles forraw:"v…"(git-describe) and report it; not finding it is aDRIFTof its own check, never aFAILof the run. - Reporting job (
needsboth,if: always(), only whengithub.ref_name == github.event.repository.default_branch; the workflow has no other triggers thanscheduleandworkflow_dispatch;permissions: issues: writefor this job only,contents: readelsewhere). It downloads both result files (artifacts) and withgh:- any
FAIL/DRIFT→ find the open issue with labelapi-monitor; if none, create it (gh label create api-monitor --forcefirst); if it exists and the fingerprint (sorted names of FAIL/DRIFT checks – not UNAVAILABLE, which would make it flap – stored as an HTML comment in the issue body) differs, add a comment and update the fingerprint; identical fingerprint → do nothing; - every check
OK→ close the open issue with a comment linking the run; - no
FAIL/DRIFTbut someUNAVAILABLE→ leave the issue as it is (an outage proves neither break nor recovery); - a job that did not produce its result file (crashed script, image
pull failure) counts as one
FAILcheck named after the job. Issue text in Czech (repo convention for issues), unwrapped GFM: deployed version, table of non-OK checks, run link, how to reproduce locally. Alternative:actions/github-scriptor a marketplace action – rejected:ghis preinstalled and needs no third-party action pin.
- any
- Schedule
cron: "17 5 * * *"(07:17 CEST) – off the full hour, before the working day. Plusworkflow_dispatch. Concurrency group per ref,cancel-in-progress: false. - Not a PR check. The new workflow does not run on
pull_request: a PR must not go red because digiarchiv is down. Contributors run the scripts locally or dispatch the workflow on their branch.
Risks / Trade-offs
- False alarms from data changes (a record deleted in the test bbox, an entity count dropping to 0) → thresholds are "≥ 1" and "≥ 50 % of the bundled codelist", not exact counts; the bbox is chosen with margin.
- Scheduled workflow auto-disabled after 60 days of inactivity →
documented in
AGENTS.md; any push tomain(dependabot included) resets the timer. - Tests only
main's plugin code – a fix waiting on a version branch is not exercised by the schedule → manual dispatch on that branch. - Anonymous only →
pristupnostB–D paths and login success are not covered; the login error path is. - Load on digiarchiv – a few dozen small requests a day; negligible.
Verification
- Both scripts run locally against the live API and pass
(
uv run -q --no-project --with requests python tests/api_contract.py; live test indocker run qgis/qgis:ltr). - #67 regression check: run the live test against the plugin as of the
commit before the #67 fix (
git archive) – it must FAIL onvedouci/nalezce; the contract test must flag a facet-shape DRIFT when its expectation is temporarily set to the old{"name":…}shape. - Outage simulation: point the scripts at an unroutable host
(environment override of the base URLs) → all checks
UNAVAILABLE, exit 0. - Reporting logic tested with a dry-run mode (
API_MONITOR_DRY_RUN=1prints theghcommands instead of running them) for: new issue, same fingerprint, changed fingerprint, recovery. - The full
AGENTS.mdcheck set passes on the new files (check_sources, bandit, detect-secrets--all-files, flake8--isolatedonamcr_viewer/, ruff on the repo);actionlinton the new workflow. - After merge: one manual
workflow_dispatchonmainand inspection of the summary.