- job: cap concurrent persona_analyze_account (real headless-Chromium scrape) jobs per owner to 1 in flight via CountActiveByOwnerAndTemplate; rejects with the new ErrTooManyActive (429) instead of letting a member queue up many scrapes and pin the shared 1-vCPU host, or double-schedule the same persona while a previous run is still executing. - harbor-worker@.service: CPUQuota=70% so nginx/gateway keep headroom while a scrape job briefly saturates Chromium. - backup.sh: default BACKUP_DIR now matches the systemd unit (/var/backups/harbor) instead of the stale /var/backups/haixun default. - add harbor-job-health.timer (every 5m): queries Mongo for jobs that failed in the last 15 minutes and writes a node-exporter textfile-collector metric (haixun_job_failures_recent), with new Prometheus alerts (PersonaScrapeJobsFailing / BackgroundJobsFailing / JobHealthMetricStale) so a stuck worker surfaces on its own instead of only via member reports — this would have caught the recent Playwright headless-shell bug directly. - add harbor-offsite-backup.timer + backup/offsite-sync.sh: pushes backup.sh's local backups to an off-host restic repository (client-side encrypted). Installed and enabled by default but a safe no-op until RESTIC_REPOSITORY/RESTIC_PASSWORD are set in harbor.env. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|---|---|---|
| .. | ||
| alertmanager | ||
| alloy | ||
| blackbox | ||
| grafana/provisioning/datasources | ||
| loki | ||
| prometheus | ||
| job-health-check.sh | ||