thread-master/apps/backend/internal/module/job/repository
王性驊 8c1be42f4a prod: throttle persona scrape jobs, cap worker CPU, add job-failure alerting + offsite backup
- job: cap concurrent persona_analyze_account (real headless-Chromium scrape)
  jobs per owner to 1 in flight via CountActiveByOwnerAndTemplate; rejects
  with the new ErrTooManyActive (429) instead of letting a member queue up
  many scrapes and pin the shared 1-vCPU host, or double-schedule the same
  persona while a previous run is still executing.
- harbor-worker@.service: CPUQuota=70% so nginx/gateway keep headroom while
  a scrape job briefly saturates Chromium.
- backup.sh: default BACKUP_DIR now matches the systemd unit (/var/backups/harbor)
  instead of the stale /var/backups/haixun default.
- add harbor-job-health.timer (every 5m): queries Mongo for jobs that failed
  in the last 15 minutes and writes a node-exporter textfile-collector
  metric (haixun_job_failures_recent), with new Prometheus alerts
  (PersonaScrapeJobsFailing / BackgroundJobsFailing / JobHealthMetricStale)
  so a stuck worker surfaces on its own instead of only via member reports —
  this would have caught the recent Playwright headless-shell bug directly.
- add harbor-offsite-backup.timer + backup/offsite-sync.sh: pushes
  backup.sh's local backups to an off-host restic repository (client-side
  encrypted). Installed and enabled by default but a safe no-op until
  RESTIC_REPOSITORY/RESTIC_PASSWORD are set in harbor.env.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-17 17:38:36 +00:00
..
memory.go prod: throttle persona scrape jobs, cap worker CPU, add job-failure alerting + offsite backup 2026-07-17 17:38:36 +00:00
mongo.go prod: throttle persona scrape jobs, cap worker CPU, add job-failure alerting + offsite backup 2026-07-17 17:38:36 +00:00