Jump to content

Server Admin Log

From Wikitech

2026-09-11

  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-10

  • 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
  • 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
  • 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
  • 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
  • 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
  • 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
  • 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
  • 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
  • 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
  • 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
  • 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
  • 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
  • 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
  • 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
  • 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
  • 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
  • 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
  • 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
  • 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
  • 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
  • 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
  • 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
  • 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
  • 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for CodeMirror: turn on the new 2017 editor integration (T432558), [CodeMirror] enable for new and logged-out users by default on enwiki (T288161) (duration: 10m 59s)
  • 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
  • 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for CodeMirror: turn on the new 2017 editor integration (T432558), [CodeMirror] enable for new and logged-out users by default on enwiki (T288161) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for CodeMirror: turn on the new 2017 editor integration (T432558), [CodeMirror] enable for new and logged-out users by default on enwiki (T288161)
  • 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
  • 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
  • 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
  • 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
  • 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
  • 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for Enable discord preview extension code on testwiki (tk 2) (T437344 T431352) (duration: 13m 23s)
  • 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
  • 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
  • 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
  • 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
  • 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
  • 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
  • 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
  • 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
  • 21:37 derenrich@deploy1003: derenrich: Backport for Enable discord preview extension code on testwiki (tk 2) (T437344 T431352) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
  • 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
  • 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
  • 21:33 derenrich@deploy1003: Started scap sync-world: Backport for Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)
  • 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
  • 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
  • 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
  • 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for Enable ReadingLists on mediawiki and wikitech (T437113) (duration: 09m 45s)
  • 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
  • 21:26 jdlrobson@deploy1003: jdlrobson: Backport for Enable ReadingLists on mediawiki and wikitech (T437113) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for Enable ReadingLists on mediawiki and wikitech (T437113)
  • 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
  • 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — T437516
  • 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for Instrument donor ID dialog: hooks defined (T435565), Instrument the donor account dialog (T435565), DonorIdentification: Adjust experiment behavior (T435534), Allow donor consent workflow to work on non-Minerva skins (T436572) (duration: 13m 54s)
  • 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
  • 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
  • 21:07 jdlrobson@deploy1003: jdlrobson: Backport for Instrument donor ID dialog: hooks defined (T435565), Instrument the donor account dialog (T435565), DonorIdentification: Adjust experiment behavior (T435534), Allow donor consent workflow to work on non-Minerva skins (T436572) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
  • 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for Instrument donor ID dialog: hooks defined (T435565), Instrument the donor account dialog (T435565), DonorIdentification: Adjust experiment behavior (T435534), Allow donor consent workflow to work on non-Minerva skins (T436572)
  • 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
  • 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for Remove mode from RestModuleOverrides (T434267) (duration: 23m 49s)
  • 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
  • 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
  • 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
  • 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
  • 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
  • 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for Remove mode from RestModuleOverrides (T434267) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
  • 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for Remove mode from RestModuleOverrides (T434267)
  • 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
  • 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
  • 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
  • 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
  • 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
  • 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
  • 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for Bumping portals to master (T128546) (duration: 11m 24s)
  • 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
  • 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
  • 20:14 jdrewniak@deploy1003: jdrewniak: Backport for Bumping portals to master (T128546) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
  • 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for Bumping portals to master (T128546)
  • 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
  • 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for Bumping portals to master (T128546)
  • 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
  • 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — T437516
  • 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
  • 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
  • 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
  • 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
  • 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
  • 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
  • 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
  • 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
  • 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
  • 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
  • 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
  • 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
  • 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
  • 19:17 thcipriani: Gerrit downtime incoming for upgrade
  • 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
  • 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
  • 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
  • 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
  • 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
  • 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
  • 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
  • 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
  • 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
  • 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
  • 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
  • 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
  • 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
  • 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
  • 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
  • 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
  • 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs T430838
  • 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
  • 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
  • 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
  • 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
  • 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
  • 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588), refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588)
  • 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
  • 16:33 urbanecm@deploy1003: urbanecm: Backport for refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588), refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588) synced to the te
  • 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588), refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588), Provider: Cache the valid configuration in the process (T437588)
  • 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
  • 15:04 samtar@deploy1003: Finished scap sync-world: Backport for IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955) (duration: 08m 08s)
  • 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
  • 14:59 samtar@deploy1003: samtar: Continuing with deployment
  • 14:58 samtar@deploy1003: samtar: Backport for IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 14:56 samtar@deploy1003: Started scap sync-world: Backport for IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)
  • 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
  • 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
  • 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
  • 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
  • 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
  • 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
  • 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
  • 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
  • 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
  • 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 12:48 klausman@dns1004: END - running authdns-update
  • 12:46 klausman@dns1004: START - running authdns-update
  • 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
  • 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
  • 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
  • 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
  • 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
  • 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
  • 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
  • 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
  • 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
  • 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
  • 11:52 cgoubert@dns1004: END - running authdns-update
  • 11:49 cgoubert@dns1004: START - running authdns-update
  • 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
  • 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
  • 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
  • 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
  • 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
  • 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
  • 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
  • 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
  • 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
  • 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
  • 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
  • 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
  • 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
  • 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
  • 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
  • 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
  • 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
  • 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
  • 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
  • 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
  • 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
  • 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
  • 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
  • 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
  • 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
  • 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
  • 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299) (duration: 09m 56s)
  • 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
  • 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
  • 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
  • 08:29 mszwarc@deploy1003: mszwarc: Backport for UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
  • 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
  • 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
  • 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
  • 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)
  • 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
  • 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
  • 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
  • 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
  • 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
  • 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
  • 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
  • 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
  • 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
  • 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
  • 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
  • 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
  • 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
  • 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
  • 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
  • 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
  • 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
  • 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
  • 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
  • 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
  • 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
  • 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
  • 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
  • 00:11 dzahn@dns1004: END - running authdns-update
  • 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
  • 00:08 dzahn@dns1004: START - running authdns-update

2026-09-09

  • 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
  • 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
  • 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
  • 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
  • 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
  • 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
  • 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
  • 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
  • 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
  • 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
  • 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
  • 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
  • 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
  • 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
  • 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
  • 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for Move testcommonswiki links tables to x4 (T398709) (duration: 11m 15s)
  • 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
  • 23:16 ladsgroup@deploy1003: ladsgroup: Backport for Move testcommonswiki links tables to x4 (T398709) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
  • 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for Move testcommonswiki links tables to x4 (T398709)
  • 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
  • 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
  • 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
  • 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
  • 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
  • 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
  • 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
  • 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
  • 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
  • 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
  • 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for Revert "Enable discord preview extension code on testwiki" (duration: 10m 21s)
  • 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
  • 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
  • 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
  • 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
  • 22:26 derenrich@deploy1003: derenrich, egardner: Backport for Revert "Enable discord preview extension code on testwiki" synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
  • 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
  • 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
  • 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
  • 22:22 derenrich@deploy1003: Started scap sync-world: Backport for Revert "Enable discord preview extension code on testwiki"
  • 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
  • 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for Enable discord preview extension code on testwiki (T437344) (duration: 13m 40s)
  • 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
  • 22:10 derenrich@deploy1003: derenrich: Backport for Enable discord preview extension code on testwiki (T437344) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 22:05 derenrich@deploy1003: Started scap sync-world: Backport for Enable discord preview extension code on testwiki (T437344)
  • 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
  • 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
  • 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
  • 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
  • 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
  • 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
  • 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
  • 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
  • 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
  • 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
  • 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
  • 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
  • 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
  • 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for wikifunctions: Move client fragments to mainstash (T432849) (duration: 12m 29s)
  • 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
  • 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
  • 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
  • 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
  • 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
  • 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
  • 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
  • 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
  • 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
  • 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 21:14 jforrester@deploy1003: jforrester: Backport for wikifunctions: Move client fragments to mainstash (T432849) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 21:09 jforrester@deploy1003: Started scap sync-world: Backport for wikifunctions: Move client fragments to mainstash (T432849)
  • 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
  • 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
  • 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
  • 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
  • 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
  • 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
  • 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
  • 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
  • 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
  • 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
  • 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
  • 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
  • 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
  • 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
  • {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", [[gerrit:1338279|Revert^2}}
  • 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
  • 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
  • 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
  • {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", [[gerrit:1338279|Revert^2 "Filter}}
  • 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
  • {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", Revert^2 "Filter out non-http(s) license urls", [[gerrit:1338279|Revert^2 "}}
  • 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
  • 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
  • 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
  • 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
  • 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
  • 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
  • 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
  • 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
  • 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
  • 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
  • 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
  • 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
  • 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
  • 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
  • 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
  • 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
  • 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
  • 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
  • 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
  • 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
  • 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs T430838
  • 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
  • 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
  • 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
  • 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
  • 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
  • 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
  • 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
  • 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
  • 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
  • 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
  • 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
  • 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
  • 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
  • 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
  • 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
  • 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
  • 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
  • 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
  • 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
  • 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
  • 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
  • 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
  • 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
  • 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
  • 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
  • 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
  • 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
  • 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
  • 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
  • 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
  • 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
  • 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
  • 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
  • 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
  • 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
  • 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
  • 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for testwiki: allow testing new AccountSetup experiment (T436872) (duration: 09m 28s)
  • 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
  • 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for testwiki: allow testing new AccountSetup experiment (T436872) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for testwiki: allow testing new AccountSetup experiment (T436872)
  • 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
  • 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
  • 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
  • 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
  • 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
  • 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
  • 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
  • 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
  • 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for Disable parsoidCachePrewarm jobs everywhere (T436001 T436206) (duration: 09m 43s)
  • 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
  • 15:40 ladsgroup@deploy1003: ladsgroup: Backport for Disable parsoidCachePrewarm jobs everywhere (T436001 T436206) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)
  • 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; T392944)
  • 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
  • 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
  • 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
  • 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
  • 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
  • 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
  • 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
  • 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
  • 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
  • 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
  • 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
  • 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
  • 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
  • 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
  • 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
  • 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
  • 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
  • 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry T416452
  • 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
  • 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
  • 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
  • 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
  • 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
  • 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
  • 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for ArticleGuidance: Add the redirect configuration keys (T434487), Replace experiment with instrument and config-driven redirect (T434487) (duration: 10m 15s)
  • 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
  • 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
  • 13:09 sbisson@deploy1003: sbisson: Backport for ArticleGuidance: Add the redirect configuration keys (T434487), Replace experiment with instrument and config-driven redirect (T434487) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:04 sbisson@deploy1003: Started scap sync-world: Backport for ArticleGuidance: Add the redirect configuration keys (T434487), Replace experiment with instrument and config-driven redirect (T434487)
  • 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
  • 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
  • 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
  • 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
  • 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409) (duration: 14m 39s)
  • 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
  • 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)
  • 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [Growth] Enable iterative Add Link task pool population on all wikis (T392944) (duration: 21m 58s)
  • 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
  • 11:35 urbanecm@deploy1003: urbanecm: Backport for [Growth] Enable iterative Add Link task pool population on all wikis (T392944) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [Growth] Enable iterative Add Link task pool population on all wikis (T392944)
  • 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
  • 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry T416452
  • 10:43 moritzm: installing Bird security updates
  • 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
  • 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
  • 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry T416452
  • 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
  • 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
  • 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
  • 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
  • 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
  • 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
  • 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
  • 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
  • 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
  • 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
  • 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
  • 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
  • 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
  • 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
  • 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
  • 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
  • 08:30 brouberol@dns1004: END - running authdns-update
  • 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry T416452
  • 08:28 brouberol@dns1004: START - running authdns-update
  • 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
  • 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
  • 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
  • 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
  • 07:40 chlod: UTC morning backport window done
  • 07:37 chlod@deploy1003: Finished scap sync-world: Backport for thwikibooks: update tagline and wordmark (T436426) (duration: 21m 36s)
  • 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
  • 07:20 chlod@deploy1003: chlod, hamishz: Backport for thwikibooks: update tagline and wordmark (T436426) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 07:15 chlod@deploy1003: Started scap sync-world: Backport for thwikibooks: update tagline and wordmark (T436426)
  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
  • 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm

2026-09-08

  • 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
  • 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
  • 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
  • 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
  • 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
  • 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
  • 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:05 Amir1: dropped 57 tables on db1260 (T437278)
  • 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
  • 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
  • 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
  • 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
  • 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
  • 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
  • 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
  • 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
  • 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
  • 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
  • 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
  • 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
  • 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
  • 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
  • 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
  • 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
  • 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
  • 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
  • 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
  • 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for Bumping portals to master (T128546) (duration: 09m 53s)
  • 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
  • 21:58 Amir1: drop links tables from db2210 (T437278)
  • 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
  • 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
  • 21:56 jdrewniak@deploy1003: jdrewniak: Backport for Bumping portals to master (T128546) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
  • 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for Bumping portals to master (T128546)
  • 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
  • 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
  • 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for Assets build - 2026-09-08 21:25:19+00:00 (duration: 05m 27s)
  • 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
  • 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for Assets build - 2026-09-08 21:25:19+00:00 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for Assets build - 2026-09-08 21:25:19+00:00
  • 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
  • 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
  • 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
  • 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
  • 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
  • 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
  • 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
  • 21:24 reedy@deploy1003: Finished scap sync-world: Backport for InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103) (duration: 09m 12s)
  • 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
  • 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
  • 21:19 reedy@deploy1003: reedy: Continuing with deployment
  • 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
  • 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
  • 21:19 reedy@deploy1003: reedy: Backport for InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
  • 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
  • 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
  • 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
  • 21:15 reedy@deploy1003: Started scap sync-world: Backport for InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)
  • {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", [[gerrit:1338033|Revert "Filter out}}
  • 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
  • 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
  • 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
  • 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
  • {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", [[gerrit:1338033|Revert "Filter out non-http(s) lice}}
  • 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
  • {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", Revert "Filter out non-http(s) license urls", [[gerrit:1338033|Revert "Filter out n}}
  • 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
  • 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
  • 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
  • 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
  • 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
  • 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
  • {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), [[gerrit:1337996|Filter out non-http(s) license}}
  • 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
  • {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), Filter out non-http(s) license urls (T435999), [[gerrit:1337996|Filter out non-}}
  • 20:15 aaron@deploy1003: Finished scap sync-world: Backport for Add wmf-analytics-commons external module to commonswiki (T434927) (duration: 10m 16s)
  • 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
  • 20:10 aaron@deploy1003: aaron: Continuing with deployment
  • 20:09 aaron@deploy1003: aaron: Backport for Add wmf-analytics-commons external module to commonswiki (T434927) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
  • 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
  • 20:05 aaron@deploy1003: Started scap sync-world: Backport for Add wmf-analytics-commons external module to commonswiki (T434927)
  • 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
  • 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
  • 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
  • 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
  • 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
  • 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
  • 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
  • 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
  • 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
  • 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
  • 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
  • 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
  • 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
  • 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
  • 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
  • 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
  • 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
  • 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
  • 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
  • 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
  • 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
  • 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
  • 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
  • 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
  • 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
  • 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
  • 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
  • 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
  • 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
  • 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
  • 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
  • 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
  • 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
  • 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
  • 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs T430838
  • 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
  • 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
  • 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
  • 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
  • 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
  • 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
  • 17:11 Amir1: dropping links tables from db1247 (s4 replica) - (T437278)
  • 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
  • 16:51 zabe@deploy1003: Finished scap sync-world: Backport for Use local database for category table in SpecialWantedCategories (T437286), Use local database for category table in SpecialWantedCategories (T437286) (duration: 10m 19s)
  • 16:46 zabe@deploy1003: zabe: Continuing with deployment
  • 16:45 zabe@deploy1003: zabe: Backport for Use local database for category table in SpecialWantedCategories (T437286), Use local database for category table in SpecialWantedCategories (T437286) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 16:41 zabe@deploy1003: Started scap sync-world: Backport for Use local database for category table in SpecialWantedCategories (T437286), Use local database for category table in SpecialWantedCategories (T437286)
  • 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
  • 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
  • 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
  • 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
  • 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
  • 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
  • 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
  • 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
  • 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
  • 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
  • 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
  • 14:55 logmsgbot: dreamyjazz Deployed security patch for T436429
  • 14:46 logmsgbot: dreamyjazz Deployed security patch for T436429
  • 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
  • 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
  • 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
  • 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
  • 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
  • 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
  • 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
  • 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
  • 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
  • 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
  • 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan 1788220800 --verbose # T437158
  • 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
  • 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
  • 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
  • 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade T426197
  • 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
  • 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
  • 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
  • 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
  • 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
  • 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
  • 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
  • 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # T435066
  • 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
  • 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
  • 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
  • 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
  • 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
  • 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
  • 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
  • 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
  • 13:44 stran@deploy1003: Finished scap sync-world: Backport for Split out edit and block-based filters from activity filters (T436508) (duration: 34m 00s)
  • 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
  • 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
  • 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
  • 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
  • 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
  • 13:32 stran@deploy1003: stran: Continuing with deployment
  • 13:29 stran@deploy1003: stran: Backport for Split out edit and block-based filters from activity filters (T436508) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
  • 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
  • 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
  • 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
  • 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
  • 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
  • 13:21 moritzm: installing qemu security updates
  • 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
  • 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
  • 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
  • 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
  • 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
  • 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
  • 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
  • 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
  • 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
  • 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:10 stran@deploy1003: Started scap sync-world: Backport for Split out edit and block-based filters from activity filters (T436508)
  • 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
  • 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
  • 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
  • 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
  • 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
  • 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
  • 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
  • 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
  • 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
  • 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
  • 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
  • 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
  • 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
  • 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
  • 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
  • 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
  • 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
  • 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
  • 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
  • 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
  • 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
  • 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
  • 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
  • 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
  • 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
  • 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
  • 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
  • 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
  • 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
  • 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
  • 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
  • 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
  • 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
  • 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
  • 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
  • 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
  • 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
  • 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
  • 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
  • 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
  • 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
  • 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
  • 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
  • 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
  • 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
  • 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
  • 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
  • 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
  • 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
  • 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
  • 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
  • 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
  • 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
  • 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
  • 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
  • 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
  • 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
  • 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
  • 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
  • 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
  • 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
  • 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
  • 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
  • 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
  • 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
  • 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
  • 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
  • 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
  • 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
  • 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
  • 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
  • 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
  • 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
  • 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
  • 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
  • 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
  • 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
  • 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
  • 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
  • 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 12:06 zabe@deploy1003: Finished scap sync-world: Backport for Skip DB read-only check for CentralAuthSessionManager (T437273), Skip DB read-only check for CentralAuthSessionManager (T437273) (duration: 09m 54s)
  • 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
  • 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
  • 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for Skip DB read-only check for CentralAuthSessionManager (T437273), Skip DB read-only check for CentralAuthSessionManager (T437273) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
  • 11:56 zabe@deploy1003: Started scap sync-world: Backport for Skip DB read-only check for CentralAuthSessionManager (T437273), Skip DB read-only check for CentralAuthSessionManager (T437273)
  • 11:46 marostegui@dns1004: END - running authdns-update
  • 11:44 marostegui@dns1004: START - running authdns-update
  • 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
  • 11:40 Amir1: dropping unneeded tables from x4 - db1260 (T437278)
  • 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
  • 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
  • 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
  • 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
  • 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
  • 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
  • 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
  • 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
  • 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
  • 10:43 samtar@deploy1003: Finished scap sync-world: Backport for IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355) (duration: 10m 57s)
  • 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
  • 10:38 samtar@deploy1003: samtar: Continuing with deployment
  • 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
  • 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
  • 10:36 samtar@deploy1003: samtar: Backport for IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
  • 10:32 samtar@deploy1003: Started scap sync-world: Backport for IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)
  • 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
  • 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
  • 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
  • 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
  • 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
  • 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
  • 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
  • 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
  • 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
  • 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
  • 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
  • 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
  • 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
  • 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
  • 09:45 zabe@deploy1003: Finished scap sync-world: Backport for Use virtual domain in NameTableStore for collation (T405812), Use virtual domain in NameTableStore for collation (T405812) (duration: 13m 15s)
  • 09:45 ayounsi@dns1004: END - running authdns-update
  • 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
  • 09:43 ayounsi@dns1004: START - running authdns-update
  • 09:39 zabe@deploy1003: zabe: Continuing with deployment
  • 09:37 zabe@deploy1003: zabe: Backport for Use virtual domain in NameTableStore for collation (T405812), Use virtual domain in NameTableStore for collation (T405812) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
  • 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
  • 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
  • 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 09:32 zabe@deploy1003: Started scap sync-world: Backport for Use virtual domain in NameTableStore for collation (T405812), Use virtual domain in NameTableStore for collation (T405812)
  • 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
  • 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
  • 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
  • 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
  • 09:23 moritzm: installing rsync security updates
  • 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
  • 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
  • 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again T404715', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
  • 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 T404715', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
  • 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters T404715', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
  • 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
  • 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance T404715', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
  • 09:02 marostegui: Starting x4 split from s4, RO time on commons needed T404715
  • 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
  • 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
  • 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
  • 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
  • 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
  • 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
  • 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
  • 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
  • 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
  • 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
  • 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
  • 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
  • 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
  • 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
  • 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
  • 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
  • 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
  • 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
  • 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
  • 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - T436056
  • 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
  • 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
  • 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
  • 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
  • 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
  • 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
  • 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
  • 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
  • 07:14 jmm@dns1004: END - running authdns-update
  • 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
  • 07:12 jmm@dns1004: START - running authdns-update
  • 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
  • 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
  • 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - T436056
  • 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
  • 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs T430838 (duration: 36m 30s)
  • 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs T430838
  • 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
  • 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-07

  • 21:52 zabe@deploy1003: Finished scap sync-world: Backport for Move globalimagelinks queries to x4 (T437108) (duration: 11m 00s)
  • 21:47 zabe@deploy1003: zabe: Continuing with deployment
  • 21:45 zabe@deploy1003: zabe: Backport for Move globalimagelinks queries to x4 (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:41 zabe@deploy1003: Started scap sync-world: Backport for Move globalimagelinks queries to x4 (T437108)
  • 21:37 zabe@deploy1003: Finished scap sync-world: Backport for Move production reads for commons link tables to x4 for all requests (T437108) (duration: 09m 34s)
  • 21:33 zabe@deploy1003: zabe: Continuing with deployment
  • 21:32 zabe@deploy1003: zabe: Backport for Move production reads for commons link tables to x4 for all requests (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:28 zabe@deploy1003: Started scap sync-world: Backport for Move production reads for commons link tables to x4 for all requests (T437108)
  • 21:03 zabe@deploy1003: Finished scap sync-world: Backport for Move production reads for commons link tables to x4 for 50% of requests (T437108) (duration: 10m 27s)
  • 20:58 zabe@deploy1003: zabe: Continuing with deployment
  • 20:57 zabe@deploy1003: zabe: Backport for Move production reads for commons link tables to x4 for 50% of requests (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:52 zabe@deploy1003: Started scap sync-world: Backport for Move production reads for commons link tables to x4 for 50% of requests (T437108)
  • 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 (T437108)', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
  • 20:23 zabe@deploy1003: Finished scap sync-world: Backport for Move production reads for commons link tables to x4 for 25% of requests (T437108) (duration: 09m 28s)
  • 20:19 zabe@deploy1003: zabe: Continuing with deployment
  • 20:18 zabe@deploy1003: zabe: Backport for Move production reads for commons link tables to x4 for 25% of requests (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:14 zabe@deploy1003: Started scap sync-world: Backport for Move production reads for commons link tables to x4 for 25% of requests (T437108)
  • 20:12 zabe@deploy1003: Finished scap sync-world: Backport for Stop setting wgCampaignEventsEnableWorklists (T429510) (duration: 10m 06s)
  • 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
  • 20:06 zabe@deploy1003: zabe, daimona: Backport for Stop setting wgCampaignEventsEnableWorklists (T429510) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:02 zabe@deploy1003: Started scap sync-world: Backport for Stop setting wgCampaignEventsEnableWorklists (T429510)
  • 19:59 zabe@deploy1003: Finished scap sync-world: Backport for Move production reads for commons link tables to x4 for 10% of requests (T437108) (duration: 11m 27s)
  • 19:55 zabe@deploy1003: zabe: Continuing with deployment
  • 19:52 zabe@deploy1003: zabe: Backport for Move production reads for commons link tables to x4 for 10% of requests (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 19:48 zabe@deploy1003: Started scap sync-world: Backport for Move production reads for commons link tables to x4 for 10% of requests (T437108)
  • 19:31 zabe@deploy1003: Finished scap sync-world: Backport for Move production reads for commons link tables to x4 for 1% of requests (T437108) (duration: 11m 40s)
  • 19:26 zabe@deploy1003: zabe: Continuing with deployment
  • 19:23 zabe@deploy1003: zabe: Backport for Move production reads for commons link tables to x4 for 1% of requests (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 19:19 zabe@deploy1003: Started scap sync-world: Backport for Move production reads for commons link tables to x4 for 1% of requests (T437108)
  • 19:07 zabe@deploy1003: Finished scap sync-world: Backport for Set RemoteVirtualDomainsMapping for test-commons (duration: 14m 17s)
  • 19:00 zabe@deploy1003: zabe: Continuing with deployment
  • 18:57 zabe@deploy1003: zabe: Backport for Set RemoteVirtualDomainsMapping for test-commons synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 18:53 zabe@deploy1003: Started scap sync-world: Backport for Set RemoteVirtualDomainsMapping for test-commons
  • 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
  • 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
  • 18:31 zabe@deploy1003: Finished scap sync-world: Backport for SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824) (duration: 09m 12s)
  • 18:27 zabe@deploy1003: zabe: Continuing with deployment
  • 18:26 zabe@deploy1003: zabe: Backport for SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 18:22 zabe@deploy1003: Started scap sync-world: Backport for SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)
  • 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
  • 16:01 zabe@deploy1003: Finished scap sync-world: Backport for Disable query pages not yet compatible with commons split (T437108) (duration: 10m 22s)
  • 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
  • 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
  • 15:56 zabe@deploy1003: zabe: Continuing with deployment
  • 15:55 zabe@deploy1003: zabe: Backport for Disable query pages not yet compatible with commons split (T437108) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 15:51 zabe@deploy1003: Started scap sync-world: Backport for Disable query pages not yet compatible with commons split (T437108)
  • 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
  • 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
  • 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
  • 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
  • 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
  • 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
  • 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
  • 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
  • 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
  • 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
  • 15:11 moritzm: installing rsync security updates
  • 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
  • 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
  • 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
  • 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
  • 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
  • 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
  • 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
  • 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
  • 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
  • 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
  • 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
  • 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
  • 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
  • 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
  • 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
  • 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
  • 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
  • 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
  • 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
  • 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
  • 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
  • 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
  • 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
  • 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
  • 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
  • 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
  • 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971), User: Use ActorStore on User::load (T428517), Fix empty bases in m(under|over) (T436876), Use Core-compatible output for cancellation (T379359), [[gerrit:1336609|Page: Fix absence caching when combined with mul}}
  • 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
  • 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
  • 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
  • 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
  • {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971), User: Use ActorStore on User::load (T428517), Fix empty bases in m(under|over) (T436876), Use Core-compatible output for cancellation (T379359), [[gerrit:1336609|Page: Fix absence caching when combined with multiple properties}}
  • 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
  • 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
  • 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
  • 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
  • 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
  • 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
  • 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
  • {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971), User: Use ActorStore on User::load (T428517), Fix empty bases in m(under|over) (T436876), Use Core-compatible output for cancellation (T379359), [[gerrit:1336609|Page: Fix absence caching when combined with mult}}
  • 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
  • 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
  • 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
  • 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
  • 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
  • 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
  • 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
  • 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
  • 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
  • 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
  • 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
  • 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
  • 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
  • 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
  • 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
  • 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
  • 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
  • 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
  • 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
  • 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
  • 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
  • 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
  • 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:25 moritzm: installing openssh security updates
  • 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
  • 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
  • 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
  • 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
  • 13:22 stran@deploy1003: Finished scap sync-world: Backport for SI: Use new InfoChip text class in case status updater (T437020) (duration: 10m 06s)
  • 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
  • 13:17 stran@deploy1003: stran: Continuing with deployment
  • 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 13:16 stran@deploy1003: stran: Backport for SI: Use new InfoChip text class in case status updater (T437020) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
  • 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
  • 13:12 stran@deploy1003: Started scap sync-world: Backport for SI: Use new InfoChip text class in case status updater (T437020)
  • 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
  • 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
  • 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
  • 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
  • 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
  • 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
  • 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
  • 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
  • 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
  • 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
  • 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
  • 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
  • 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
  • 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
  • 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
  • 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
  • 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
  • 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
  • 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in T435499 or contact the oncall SREs
  • 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
  • 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
  • 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
  • 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
  • 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
  • 11:54 jmm@dns1004: END - running authdns-update
  • 11:52 jmm@dns1004: START - running authdns-update
  • 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
  • 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
  • 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
  • 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
  • 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
  • 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
  • 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
  • 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
  • 11:41 moritzm: installing bash updates from bookworm point release
  • 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
  • 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
  • 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
  • 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
  • 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
  • 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
  • 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
  • 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
  • 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
  • 11:18 zabe@deploy1003: Finished scap sync-world: Backport for rdbms: Discourage setting db domain to false for remote virtual domains (T422940), Migrate querying categorylinks to virtual domain (T405812) (duration: 14m 08s)
  • 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
  • 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
  • 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • 11:10 zabe@deploy1003: zabe: Continuing with deployment
  • 11:10 zabe@deploy1003: zabe: Backport for rdbms: Discourage setting db domain to false for remote virtual domains (T422940), Migrate querying categorylinks to virtual domain (T405812) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
  • 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
  • 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
  • 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
  • 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • 11:04 zabe@deploy1003: Started scap sync-world: Backport for rdbms: Discourage setting db domain to false for remote virtual domains (T422940), Migrate querying categorylinks to virtual domain (T405812)
  • 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
  • 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for T436913 (duration: 35m 20s)
  • 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
  • 10:53 jmm@dns1004: END - running authdns-update
  • 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
  • 10:50 jmm@dns1004: START - running authdns-update
  • 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # T436659
  • 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # T436659
  • 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
  • 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
  • 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
  • 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
  • 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
  • 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
  • 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
  • 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
  • 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
  • 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
  • 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
  • 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
  • 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
  • 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
  • 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
  • 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
  • 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
  • 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
  • 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
  • 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
  • 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
  • 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for T436913
  • 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
  • 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
  • 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
  • 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
  • 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
  • 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
  • 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
  • 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
  • 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
  • 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
  • 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
  • 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
  • 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
  • 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
  • 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
  • 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
  • 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for Enable thumb.wikimedia.org everywhere (T427465) (duration: 10m 11s)
  • 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
  • 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
  • 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
  • 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
  • 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
  • 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
  • 09:46 ladsgroup@deploy1003: ladsgroup: Backport for Enable thumb.wikimedia.org everywhere (T427465) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
  • 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
  • 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for Enable thumb.wikimedia.org everywhere (T427465)
  • 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
  • 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
  • 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
  • 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
  • 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for Mentorship: Don't renew mentors' away status on every cleaner run (T436659), Ignore PersonalDashboard stubs (T428679) (duration: 20m 12s)
  • 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
  • 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
  • 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
  • 09:10 moritzm: rebuild software RAID following disk replacement T437036
  • 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
  • 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
  • 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
  • 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
  • 09:02 moritzm: installing giflib security updates
  • 09:01 urbanecm@deploy1003: urbanecm: Backport for Mentorship: Don't renew mentors' away status on every cleaner run (T436659), Ignore PersonalDashboard stubs (T428679) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
  • 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
  • 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
  • 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for Mentorship: Don't renew mentors' away status on every cleaner run (T436659), Ignore PersonalDashboard stubs (T428679)
  • 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
  • 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
  • 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
  • 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
  • 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
  • 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
  • 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
  • 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 T437188', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
  • 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
  • 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary T437188', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
  • 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
  • 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
  • 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - T437188
  • 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
  • 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
  • 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
  • 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
  • 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 T437188', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
  • 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 T437188
  • 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
  • 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
  • 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
  • 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
  • 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
  • 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
  • 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
  • 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
  • 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
  • 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
  • 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
  • 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
  • 07:47 kartik@deploy1003: Finished scap sync-world: Backport for ArticleGuidance: Configure feedback links to local talk pages (T433483) (duration: 41m 51s)
  • 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
  • 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
  • 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
  • 07:23 kartik@deploy1003: abi, kartik: Backport for ArticleGuidance: Configure feedback links to local talk pages (T433483) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 07:05 kartik@deploy1003: Started scap sync-world: Backport for ArticleGuidance: Configure feedback links to local talk pages (T433483)
  • 06:14 moritzm: installing Chromium security updates
  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-06

  • 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-05

  • 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-04

  • 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
  • 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
  • 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
  • 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
  • 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
  • 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
  • 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
  • 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
  • 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
  • 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
  • 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
  • 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
  • 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
  • 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
  • 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
  • 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
  • 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
  • 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
  • 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
  • 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
  • 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
  • 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
  • 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
  • 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
  • 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
  • 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
  • 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
  • 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
  • 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
  • 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
  • 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
  • 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
  • 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
  • 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
  • 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
  • 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
  • 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
  • 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
  • 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
  • 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
  • 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
  • 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
  • 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
  • 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
  • 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
  • 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
  • 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
  • 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
  • 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
  • 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
  • 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
  • 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
  • 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
  • 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
  • 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
  • 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
  • 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
  • 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
  • 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
  • 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
  • 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
  • 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
  • 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
  • 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
  • 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
  • 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
  • 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
  • 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
  • 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
  • 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
  • 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
  • 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
  • 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
  • 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
  • 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
  • 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
  • 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
  • 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
  • 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
  • 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
  • 10:00 marostegui@cumin1003: Removing db1182 from zarcillo T434869
  • 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
  • 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
  • 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
  • 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
  • 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
  • 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
  • 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
  • 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
  • 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
  • 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
  • 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
  • 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl T434869', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
  • 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
  • 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
  • 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
  • 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
  • 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
  • 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
  • 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
  • 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
  • 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
  • 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
  • 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
  • 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
  • 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
  • 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
  • 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
  • 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
  • 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
  • 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
  • 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
  • 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
  • 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
  • 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
  • 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for T436913 (duration: 34m 26s)
  • 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
  • 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
  • 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
  • 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
  • 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
  • 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
  • 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
  • 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
  • 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
  • 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
  • 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
  • 08:07 btullis@deploy1003: Started scap sync-world: Trying again for T436913
  • 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
  • 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
  • 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
  • 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
  • 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
  • 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
  • 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
  • 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
  • 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
  • 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
  • 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
  • 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
  • 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
  • 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
  • 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
  • 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
  • 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
  • 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
  • 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
  • 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
  • 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"

2026-09-03

  • 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # T435363
  • 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
  • 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for Add exclusions to Apple app site association file (T435363) (duration: 12m 24s)
  • 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
  • 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
  • 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
  • 20:46 arlolra@deploy1003: arlolra, tsev: Backport for Add exclusions to Apple app site association file (T435363) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
  • 20:42 arlolra@deploy1003: Started scap sync-world: Backport for Add exclusions to Apple app site association file (T435363)
  • 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
  • 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
  • 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919) (duration: 10m 23s)
  • 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
  • 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
  • 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
  • 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
  • 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
  • 20:28 arlolra@deploy1003: Started scap sync-world: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919)
  • 20:23 catrope@deploy1003: Finished scap sync-world: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135) (duration: 13m 41s)
  • 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
  • 20:16 catrope@deploy1003: catrope: Continuing with deployment
  • 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
  • 20:13 catrope@deploy1003: catrope: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
  • 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
  • 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
  • 20:09 catrope@deploy1003: Started scap sync-world: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135)
  • 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
  • 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
  • 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - T435419 (duration: 02m 59s)
  • 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - T435419
  • 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
  • 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
  • 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
  • 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs T430837
  • 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
  • 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
  • 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
  • 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
  • 17:39 ryankemper: T421642 [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
  • 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
  • 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
  • 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
  • 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
  • 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
  • 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
  • 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
  • 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
  • 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
  • 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
  • 17:01 btullis@deploy1003: Started scap sync-world: Trying again for T436913
  • 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
  • 16:58 dancy: Running scap clean-images on deploy1003
  • 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
  • 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 16:39 btullis@deploy1003: Started scap sync-world: Trying again for T436913
  • 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
  • 16:14 ryankemper: T430880 Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
  • 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
  • 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
  • 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
  • 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
  • 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for T436913
  • 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for Growth: Remove now removed config variable (duration: 09m 29s)
  • 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
  • 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for Growth: Remove now removed config variable
  • 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
  • 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
  • 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
  • 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
  • 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
  • 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
  • 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
  • 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
  • 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
  • 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
  • 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
  • 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
  • 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
  • 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
  • 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
  • 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
  • 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
  • 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
  • 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
  • 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
  • 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": T425441
  • 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
  • 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
  • 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" T425441
  • 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
  • 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
  • 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
  • 14:19 arnaudb@dns1006: END - running authdns-update
  • 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
  • 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
  • 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
  • 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
  • 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
  • 14:17 arnaudb@dns1006: START - running authdns-update
  • 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
  • 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
  • 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
  • 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
  • 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
  • 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
  • 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
  • 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
  • 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
  • 14:07 samtar@deploy1003: Finished scap sync-world: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087) (duration: 09m 36s)
  • 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
  • 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
  • 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
  • 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
  • 14:02 samtar@deploy1003: btullis, samtar: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
  • 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
  • 13:58 samtar@deploy1003: Started scap sync-world: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)
  • 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
  • 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
  • 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
  • 13:51 moritzm: installing sqlite3 security updates
  • 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
  • 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
  • 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
  • 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
  • 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
  • 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
  • 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
  • 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
  • 13:42 samtar@deploy1003: Finished scap sync-world: Backport for Remove revisionId from term fallback cache lines for properties (T434204) (duration: 13m 50s)
  • 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
  • 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
  • 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
  • 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
  • 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
  • 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
  • 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
  • 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
  • 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
  • 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
  • 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for Remove revisionId from term fallback cache lines for properties (T434204) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
  • 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
  • 13:28 samtar@deploy1003: Started scap sync-world: Backport for Remove revisionId from term fallback cache lines for properties (T434204)
  • 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
  • 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
  • 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
  • 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
  • 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:14 moritzm: installing bash updates from trixie point release
  • 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
  • 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
  • 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
  • 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
  • 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
  • 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
  • 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
  • 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
  • 13:08 jelto@dns1004: END - running authdns-update
  • 13:06 jelto@dns1004: START - running authdns-update
  • 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
  • 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
  • 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
  • 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
  • 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
  • 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
  • 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
  • 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
  • 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
  • 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
  • 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
  • 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
  • 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
  • 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
  • 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
  • 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
  • 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
  • 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
  • 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
  • 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
  • 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
  • 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
  • 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
  • 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
  • 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
  • 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
  • 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
  • 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
  • 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
  • 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
  • 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
  • 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
  • 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
  • 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
  • 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
  • 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
  • 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
  • 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
  • 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
  • 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
  • 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
  • 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
  • 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
  • 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
  • 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
  • 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
  • 11:58 kart_: cxserver: Use urldownloader LVS endpoint (T429175)
  • 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
  • 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
  • 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
  • 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
  • 11:55 moritzm: installing rsync security updates
  • 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
  • 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
  • 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
  • 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
  • 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
  • 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
  • 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
  • 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
  • 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
  • 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
  • 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
  • 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
  • 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
  • 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
  • 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
  • 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
  • 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
  • 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
  • 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
  • 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
  • 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
  • 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
  • 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
  • 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
  • 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
  • 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
  • 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
  • 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
  • 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
  • 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
  • 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
  • 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
  • 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
  • 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
  • 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 T434778
  • 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
  • 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
  • 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
  • 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
  • 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
  • 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
  • 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
  • 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
  • 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
  • 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
  • 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
  • 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
  • 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
  • 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
  • 09:39 hnowlan: fixed currently oncall pane in klaxon
  • 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
  • 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 T434751
  • 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
  • 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
  • 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 T434775
  • 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
  • 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
  • 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
  • 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
  • 09:30 zabe@deploy1003: Finished scap sync-world: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791) (duration: 09m 30s)
  • 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 09:25 zabe@deploy1003: zabe: Continuing with deployment
  • 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
  • 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
  • 09:25 zabe@deploy1003: zabe: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl T436904', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
  • 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
  • 09:21 zabe@deploy1003: Started scap sync-world: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)
  • 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
  • 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
  • 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable T435810
  • 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
  • 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
  • 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
  • 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 T434775
  • 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
  • 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
  • 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
  • 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
  • 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
  • 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
  • 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
  • 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
  • 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 T434776
  • 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
  • 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
  • 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server (duration: 01m 21s)
  • 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server
  • 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server (duration: 01m 28s)
  • 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server
  • 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
  • 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
  • 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
  • 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 T434287
  • 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
  • 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
  • 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
  • 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
  • 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
  • 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
  • 07:29 chlod: UTC morning backport window done
  • 07:27 chlod@deploy1003: Finished scap sync-world: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651) (duration: 11m 54s)
  • 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
  • 07:22 XioNoX: push pfw policies - T436729
  • 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 07:15 chlod@deploy1003: Started scap sync-world: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)
  • 07:15 marostegui: Power off db1228 for maintenance
  • 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
  • 07:01 arnaudb@dns1006: END - running authdns-update
  • 06:58 arnaudb@dns1006: START - running authdns-update
  • 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
  • 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
  • 06:46 moritzm: installing libxml2 security updates
  • 06:27 hashar: Upgrading CI Jenkins on contint1003 # T436812
  • 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
  • 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
  • 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
  • 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
  • 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
  • 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
  • 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
  • 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
  • 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm

2026-09-02

  • 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
  • 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for private/readme.php: Remove now removed secrets (T436880) (duration: 10m 21s)
  • 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
  • 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
  • 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for private/readme.php: Remove now removed secrets (T436880) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for private/readme.php: Remove now removed secrets (T436880)
  • 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
  • 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
  • 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
  • 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
  • 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
  • 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
  • 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for Campaigns should not override existing campaign query strings (T436681) (duration: 11m 03s)
  • 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
  • 22:30 jdlrobson@deploy1003: jdlrobson: Backport for Campaigns should not override existing campaign query strings (T436681) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for Campaigns should not override existing campaign query strings (T436681)
  • 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054) (duration: 14m 14s)
  • 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
  • 21:55 krinkle@deploy1003: krinkle: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:50 krinkle@deploy1003: Started scap sync-world: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054)
  • 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes T431608
  • 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637) (duration: 09m 48s)
  • 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
  • 21:34 jforrester@deploy1003: jforrester: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes T431608
  • 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637)
  • 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
  • 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
  • 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
  • 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
  • 20:18 dancy@deploy1003: Started scap sync-world: testing
  • 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
  • 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
  • 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851) (duration: 64m 27s)
  • 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade T431608
  • 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
  • 18:57 jforrester@deploy1003: jforrester: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
  • 18:53 jforrester@deploy1003: Started scap sync-world: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)
  • 18:31 sukhe@dns1004: END - running authdns-update
  • 18:28 sukhe@dns1004: START - running authdns-update
  • 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs T430837
  • 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
  • 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
  • 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
  • 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
  • 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
  • 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
  • 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
  • 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
  • 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
  • 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
  • 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
  • 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
  • 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
  • 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
  • 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
  • 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
  • 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
  • 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
  • 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
  • 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
  • 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
  • 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
  • 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
  • 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
  • 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
  • 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781 (duration: 48m 09s)
  • 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - T436781
  • 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
  • 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
  • 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
  • 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
  • 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
  • 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
  • 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
  • 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
  • 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
  • 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
  • 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
  • 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
  • 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
  • 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
  • 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
  • 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
  • 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
  • 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
  • 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
  • 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
  • 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
  • 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
  • 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
  • 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
  • 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
  • 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
  • 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
  • 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
  • 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
  • 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
  • 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
  • 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781
  • 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
  • 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
  • 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
  • 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
  • 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
  • 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
  • 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
  • 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
  • 14:32 moritzm: installing pdns-recursor security updates
  • 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
  • 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
  • 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
  • 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
  • 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
  • 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
  • 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
  • 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
  • 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
  • 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
  • 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
  • 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
  • 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
  • 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
  • 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 T427897
  • 13:36 samtar@deploy1003: Finished scap sync-world: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970) (duration: 09m 52s)
  • 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
  • 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:26 samtar@deploy1003: Started scap sync-world: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970)
  • 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
  • 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
  • 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
  • 13:24 moritzm: installing wireshark security updates
  • 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
  • 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
  • 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
  • 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
  • 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
  • 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia T436505
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
  • 12:44 atsuko@dns1004: END - running authdns-update
  • 12:41 atsuko@dns1004: START - running authdns-update
  • 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
  • 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
  • 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
  • 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517) (duration: 12m 50s)
  • 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
  • 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
  • 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
  • 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
  • 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517)
  • 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
  • 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
  • 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
  • 11:31 marostegui@cumin1003: Removing db1172 from zarcillo T436763
  • 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
  • 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
  • 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
  • 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
  • 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
  • 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
  • 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
  • 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
  • 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
  • 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
  • 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
  • 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
  • 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
  • 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
  • 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
  • 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
  • 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
  • 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
  • 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
  • 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
  • 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
  • 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
  • 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
  • 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
  • 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
  • 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
  • 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
  • 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
  • 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
  • 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl T436763', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
  • 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
  • 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
  • 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for T417800 (duration: 05m 37s)
  • 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for T417800
  • 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
  • 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
  • 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
  • 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
  • 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:56 jmm@dns1004: END - running authdns-update
  • 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:53 jmm@dns1004: START - running authdns-update
  • 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
  • 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl T435892', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
  • 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
  • 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
  • 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:30 moritzm: installing openjdk-21 security updates
  • 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
  • 09:10 moritzm: installing openjdk-8 security updates
  • 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
  • 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 08:46 tappof: bump space for prometheus k8s-dse in eqiad
  • 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
  • 08:36 moritzm: installing libgraphite2 security updates
  • 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 08:23 Msz2001: UTC morning backport window done
  • 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for Update stream config for user_info_card_interaction (T435585) (duration: 14m 36s)
  • 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
  • 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
  • 08:14 mszwarc@deploy1003: mszwarc: Backport for Update stream config for user_info_card_interaction (T435585) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
  • 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
  • 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
  • 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for Update stream config for user_info_card_interaction (T435585)
  • 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
  • 08:07 fabfur: depooling and silencing cp5022 (T414411)
  • 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
  • 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)}}
  • 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
  • 07:49 mszwarc@deploy1003: mszwarc: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585) synced to the
  • 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
  • 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
  • 07:30 jmm@dns1004: END - running authdns-update
  • {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
  • 07:27 jmm@dns1004: START - running authdns-update
  • 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734) (duration: 16m 04s)
  • 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
  • 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
  • 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
  • 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
  • 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
  • 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
  • 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734)
  • 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
  • 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
  • 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
  • 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
  • 06:23 slyngshede@dns1004: END - running authdns-update
  • 06:21 marostegui: Drop cu* tables from s3 bswiktionary T435965
  • 06:20 slyngshede@dns1004: START - running authdns-update
  • 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
  • 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - T436675
  • 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses (duration: 04m 42s)
  • 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
  • 05:03 tstarling@deploy1003: tstarling: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 05:01 tstarling@deploy1003: Started scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses
  • 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
  • 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
  • 04:29 tstarling@deploy1003: tstarling: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 04:25 tstarling@deploy1003: Started scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses
  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image

2026-09-01

  • 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for Merge branch 'master' into wmf_deploy (duration: 18m 05s)
  • 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
  • 21:47 jdlrobson@deploy1003: jdlrobson: Backport for Merge branch 'master' into wmf_deploy synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for Merge branch 'master' into wmf_deploy
  • 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248) (duration: 23m 55s)
  • 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
  • 21:20 jdlrobson@deploy1003: jdlrobson: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248)
  • 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
  • 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
  • 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
  • 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
  • 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
  • 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
  • 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
  • 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
  • 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
  • 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
  • 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
  • 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
  • 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
  • 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
  • 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
  • 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
  • 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
  • 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
  • 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
  • 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
  • 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
  • 19:06 sukhe@dns1004: END - running authdns-update
  • 19:03 sukhe@dns1004: START - running authdns-update
  • 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs T430837
  • 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade T431608
  • 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
  • 16:52 dancy@deploy1003: Started scap sync-world: testing
  • 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
  • 16:48 dancy@deploy1003: Started scap sync-world: testing
  • 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
  • 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
  • 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
  • 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
  • 15:51 moritzm: installing mesa security updates
  • 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
  • 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
  • 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
  • 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
  • 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
  • 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
  • 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
  • 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
  • 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
  • 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
  • 14:55 hashar: Restarted Jenkins on releases1003
  • 14:51 hashar: Restarted CI Jenkins on contint1003
  • 14:48 hashar: Restarting Gerrit primary on gerrit2003
  • 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
  • 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
  • 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
  • 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
  • 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
  • 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
  • 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
  • 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
  • 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
  • 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
  • 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
  • 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
  • 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
  • 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
  • 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
  • 14:24 moritzm: installing curl security updates
  • 14:24 jmm@dns1004: END - running authdns-update
  • 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
  • 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
  • 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
  • 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
  • 14:21 jmm@dns1004: START - running authdns-update
  • 14:21 jmm@dns1004: END - running authdns-update
  • 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
  • 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
  • 14:19 jmm@dns1004: START - running authdns-update
  • 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
  • 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
  • 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
  • 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
  • 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
  • 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
  • 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
  • 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
  • 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
  • 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
  • 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
  • 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
  • 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
  • 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
  • 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
  • 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
  • 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
  • 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
  • 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
  • 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
  • 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
  • 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490) (duration: 37m 37s)
  • 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
  • 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
  • 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # T418521
  • 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
  • 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
  • 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
  • 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
  • 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
  • 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
  • 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
  • 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
  • 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
  • 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
  • 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
  • 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
  • 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
  • 13:58 ladsgroup@dns1004: END - running authdns-update
  • 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
  • 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
  • 13:56 ladsgroup@dns1004: START - running authdns-update
  • 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
  • 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 13:56 ladsgroup@dns1004: END - running authdns-update
  • 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 13:53 ladsgroup@dns1004: START - running authdns-update
  • 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
  • 13:48 kharlan@deploy1003: kharlan: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
  • 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 13:41 jmm@dns1004: END - running authdns-update
  • 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
  • 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
  • 13:38 jmm@dns1004: START - running authdns-update
  • 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
  • 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
  • 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
  • 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
  • 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
  • 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
  • 13:26 kharlan@deploy1003: Started scap sync-world: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490)
  • 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
  • 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
  • 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
  • 13:23 aude@deploy1003: Finished scap sync-world: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922) (duration: 20m 24s)
  • 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
  • 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
  • 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
  • 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
  • 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
  • 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
  • 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
  • 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
  • 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
  • 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
  • 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM T429175
  • 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
  • 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
  • 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
  • 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
  • 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
  • 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
  • 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
  • 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
  • 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
  • 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
  • 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
  • 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
  • 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
  • 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
  • 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
  • 13:11 aude@deploy1003: aude: Continuing with deployment
  • 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
  • 13:07 aude@deploy1003: aude: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
  • 13:02 aude@deploy1003: Started scap sync-world: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922)
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
  • 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
  • 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
  • 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
  • 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
  • 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
  • 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
  • 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
  • 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
  • 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
  • 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
  • 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
  • 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848) (duration: 16m 25s)
  • 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
  • 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
  • 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
  • 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
  • 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
  • 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
  • 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
  • 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
  • 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
  • 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
  • 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
  • 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
  • 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
  • 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
  • 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
  • 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
  • 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
  • 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
  • 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
  • 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
  • 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
  • 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
  • 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)
  • 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
  • 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
  • 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
  • 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
  • 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
  • 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
  • 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
  • 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
  • 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
  • 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
  • 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
  • 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
  • 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
  • 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
  • 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
  • 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
  • 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
  • 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
  • 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
  • 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
  • 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
  • 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
  • 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
  • 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
  • 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
  • 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
  • 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
  • 11:12 moritzm: installing openjdk-21 security updates
  • 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
  • 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
  • 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
  • 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
  • 10:45 moritzm: installing Python 3.11 security updates
  • 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
  • 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
  • 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
  • 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
  • 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
  • 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
  • 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
  • 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
  • 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
  • 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
  • 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
  • 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
  • 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
  • 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
  • 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
  • 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
  • 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
  • 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
  • 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
  • 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
  • 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
  • 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
  • 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
  • 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
  • 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
  • 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
  • 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
  • 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production T427059', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
  • 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
  • 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
  • 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
  • 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
  • 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
  • 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
  • 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
  • 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
  • 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
  • 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
  • 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
  • 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
  • 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
  • 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
  • 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
  • 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
  • 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
  • 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
  • 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
  • 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
  • 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
  • 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
  • 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
  • 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
  • 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
  • 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
  • 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
  • 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
  • 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
  • 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
  • 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
  • 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
  • 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
  • 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
  • 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
  • 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
  • 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
  • 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
  • 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
  • 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
  • 06:29 moritzm: installing Java 17 security updates
  • 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
  • 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
  • 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
  • 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs T430837 (duration: 37m 30s)
  • 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs T430837
  • 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
  • 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
  • 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
  • 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
  • 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
  • 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
  • 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
  • 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply

Other archives

See Server Admin Log/Archives.