<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>GPU Rent Hub: incident history</title>
  <link>https://gpurenthub.com/status/</link>
  <atom:link href="https://gpurenthub.com/status-feed.xml" rel="self" type="application/rss+xml"/>
  <description>Every incident that affected more than one GPU Rent Hub customer, since March 2021.</description>
  <language>en</language>
  <item>
    <title>IAD-1 edge router failover Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-21-aug-2026-09-15-09-40-edt-iad-1-edge-router-failover-degraded</guid>
    <pubDate>21 Aug 2026 09:15 – 09:40 EDT</pubDate>
    <description>Primary edge lost a line card. Failover worked; 25 minutes of elevated latency (+30 ms) to some East-coast destinations while the secondary re-learned routes. Written by RK, on call.</description>
  </item>
  <item>
    <title>Monero deposit detection delayed Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-6-aug-2026-11-00-13-20-cdt-monero-deposit-detection-delayed-degraded</guid>
    <pubDate>6 Aug 2026 11:00 – 13:20 CDT</pubDate>
    <description>Our monerod fell behind the chain after a restart. Deposits were credited late (up to 2h20) but all were credited at the rate of their actual first confirmation, not the late detection time. Written by MT, on call.</description>
  </item>
  <item>
    <title>DFW-1 upstream fibre cut Outage</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-20-jul-2026-07-12-09-41-cdt-dfw-1-upstream-fibre-cut-outage</guid>
    <pubDate>20 Jul 2026 07:12 – 09:41 CDT</pubDate>
    <description>A construction crew cut the Lumen path into DFW-1. Cogent and HE stayed up but our BGP re-convergence took 14 minutes longer than it should have because of a stale prefix-list on the edge. Compute was unaffected; customers saw packet loss of 30–60% for 2h29 total. Root cause on the prefix-list has been fixed and a third diverse path (Zayo) was ordered. 07:12 loss detected on Lumen · 07:26 traffic drained · 08:50 Lumen splice complete · 09:41 resolved Written by RK, on call.</description>
  </item>
  <item>
    <title>DFW-1 cage 2 PDU replacement overran Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-7-jul-2026-02-00-03-10-cdt-dfw-1-cage-2-pdu-replacement-overran-degrad</guid>
    <pubDate>7 Jul 2026 02:00 – 03:10 CDT</pubDate>
    <description>Planned 60-minute window ran 70 minutes. 96 A6000/A40 instances rebooted once instead of the announced zero-reboot swap; the A-feed transfer switch did not hold. Affected customers were credited 3 days. Written by AV, on call.</description>
  </item>
  <item>
    <title>API latency Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-21-jun-2026-14-30-16-05-cdt-api-latency-degraded</guid>
    <pubDate>21 Jun 2026 14:30 – 16:05 CDT</pubDate>
    <description>Catalogue endpoint p99 went above 4 s after a stock-sync job locked the inventory table. Dashboards were slow; no instance was affected. Job moved to a read replica. Written by MT, on call.</description>
  </item>
  <item>
    <title>DFW-1 cooling loop alarm Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-2-mar-2026-22-10-23-55-cst-dfw-1-cooling-loop-alarm-degraded</guid>
    <pubDate>2 Mar 2026 22:10 – 23:55 CST</pubDate>
    <description>Facility CRAH unit fault raised inlet temperatures in cage 4 to 31 °C. We throttled 8-way H100/H200 nodes to 80% power for 1h45 rather than risk thermal shutdown. No reboots. Customers on affected nodes were credited 2 days. Written by RK, on call.</description>
  </item>
  <item>
    <title>Dashboard unavailable Outage</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-8-nov-2025-04-00-06-30-cst-dashboard-unavailable-outage</guid>
    <pubDate>8 Nov 2025 04:00 – 06:30 CST</pubDate>
    <description>A bad deploy of the dashboard broke login for everyone for 2.5 hours. The API and all instances were fine, which is cold comfort if you needed the console at 4 AM. We added a canary stage after this. Written by AV, on call.</description>
  </item>
  <item>
    <title>PDX-1 network flap Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-21-may-2025-16-40-17-05-pdt-pdx-1-network-flap-degraded</guid>
    <pubDate>21 May 2025 16:40 – 17:05 PDT</pubDate>
    <description>Upstream (HE) maintenance not communicated to us. 25 minutes of intermittent loss for PDX-1. Written by MT, on call.</description>
  </item>
  <item>
    <title>DFW-1 utility power event Outage</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-14-aug-2024-01-00-05-45-cdt-dfw-1-utility-power-event-outage</guid>
    <pubDate>14 Aug 2024 01:00 – 05:45 CDT</pubDate>
    <description>Utility feed dropped during a storm; generators started but the B-side UPS failed to carry cage 3 during transfer. 412 instances lost power for up to 4h45 (hard reboot). No data loss reported. Facility replaced the UPS string; we moved to 2N UPS on all cages by October 2024. Credited 7 days. Written by RK, on call.</description>
  </item>
  <item>
    <title>IAD-1 upstream degradation Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-3-feb-2023-10-20-12-00-est-iad-1-upstream-degradation-degraded</guid>
    <pubDate>3 Feb 2023 10:20 – 12:00 EST</pubDate>
    <description>Zayo path congested; we had one upstream at IAD-1 at the time. Second carrier (Lumen) was live by March 2023. Written by AV, on call.</description>
  </item>
  <item>
    <title>IAD-1 opening delayed Notice</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-30-nov-2022-all-day-iad-1-opening-delayed-notice</guid>
    <pubDate>30 Nov 2022 all day</pubDate>
    <description>First IAD-1 deployments were scheduled for 28 Nov and slipped two days waiting on a cross-connect. Pre-orders were credited a week. Written by MT, on call.</description>
  </item>
  <item>
    <title>Bitcoin deposit detection paused Degraded</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-12-oct-2021-13-00-14-10-cdt-bitcoin-deposit-detection-paused-degraded</guid>
    <pubDate>12 Oct 2021 13:00 – 14:10 CDT</pubDate>
    <description>Our first payment incident. bitcoind ran out of disk. Deposits were credited late, at the correct rate. Disk alarms were, in retrospect, a good idea. Written by RK, on call.</description>
  </item>
  <item>
    <title>DFW-1 switch failure Outage</title>
    <link>https://gpurenthub.com/status/#incidents</link>
    <guid isPermaLink="false">gpurenthub-incident-4-apr-2021-19-30-21-15-cdt-dfw-1-switch-failure-outage</guid>
    <pubDate>4 Apr 2021 19:30 – 21:15 CDT</pubDate>
    <description>The single top-of-rack switch in our first rack died. Forty RTX 3090s were offline for 1h45 while we swapped it. Every rack has had redundant ToR switches since May 2021. Written by AV, on call.</description>
  </item>
</channel>
</rss>