seaweedfs/weed/replication
Chris Lu 2386fa550a
grpc: don't tear down the shared master connection on a caller's own timeout (#9775)
A Canceled/DeadlineExceeded from the caller's per-request context was
treated like a dead channel: it closed the shared cached ClientConn and
cancelled every other in-flight RPC on it with "the client connection is
closing". Under a burst of concurrent chunk assigns (e.g. a large S3
multipart upload) one slow assign hitting its 10s attempt timeout could
poison the connection for all the rest, cascading into a flood of 500s.

Thread the caller's context into shouldInvalidateConnection and only
invalidate on Canceled/DeadlineExceeded while that context is still live,
which isolates the genuine stale-channel signal (a peer restart behind a
k8s Service VIP). To carry the context, add a ctx parameter to the
existing WithGrpcClient, WithMasterClient, and WithMasterServerClient; the
master assign and volume-lookup paths pass their per-attempt context and
every other caller passes context.Background().
2026-06-01 15:11:02 -07:00
..
repl_util go fmt 2026-04-10 17:31:14 -07:00
sink grpc: don't tear down the shared master connection on a caller's own timeout (#9775) 2026-06-01 15:11:02 -07:00
source grpc: don't tear down the shared master connection on a caller's own timeout (#9775) 2026-06-01 15:11:02 -07:00
sub notification.kafka: add SASL authentication and TLS support (#8832) 2026-03-29 13:45:54 -07:00
replicator_test.go Adjust rename events metadata format (#8854) 2026-03-30 18:25:11 -07:00
replicator.go fix: replication sinks upload ciphertext for SSE-encrypted objects (#8931) 2026-04-06 00:32:27 -07:00