한 대의 서버에서 모든 컴포넌트가 연결되는 것을 확인하면 distributed runtime이 완성된 것처럼 느껴진다.
내 PaaS도 한동안 그 단계까지는 꽤 안정적으로 보였다.
각 물리 노드에는 다음 세 컴포넌트를 둔다.
Node A
├─ PaaS(A)
├─ Server Ops(A)
└─ Security Guard(A)
Server Ops는 local runtime의 사실을 관찰하고, Security Guard는 그 사실을 해석하고, PaaS는 workload placement와 runtime intent를 관리한다.
문제는 이 구조가 두 번째 실제 노드에서도 같은 의미로 동작하는가였다.
fixture와 unit test만으로는 확인하기 어려운 질문이었다.
그래서 두 번째 Linux 서버에 같은 node-local triad를 실제로 붙여봤다.
두 번째 노드에서도 identity를 새로 만들지 않았다
새 노드를 bootstrap한다고 해서 매번 새로운 identity를 발급하면 runtime lineage가 끊긴다.
두 번째 서버에는 이미 Server Ops가 사용하던 durable node identity가 있었다.
그래서 기존 identity를 그대로 재사용했다.
PHYSICAL NODE
→ durable node identity
restart
→ reuse identity
new local component
→ join the same node identity
PaaS, Ops, Guard가 각각 자기 식별자를 임의로 만들어내는 것이 아니라 같은 물리 노드에 속한다는 사실을 공유해야 한다.
그래야 다음 의미가 유지된다.
Ops(B)
= Node B observation authority
Guard(B)
= Node B security interpretation authority
PaaS(B)
= Node B local platform participant
Handshake보다 authority를 확인했다
세 컴포넌트의 연결 자체는 모두 성공했다.
PaaS <-> Ops
PaaS <-> Guard
Ops <-> Guard
하지만 socket이 열렸다는 것만으로 acceptance를 주지는 않았다.
확인하고 싶은 것은 메시지가 실제 production path를 지나면서도 authority가 섞이지 않는지였다.
실제 흐름은 다음처럼 검증했다.
PaaS -> Ops
→ operation intent
PaaS -> Guard
→ security observation
Ops -> PaaS
→ node runtime signal
Guard -> PaaS
→ security signal
Ops -> Guard
→ process/resource facts
Guard -> Ops
→ containment intent
→ containment result
중요한 것은 모든 flow가 "성공"을 반환하는 것이 아니었다.
오히려 containment에서는 의도적으로 REJECTED / ACTUATOR_NOT_BOUND가 나왔다.
REJECTED가 PASS일 수 있다
이번 검증 시점에는 실제 destructive actuator를 연결하지 않았다.
그 상태에서 Guard가 containment를 요청했을 때 Ops가 임의로 무엇인가를 죽이면 오히려 잘못된 동작이다.
기대한 결과는:
ContainmentIntent
→ target/contract accepted
→ no actuator bound
→ REJECTED / ACTUATOR_NOT_BOUND
이었다.
실제 결과도 그랬고 Guard는 이 terminal result를 lifecycle store에 기록했다.
따라서 acceptance는 단순한 SUCCESS 여부가 아니었다.
SUCCESS == PASS
가 아니라:
EXPECTED SAFE RESULT
== PASS
였다.
보안이나 운영 시스템에서는 실패 응답도 contract가 정확하게 지켜졌다는 evidence가 될 수 있다.
Runtime fact를 Security decision으로 자동 승격하지 않았다
Ops에서 Guard로 process event와 resource pressure도 보냈다.
여기서 또 하나 확인할 것이 있었다.
관찰 사실이 들어왔다고 해서 Guard가 곧바로 "공격"이라고 판단하면 안 된다.
ProcessEvent
ResourcePressure
= observation
NOT
= malicious classification
실제 두 번째 노드에서도 이 이벤트들은 neutral runtime facts로 받아들여졌다.
분류와 containment decision은 별도 semantics를 따라야 한다는 경계가 유지됐다.
실제 노드 acceptance가 끝났는데도 repair code를 다시 읽었다
두 번째 물리 노드에서 local triad flow는 통과했다.
그 뒤 runtime failure mode 관련 repair code를 다시 리뷰했다.
대상은 대략 다음 네 가지였다.
PaaS stale local projection
Ops/Guard stream replacement race
stale Unix-socket crash restart
Guard containment retention
repair의 방향 자체는 대부분 맞았다.
그런데 실제 코드를 컴파일하고 production-shaped path를 다시 따라가니 몇 가지 결함과 검증 gap이 드러났다.
거의 실행되지 않는 branch도 contract는 맞아야 한다
Server Ops의 stale Unix-socket recovery에는 OSError fallback branch가 있었다.
현재 CPython에서는 더 구체적인 ConnectionRefusedError branch가 먼저 잡혀 사실상 도달하기 어려운 코드였다.
하지만 fallback에서 stale socket unlink helper를 호출할 때 필요한 argument 하나가 빠져 있었다.
intended:
unlink_stale_socket(path, expected_identity)
actual:
unlink_stale_socket(path)
현재 환경에서 거의 실행되지 않는다는 사실은 correctness 근거가 아니다.
platform behavior가 달라지는 순간 intended recovery 대신 TypeError가 날 수 있었다.
crash-restart path의 작은 결함은 평상시에는 잘 보이지 않는다.
논리적으로 맞는 repair가 TypeScript contract에서는 깨질 수 있다
PaaS 쪽 repair는 disconnect 시 local projection을 지우도록 만들었다.
의도는 맞았다.
Ops disconnect
→ nodeRuntime/workloads invalidate
Guard disconnect
→ security invalidate
그런데 TypeScript 설정에 exactOptionalPropertyTypes: true가 켜져 있었다.
optional field에 명시적으로 undefined를 대입하는 구현이 compile error가 됐다.
SEMANTICALLY RIGHT
!=
COMPILABLE IN THIS PROJECT
실제 typecheck를 돌리지 않았다면 source review만으로 지나갈 수 있는 종류였다.
Unix socket 테스트는 운영체제 경계도 가진다
Security Guard의 stale socket test는 긴 temporary path를 사용했다.
Linux에서는 괜찮아 보였지만 macOS의 AF_UNIX path 길이 제한에서는 socket bind 자체가 실패했다.
test logic correct
+
test path too long
=
false failure on another OS
짧은 temp path로 바꾸고 나서야 실제 검증 대상만 볼 수 있었다.
portable runtime code를 만든다면 test harness 역시 platform constraint를 가진다.
Private method 테스트는 실제 wiring을 증명하지 못했다
PaaS의 stale projection test는 처음에 private invalidation method를 직접 호출했다.
그 method가 올바르게 state를 지운다는 것은 증명했다.
하지만 진짜 질문은 이것이었다.
socket이 실제로 끊겼을 때 production reconnect path가 그 method를 호출하는가?
그래서 실제 Unix-socket peer를 띄우고 production runtime을 연결한 뒤 다음 transition을 검증했다.
Ops signal
→ projection populated
Ops socket close
→ Ops projection cleared
→ Guard projection preserved
Guard socket close
→ Guard projection cleared
Ops reconnect
→ fresh projection repopulated
Guard reconnect
→ fresh security projection repopulated
METHOD TEST
!=
WIRING TEST
였다.
Crash-restart도 stale file을 직접 남겨서 확인했다
Unix domain socket은 process가 비정상 종료되면 pathname이 남을 수 있다.
다음 startup에서 단순히 "파일이 있으면 삭제"하면 위험하다.
그 파일이 실제 live listener일 수도 있기 때문이다.
그래서 recovery logic은 다음처럼 움직인다.
existing socket path
→ probe connect
live listener
→ preserve / refuse takeover
connection refused
→ lstat again
→ compare dev + inode
same stale socket identity
→ unlink
identity changed
→ preserve / abort
이 설계를 unit test만으로 끝내지 않았다.
실제로 stale Unix socket pathname을 남긴 뒤 fresh runtime component가 같은 path에서 다시 bind할 수 있는지 확인했다.
Server Ops와 Security Guard 둘 다 full component entry path에서 crash-restart recovery를 검증했다.
Distributed system 테스트는 state보다 transition을 봐야 했다
이번 검증에서 가장 유용했던 것은 정상 상태의 snapshot보다 transition이었다.
connected
→ disconnected
→ invalidated
→ reconnected
→ refreshed
그리고:
running
→ crash residue
→ restart
→ safe recovery
같은 흐름이었다.
정상 상태에서 PaaS가 Ops를 보고 Guard가 Ops를 본다는 것만으로는 stale state, replacement race, crash residue를 검증할 수 없다.
문제는 대부분 상태 사이의 경계에서 생겼다.
두 번째 물리 노드가 준 것은 "한 대 더 동작한다"는 증거만이 아니었다
두 번째 서버를 붙인 목적은 단순히 node count를 1에서 2로 늘리는 것이 아니었다.
한 노드에서 만든 authority model이 다른 실제 환경에서도 같은 의미를 유지하는지 확인하는 일이었다.
그리고 그 뒤 failure mode를 다시 검증하면서 알게 된 것은 더 단순했다.
UNIT TEST PASS
+
SINGLE NODE PASS
!=
RUNTIME TRANSITION PROVEN
distributed runtime에서는 disconnect, reconnect, crash, replacement, stale state처럼 시간에 따라 바뀌는 경계가 제품의 일부다.
다음 노드를 추가할 때 필요한 것은 테스트 개수를 더 늘리는 것만이 아니다.
실제 authority와 transport가 움직이는 transition을 검증할 수 있어야 한다.
그 단계에 와서야 "두 번째 노드에서도 된다"는 말이 단순한 데모가 아니라 운영 evidence가 된다.