> ## Content Index
> Fetch the complete content index at: https://blog.sorune.org/llms.txt
> Use this file to discover other available public pages before exploring further.

# 두 번째 서버를 붙여보니 테스트가 놓친 경계가 보였다 — Node-Local Runtime을 실제로 검증한 방법
- URL: https://blog.sorune.org/second-node-runtime-validation-boundaries/
- Published: 2026-09-22T02:20:55.000Z
- Updated: 2026-09-22T02:20:55.000Z
- Description: PaaS, Server Ops, Security Guard를 두 번째 물리 노드에 실제로 연결하면서 handshake, observation, containment, disconnect/reconnect, crash-restart를 검증했다. 단위 테스트를 넘어서 실제 runtime transition을 확인해야 했던 이유를 정리한다.
- Author: Sorune — Engineering, Photography & Lab Notes
- Tags: Engineering, PaaS, Server Ops, Security Guard, Distributed Systems, Runtime Validation

한 대의 서버에서 모든 컴포넌트가 연결되는 것을 확인하면 distributed runtime이 완성된 것처럼 느껴진다.

내 PaaS도 한동안 그 단계까지는 꽤 안정적으로 보였다.

각 물리 노드에는 다음 세 컴포넌트를 둔다.

```text
Node A
├─ PaaS(A)
├─ Server Ops(A)
└─ Security Guard(A)

```

Server Ops는 local runtime의 사실을 관찰하고, Security Guard는 그 사실을 해석하고, PaaS는 workload placement와 runtime intent를 관리한다.

문제는 이 구조가 **두 번째 실제 노드에서도 같은 의미로 동작하는가**였다.

fixture와 unit test만으로는 확인하기 어려운 질문이었다.

그래서 두 번째 Linux 서버에 같은 node-local triad를 실제로 붙여봤다.

## 두 번째 노드에서도 identity를 새로 만들지 않았다

새 노드를 bootstrap한다고 해서 매번 새로운 identity를 발급하면 runtime lineage가 끊긴다.

두 번째 서버에는 이미 Server Ops가 사용하던 durable node identity가 있었다.

그래서 기존 identity를 그대로 재사용했다.

```text
PHYSICAL NODE
→ durable node identity

restart
→ reuse identity

new local component
→ join the same node identity

```

PaaS, Ops, Guard가 각각 자기 식별자를 임의로 만들어내는 것이 아니라 같은 물리 노드에 속한다는 사실을 공유해야 한다.

그래야 다음 의미가 유지된다.

```text
Ops(B)
= Node B observation authority

Guard(B)
= Node B security interpretation authority

PaaS(B)
= Node B local platform participant

```

## Handshake보다 authority를 확인했다

세 컴포넌트의 연결 자체는 모두 성공했다.

```text
PaaS <-> Ops
PaaS <-> Guard
Ops  <-> Guard

```

하지만 socket이 열렸다는 것만으로 acceptance를 주지는 않았다.

확인하고 싶은 것은 메시지가 실제 production path를 지나면서도 authority가 섞이지 않는지였다.

실제 흐름은 다음처럼 검증했다.

```text
PaaS -> Ops
→ operation intent

PaaS -> Guard
→ security observation

Ops -> PaaS
→ node runtime signal

Guard -> PaaS
→ security signal

Ops -> Guard
→ process/resource facts

Guard -> Ops
→ containment intent
→ containment result

```

중요한 것은 모든 flow가 "성공"을 반환하는 것이 아니었다.

오히려 containment에서는 의도적으로 REJECTED / ACTUATOR\_NOT\_BOUND가 나왔다.

## REJECTED가 PASS일 수 있다

이번 검증 시점에는 실제 destructive actuator를 연결하지 않았다.

그 상태에서 Guard가 containment를 요청했을 때 Ops가 임의로 무엇인가를 죽이면 오히려 잘못된 동작이다.

기대한 결과는:

```text
ContainmentIntent
→ target/contract accepted
→ no actuator bound
→ REJECTED / ACTUATOR_NOT_BOUND

```

이었다.

실제 결과도 그랬고 Guard는 이 terminal result를 lifecycle store에 기록했다.

따라서 acceptance는 단순한 SUCCESS 여부가 아니었다.

```text
SUCCESS == PASS

```

가 아니라:

```text
EXPECTED SAFE RESULT
== PASS

```

였다.

보안이나 운영 시스템에서는 실패 응답도 contract가 정확하게 지켜졌다는 evidence가 될 수 있다.

## Runtime fact를 Security decision으로 자동 승격하지 않았다

Ops에서 Guard로 process event와 resource pressure도 보냈다.

여기서 또 하나 확인할 것이 있었다.

관찰 사실이 들어왔다고 해서 Guard가 곧바로 "공격"이라고 판단하면 안 된다.

```text
ProcessEvent
ResourcePressure
= observation

NOT
= malicious classification

```

실제 두 번째 노드에서도 이 이벤트들은 neutral runtime facts로 받아들여졌다.

분류와 containment decision은 별도 semantics를 따라야 한다는 경계가 유지됐다.

## 실제 노드 acceptance가 끝났는데도 repair code를 다시 읽었다

두 번째 물리 노드에서 local triad flow는 통과했다.

그 뒤 runtime failure mode 관련 repair code를 다시 리뷰했다.

대상은 대략 다음 네 가지였다.

```text
PaaS stale local projection
Ops/Guard stream replacement race
stale Unix-socket crash restart
Guard containment retention

```

repair의 방향 자체는 대부분 맞았다.

그런데 실제 코드를 컴파일하고 production-shaped path를 다시 따라가니 몇 가지 결함과 검증 gap이 드러났다.

## 거의 실행되지 않는 branch도 contract는 맞아야 한다

Server Ops의 stale Unix-socket recovery에는 OSError fallback branch가 있었다.

현재 CPython에서는 더 구체적인 ConnectionRefusedError branch가 먼저 잡혀 사실상 도달하기 어려운 코드였다.

하지만 fallback에서 stale socket unlink helper를 호출할 때 필요한 argument 하나가 빠져 있었다.

```text
intended:
unlink_stale_socket(path, expected_identity)

actual:
unlink_stale_socket(path)

```

현재 환경에서 거의 실행되지 않는다는 사실은 correctness 근거가 아니다.

platform behavior가 달라지는 순간 intended recovery 대신 TypeError가 날 수 있었다.

crash-restart path의 작은 결함은 평상시에는 잘 보이지 않는다.

## 논리적으로 맞는 repair가 TypeScript contract에서는 깨질 수 있다

PaaS 쪽 repair는 disconnect 시 local projection을 지우도록 만들었다.

의도는 맞았다.

```text
Ops disconnect
→ nodeRuntime/workloads invalidate

Guard disconnect
→ security invalidate

```

그런데 TypeScript 설정에 exactOptionalPropertyTypes: true가 켜져 있었다.

optional field에 명시적으로 undefined를 대입하는 구현이 compile error가 됐다.

```text
SEMANTICALLY RIGHT
!=
COMPILABLE IN THIS PROJECT

```

실제 typecheck를 돌리지 않았다면 source review만으로 지나갈 수 있는 종류였다.

## Unix socket 테스트는 운영체제 경계도 가진다

Security Guard의 stale socket test는 긴 temporary path를 사용했다.

Linux에서는 괜찮아 보였지만 macOS의 AF\_UNIX path 길이 제한에서는 socket bind 자체가 실패했다.

```text
test logic correct
+
test path too long
=
false failure on another OS

```

짧은 temp path로 바꾸고 나서야 실제 검증 대상만 볼 수 있었다.

portable runtime code를 만든다면 test harness 역시 platform constraint를 가진다.

## Private method 테스트는 실제 wiring을 증명하지 못했다

PaaS의 stale projection test는 처음에 private invalidation method를 직접 호출했다.

그 method가 올바르게 state를 지운다는 것은 증명했다.

하지만 진짜 질문은 이것이었다.

> socket이 실제로 끊겼을 때 production reconnect path가 그 method를 호출하는가?

그래서 실제 Unix-socket peer를 띄우고 production runtime을 연결한 뒤 다음 transition을 검증했다.

```text
Ops signal
→ projection populated

Ops socket close
→ Ops projection cleared
→ Guard projection preserved

Guard socket close
→ Guard projection cleared

Ops reconnect
→ fresh projection repopulated

Guard reconnect
→ fresh security projection repopulated

```

```text
METHOD TEST
!=
WIRING TEST

```

였다.

## Crash-restart도 stale file을 직접 남겨서 확인했다

Unix domain socket은 process가 비정상 종료되면 pathname이 남을 수 있다.

다음 startup에서 단순히 "파일이 있으면 삭제"하면 위험하다.

그 파일이 실제 live listener일 수도 있기 때문이다.

그래서 recovery logic은 다음처럼 움직인다.

```text
existing socket path
→ probe connect

live listener
→ preserve / refuse takeover

connection refused
→ lstat again
→ compare dev + inode

same stale socket identity
→ unlink

identity changed
→ preserve / abort

```

이 설계를 unit test만으로 끝내지 않았다.

실제로 stale Unix socket pathname을 남긴 뒤 fresh runtime component가 같은 path에서 다시 bind할 수 있는지 확인했다.

Server Ops와 Security Guard 둘 다 full component entry path에서 crash-restart recovery를 검증했다.

## Distributed system 테스트는 state보다 transition을 봐야 했다

이번 검증에서 가장 유용했던 것은 정상 상태의 snapshot보다 transition이었다.

```text
connected
→ disconnected
→ invalidated
→ reconnected
→ refreshed

```

그리고:

```text
running
→ crash residue
→ restart
→ safe recovery

```

같은 흐름이었다.

정상 상태에서 PaaS가 Ops를 보고 Guard가 Ops를 본다는 것만으로는 stale state, replacement race, crash residue를 검증할 수 없다.

문제는 대부분 상태 사이의 경계에서 생겼다.

## 두 번째 물리 노드가 준 것은 "한 대 더 동작한다"는 증거만이 아니었다

두 번째 서버를 붙인 목적은 단순히 node count를 1에서 2로 늘리는 것이 아니었다.

한 노드에서 만든 authority model이 다른 실제 환경에서도 같은 의미를 유지하는지 확인하는 일이었다.

그리고 그 뒤 failure mode를 다시 검증하면서 알게 된 것은 더 단순했다.

```text
UNIT TEST PASS
+
SINGLE NODE PASS

!=

RUNTIME TRANSITION PROVEN

```

distributed runtime에서는 disconnect, reconnect, crash, replacement, stale state처럼 **시간에 따라 바뀌는 경계**가 제품의 일부다.

다음 노드를 추가할 때 필요한 것은 테스트 개수를 더 늘리는 것만이 아니다.

실제 authority와 transport가 움직이는 transition을 검증할 수 있어야 한다.

그 단계에 와서야 "두 번째 노드에서도 된다"는 말이 단순한 데모가 아니라 운영 evidence가 된다.