Tough network system in non-cooperative environment

By constructing a three-layer architecture and a real-time situational awareness mechanism, the communication security and resilience issues in untrusted network environments are solved, enabling self-organization, self-adaptation, and self-repair capabilities in non-cooperative environments, thereby improving the anti-interference and availability of the communication system.

CN121151084APending Publication Date: 2025-12-16FUDAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511455267.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-12-16

Smart Images

  • Figure CN121151084A_ABST
    Figure CN121151084A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security, and particularly relates to a tough network system in a non-cooperative environment. The architecture of the tough network system comprises a management layer, a VPS cloud node layer and a user layer, and in combination with a pluggable anonymous communication subsystem, an interlayer communication mechanism based on gRPC and an onion network management technology based on Stem, the self-organization, self-adaption and self-repairing capabilities of the system in a complex network environment are realized; comprising the following steps: constructing a basic anonymous network topology by adopting directory authorization (DA) nodes, carrying out real-time monitoring by applying a situation awareness technology, and ensuring node security by virtue of a self-destruction mechanism, so that the dynamic toughness and high availability of a network system are realized while the anonymous communication and anti-tracking capabilities are enhanced. According to the method, the anti-interference performance and the self-adaptive capability of the network can be effectively improved, and the method can be widely applied to scenes with high requirements on communication privacy and network toughness, such as government affair systems, units, cross-border e-commerce platforms and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of network security, and particularly relates to a resilient network system for non-cooperative environment. BACKGROUND

[0002] With the continuous evolution of global network security situation, especially in the background of cross-border communication, sensitive information transmission and the widespread existence of untrusted intermediate networks, the security, reliability and availability of traditional communication architecture are facing unprecedented challenges. Current mainstream research and deployment schemes are mostly based on the assumption of trusted networks, relying on the cooperation of infrastructure providers, cloud service platforms and operators, and cannot effectively deal with the network environment where non-cooperation or malicious behavior occurs frequently.

[0003] The so-called non-cooperative network environment refers to the untrusted transmission relay in the information transmission path, including but not limited to Internet service providers who do not provide encryption transmission protection, regulatory agencies that limit information flow, and content distribution network (CDN) service providers who have the ability to detect and block traffic. In these environments, user communication is vulnerable to eavesdropping, interference and blocking, which affects data transmission efficiency and overall system availability.

[0004] In order to enhance the survivability of network systems in untrusted environments, in recent years, international standard organizations and governments have proposed to build network resilience architectures with dynamic adaptability. For example, the National Institute of Standards and Technology (NIST) has published a functional framework for resilient network systems, the European Union has promoted the standardized design of network device security through the Network Resilience Act, and China has also proposed a network architecture with "endogenous security" by relevant academicians. However, for non-cooperative and untrusted network environments, existing research mostly stays in theoretical exploration or partial technical implementation, and there is still a lack of systematic technical solutions for dynamic and complex environments.

[0005] Under this background, there is an urgent need for a network architecture and implementation mechanism that can maintain communication anonymity, security and service availability in dynamic restricted environments, thereby supporting the security communication needs of key scenarios including government systems, classified communication and cross-border data transmission. SUMMARY

[0006] The purpose of the present application is to provide a resilient network system (RN) for non-cooperative environment to address the challenges of ensuring communication security and resilience in non-cooperative and untrusted network environments.

[0007] The resilient network system for non-cooperative environment provided by the application is composed of a three-layer architecture system of a management layer, a Virtual Private Server (VPS) cloud node layer and a user layer, combines a pluggable anonymous communication subsystem, an interlayer communication mechanism based on gRPC and an onion network management technology based on Stem, and realizes the self-organization, self-adaptation and self-repairing capability of the system in a complex network environment. The key technologies include the construction of a basic anonymous network topology by using Directory Authority (DA) nodes, the real-time monitoring by using the situation awareness technology, and the node safety ensured by the self-destruction mechanism, so as to realize the dynamic resilience and high availability of the network system while strengthening the anonymous communication and anti-tracking capability. The application can effectively improve the anti-interference and self-adaptation capability of the network, and can be widely applied to the scenes such as government systems, units, cross-border e-commerce platforms and the like which have high requirements for communication privacy and network resilience.

[0008] The resilient network system for non-cooperative environment provided by the application includes the following aspects: (I) Architecture design of the resilient network system for non-cooperative environment The architecture includes three layers: a management layer, a VPS cloud node layer and a user layer. The management layer is deployed on a local trusted server and serves as the control center of the system, and undertakes the functions of identity authentication, node scheduling, directory authority management and situation awareness analysis. The cloud node layer is composed of multiple VPS, each node is deployed with an anonymous communication subsystem and a node management subsystem, and is responsible for anonymous protocol communication such as Tor and node state reporting and instruction execution. The user layer completes the initial authentication and accesses the main network through the access network mechanism. The three-layer architecture emphasizes the clear division of labor and anti-interference capability of the system, and the whole is shown in Figure 1 , wherein: The management layer, as the centralized control unit of the system, is deployed on a trusted server, and the main functional modules in the internal thereof include: a front-end interface for human-computer interaction and state display, which is linked with the back-end data to access configuration and log data; a core Management gRPC Server, which realizes comprehensive management and control through internal gRPC services, regular node gRPC services for VPS nodes and hidden service gRPC based on onion services; in addition, it also includes a RESTful API interface for external integration, a management layer internal client, and an alarm log and traffic analysis module for situation awareness, and realizes the visualization of network state through node topology and link.

[0009] The VPS cloud node layer is composed of a plurality of VPS distributed in geography, each VPS node being deployed with a node management subsystem and an anonymous communication subsystem. The node management subsystem includes gRPC API and HS gRPC API interfaces, and a node client and a Stem controller, and is responsible for responding to management layer instructions and reporting node status. The anonymous communication subsystem is deployed in the form of a container, including an access container and an anonymous container, each container being internally integrated with a Stem API for calling by the Stem controller, and running a specific onion instance and related management components such as a Stem node manager, a Stem link manager and a Stem flow manager. Meanwhile, the user space of the VPS node is deployed with a traffic detection and intrusion module, and the kernel space utilizes eBPF hooks and system call interfaces for underlying data processing and monitoring, supplemented by a self-destruction mechanism to deal with serious security threats.

[0010] The overall architecture achieves high resilience and communication security of the network system in a non-cooperative environment through the cooperative work of the above components.

[0011] The architecture emphasizes the self-organization, self-adaptation and self-repairing capability of the communication system in an untrusted and non-cooperative network environment. Each VPS node includes two key subsystems: an anonymous communication subsystem and a node management subsystem. The anonymous communication subsystem is a pluggable module that supports multiple protocols (such as Tor, Wireguard, Snowflake, VPN, etc.), enhancing the dynamic adaptability of the network; the node management subsystem is responsible for communication with the management layer and executing key instructions such as identity reporting, role registration and self-destruction.

[0012] The management layer specifies some nodes as directory authority (DA) nodes to build a functional minimum onion network. The management layer uses gRPC protocol as a communication bridge to broadcast DA information to all relay nodes, achieving automatic expansion and contraction of the network topology. In addition, to ensure the anonymity and anti-tracking capability of the system, all sensitive communications are completed based on the onion service gRPC.

[0013] In addition, by setting up a situation awareness module, the system can obtain the behavior state and external traffic characteristics of the nodes in real time, such as node load, inflow / outflow rate, port anomaly, connection request failure rate, etc., and can trigger the self-destruction mechanism when an attack occurs or an anomaly is detected, so that the node is immediately excluded from the system and the local sensitive data is cleared.

[0014] (II) Inter-layer communication design in a resilient network system in a non-cooperative environment Figure 2The timing and flow details of inter-layer communication are described in detail. Specifically, the gRPC protocol adopted by the system is divided into two types of interfaces: basic gRPC and Hidden Service gRPC (HS gRPC) based on onion domain name. The basic gRPC is used for node initialization, state synchronization, configuration distribution, and other basic control communication; the HS gRPC is used for the interaction of sensitive information, such as DA fingerprint distribution, situation awareness data upload, user authentication, and recommended routing information distribution. The VPS node maintains a continuous connection with the management layer through bidirectional streaming gRPC, forming a blocking command flow channel, so that the management surface can control the node behavior in real time. The whole communication process runs asynchronously, is scheduled based on the coroutine concurrent framework, and improves the system response efficiency in the scenario of massive VPS node deployment.

[0015] The VPS node maintains a continuous connection with the management layer through bidirectional streaming gRPC, forming a blocking command flow channel, so that the management surface can push the instructions of control behavior to the node in real time, such as role setting, configuration change or execution destruction operation. The whole communication process is designed to run asynchronously and is scheduled based on the coroutine concurrent framework, which aims to improve the system response efficiency and concurrent processing capacity in the scenario of massive VPS node deployment.

[0016] The basic gRPC interface is mainly used for regular interaction activities in the system initialization process, including node registration, DA fingerprint reporting, role allocation command reception and execution, health check report submission, and self-destruction instruction distribution and confirmation. Correspondingly, the Hidden Service gRPC interface uses the onion domain name of the onion service to build a secure communication channel, which is specially used to transmit highly sensitive data such as link information, situation awareness data, user authentication credentials, DA update information, and routing recommendation.

[0017] To further improve the communication efficiency and connection stability, the VPS node is designed to actively establish a long connection to the management surface through a bidirectional streaming gRPC interface. This connection usually remains in a blocking state during the network operation period, allowing the management layer to push operation instructions to specific nodes at any time. This mechanism ensures that the system has high-level real-time scheduling capability and fast response speed to potential attack events.

[0018] In terms of concurrent processing, this aspect introduces an asynchronous multi-coroutine mechanism, integrating the gRPC interface into an asynchronous IO framework. This integration enables the system to support parallel execution of multiple tasks, including but not limited to log analysis, link state detection, command reception and processing, traffic tagging, and session authentication, thereby significantly improving the overall service response capability and throughput of the system in a large-scale node deployment environment.

[0019] Referring toFigure 3 The figure is a resilient network module communication flow chart, which clearly shows the interaction path between the management surface and the VPS node and the division of responsibilities of each processing module. For ease of understanding, the following process can be explained step by step: (1) Node startup and registration: After the "node management client" on the VPS node starts, it initiates a registration request to the gRPC server of the management surface. The registration information includes node identification, available anonymous protocol module list, regional information, and initial software and hardware fingerprints, etc.

[0020] (2) Establish a two-way long connection: After successful registration, the node and the management surface establish a two-way long connection (two-way streaming RPC) on the basis of gRPC channel, which is used for heartbeat, index reporting, and receiving issued command flow. The management surface can immediately issue control instructions through the long connection, and the node can also report the running situation through the channel.

[0021] (3) Index reporting and situation summary: The node management client reports the heartbeat and monitoring indicators (CPU / memory, port connectivity, traffic statistics, abnormal detection alarm, etc.) to the management surface gRPC server according to the configuration period, and the management surface stores these data into the situation awareness module for strategy calculation and display.

[0022] (4) Control command issuance (Stem control): When the management surface decides to issue a running or configuration adjustment command, it issues a Stem control command (such as updating torrc, reloading HiddenService, role transfer instruction, etc.) to the node management client through gRPC. The node management client receives the command and calls the API of the local Stem module to execute specific operations.

[0023] (5) Anonymous communication module execution (onion module): The Stem module interacts with the local onion module (Tor or other anonymous instances) through its API to realize the configuration, startup, stop, or routing adjustment of hidden services, and complete the fine control of the anonymous communication stack.

[0024] (6) Command execution feedback and closed-loop confirmation: After the node completes the command execution, it returns the execution result and status to the management surface through gRPC, and the management surface records the execution feedback in the situation panel and the log to ensure operation closed loop and audit traceability.

[0025] (7) Front-end interaction and manual intervention path: The management surface also provides a management surface client and a front-end interaction interface (REST), and operation and maintenance personnel can view the situation, trigger manual operation or adjust the strategy through the front-end interface. The operation of the front-end will be translated by the management surface gRPC server into the issued command to the node.

[0026] The figure clearly shows the complete closed-loop process of "reporting information → establishing connection / returning command flow → issuing Stem control command → Stem API calling onion module" through the coupling of arrows and modules, and emphasizes the division of responsibilities of gRPC as the control channel and REST as the management interface interacting with the front end. This process helps auditors quickly understand the interaction mechanism and control loop of the system under normal and abnormal conditions.

[0027] (Three) The design of the resilient network system in the non-cooperative environment based on Stem implementation of onion network management includes: The Stem library written in Python language is used as a tool for managing and monitoring the onion network. This tool allows dynamic remote control of the onion instances running on various VPS nodes. System administrators can start, stop, reload, or reconfigure onion instance processes through the interface provided by Stem, thereby supporting role switching (e.g., from a normal relay to a bridge or exit), dynamic registration and deregistration of hidden services, verification of communication path effectiveness, and on-demand reconstruction of network topology throughout the node's life cycle.

[0028] The torrc configuration file (onion routing configuration file) of each node is automatically generated by the node management subsystem on the node based on the latest policies issued by the management interface. The generated configuration can be dynamically loaded and updated during onion instance runtime through the Stem interface without restarting the onion process, ensuring service continuity. The Directory Authority (DA) node plays the role of the onion network DA server in the network, and its running state is monitored in real time by the management interface. The management layer can obtain the current detailed state information of the DA node through the Stem library, including but not limited to service fingerprints, port connectivity test results, and network consensus file generation progress, among other key indicators.

[0029] When the management interface detects that a DA node has unexpectedly left the network, behaves abnormally (such as response timeout, error logs, etc.), or is determined to have been compromised, the management interface will issue instructions to the abnormal node through the gRPC interface, notifying it to execute a self-destruction program. At the same time, the management interface will select a new, healthy node to complete the migration of the DA role according to the pre-set strategy or real-time evaluation results. The entire role migration process is completed through real-time and precise manipulation of the Stem library, ensuring the resilience of the onion network core structure and the continuous availability of anonymous communication links in the face of node failures or attacks.

[0030] Furthermore, by utilizing the event subscription mechanism provided by the Stem library, the system can proactively monitor state changes of all hidden services within the network. These changes include, but are not limited to, successful establishment or unexpected destruction of Onion services, new link connection events, and potential denial-of-service attack warnings. This mechanism enables the system to promptly detect any anomalies in the anonymous communication environment, supporting rapid response and troubleshooting.

[0031] (iv) User verification design in resilient network systems under non-cooperative environments, specifically: To provide effective identity isolation and network buffering mechanisms before users access the main system, this invention proposes an access network mechanism. This mechanism involves constructing a small, dedicated access network independent of the main onion network. This dedicated access network consists of a group of controlled relay nodes, whose sole function is to handle the verification of user authentication credentials and the issuance of permissions. Users who successfully authenticate through the access network will obtain the DA information of the main network, thereby gaining access to the truly anonymous communication network. This design not only ensures the integrity and unpredictability of the main network structure but also effectively reduces the risk of direct exposure of the main network.

[0032] This dedicated access network consists of a group of strictly controlled VPS nodes. Their specific functions are handling new user access authentication requests, configuring fine-grained permissions, and distributing initial network information (such as the main network's entry point or DA information) to users upon successful authentication. The anonymous communication subsystems deployed within the access network nodes can also be built based on anonymous protocols such as the Onion routing protocol or Snowflake, but their use is strictly limited to the authentication process, ensuring that these nodes do not participate in the transmission of Relay or Bridge data within the main anonymous network.

[0033] Figure 4 This is a diagram illustrating the design and communication process of the access network module in a resilient network system, clearly showing the detailed steps of user authentication and access. Figure 4 Users first establish a connection with the management plane through the access network. This user client integrates a node management module, an access network container, and an anonymous network container. The specific access process is as follows: (1) Authentication request initiation: User client, such as Figure 4 As shown in section [1], authentication requests and related credentials are securely submitted to the management layer through the hidden service gRPC interface (HS gRPC API) provided by the access network container. Each access network container can be configured to run a dedicated Onion instance to handle such requests.

[0034] (2) Management verification and authorization: After receiving the user's authentication request, the gRPC server of the management layer will strictly verify the user's legitimacy based on the user policy preset in the system and the user credential information stored in the database.

[0035] (3) Issuance of authorization information: Once the user's identity verification is successful, the management team, such as Figure 4 As shown in the middle number [2], the current valid DA information of the main anonymous network, a list of recommended available relay nodes, and the access policy specially configured for the user account (such as the types of services allowed to be accessed, bandwidth restrictions, etc.) will be accurately returned to the user client that initiated the request through the secure access network HS gRPC API channel.

[0036] (4) Access to the main anonymous network: After the user client successfully receives and parses the authorization information issued by the management layer (this process corresponds to...) Figure 4 The interaction shown in the middle number [3] is completed), and the anonymous network container inside it can then use this information, such as Figure 4 As shown in number [4], an anonymous communication link to the anonymous network container of other VPS nodes is initialized and successfully established. After the communication result is returned from [5], the actual anonymous data transmission begins, as shown in number [6].

[0037] This meticulously designed access network mechanism, with its superior isolation and high controllability during the authentication process, significantly improves the overall anonymity of the resilient network system while effectively reducing potential security risks that the main anonymity network might expose due to directly handling connection requests from unknown sources. In particular, it effectively prevents unauthenticated malicious users from conducting unauthorized malicious probing and network mapping of the main network's core topology, thereby comprehensively enhancing the system's resistance to mapping and its survivability in complex adversarial environments. This method is particularly suitable for deploying layered and refined access control strategies in sensitive communication application scenarios that have extremely high requirements for the confidentiality of communication processes, the anonymity of user identities, and the control level of system access.

[0038] (v) Design of Resilient Resource Pool in Resilient Network Systems under Non-cooperative Environments The resilient resource pool refers to a collection of VPS nodes that possess anonymous communication capabilities, support dynamic role configuration, and have self-destruct capabilities. Nodes in this collection can be flexibly assigned different roles (e.g., configured as Relay nodes, DA directory authorization nodes, Bridge nodes, etc.) based on actual needs and network policies, and can be dynamically and automatically reconfigured based on real-time network situational awareness results.

[0039] likeFigure 5 As shown, the overall structure and operational status of the resilience resource pool are uniformly coordinated and scheduled by the management plane. Based on the current overall network load status, quality assessment of each communication link, link stability indicators, and overall security posture analysis results, the system performs intelligent grouping and role reassignment operations on the nodes in the resource pool. The anonymous communication subsystem deployed on each node adopts a pluggable modular architecture, allowing different nodes to select and load the most suitable anonymous communication protocol (such as one or more combinations of Tor, Snowflake, VPN, Wireguard, etc.) according to the policies issued by the management plane or their own network environment, thereby achieving protocol-level diversity and adaptability to dynamically changing network environments.

[0040] The resilient resource pool mechanism further incorporates robust health checks and node lifecycle management mechanisms. These mechanisms are closely integrated with the situational awareness module, continuously assessing the availability and potential risk status of each node in the resource pool. Once the system detects abnormal behavior in a node (e.g., continuously rejecting new connection requests, abnormal surges or drops in traffic, frequent self-destruction or restarts), the management layer will immediately take measures to temporarily remove the problematic node from the resource pool (e.g., for isolation and observation or repair) or permanently remove it (if the problem is severe or irreparable), thereby ensuring the continuous stability and security of the entire network operation.

[0041] Through the coordinated operation of the aforementioned mechanisms, the resilient resource pool, as a core foundational component for dynamic scheduling and anonymous communication assurance, provides highly flexible structural support for the primary anonymous network. This enables the network to exhibit rapid self-recovery and efficient route reconstruction capabilities when encountering adverse scenarios such as network structure damage, malicious attacks, or local node failures. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the resilient network system architecture in a non-cooperative environment according to the present invention.

[0043] Figure 2 This is a timing diagram of inter-layer communication in a resilient network system under non-cooperative conditions according to the present invention.

[0044] Figure 3 This is a flowchart illustrating the inter-layer communication process of the resilient network system in a non-cooperative environment according to the present invention.

[0045] Figure 4 This is a flowchart illustrating the access to a dedicated network and communication process in a resilient network system operating in a non-cooperative environment, as described in this invention.

[0046] Figure 5 This is a diagram illustrating the resilient resource pool mechanism and architecture in a resilient network system under non-cooperative environments, as described in this invention. Detailed Implementation

[0047] (a) Distributed test deployment Management Team: Utilizes a single local trusted server to manage the Management gRPC Server, situational awareness module, database, and front-end display. The management server should be deployed in a trusted data center or a controlled intranet environment.

[0048] VPS Cloud Node Layer: A minimum resilient network is built using 5 VPS cloud nodes, which can be deployed on different cloud service providers or availability zones to enhance topology resilience. Each VPS cloud node runs: a node management subsystem, including a gRPC client, HS gRPC support, and a Stem controller; an anonymous communication subsystem, including containerized Tor instances, with alternatives such as Snowflake or Wireguard modules; a local traffic detection module, including traffic statistics and port probing scripts; and a self-destruct module, including functions responsible for cleaning up critical keys and destroying sensitive configurations upon receiving a self-destruct command.

[0049] User-side: Use one VPS cloud node as a sample client for testing, integrating access container and anonymous container to initiate access authentication and establish anonymous links.

[0050] See architecture diagram Figure 1 .

[0051] (II) Initial Configuration Management preparation: Pre-configure the controlled node list and default configuration template on the management server; generate management-side key pairs and securely pre-install the public key in the node management subsystem of each VPS (for initial verification).

[0052] VPS Onboarding and Registration: Each VPS starts a node management service and initiates a registration request (based on gRPC) to the management layer. The registration request includes the node's hardware / software fingerprint, regional label, and currently available anonymous protocol modules. The management layer registers the VPS and returns the node ID and a list of role candidates.

[0053] DA Node Designation: The management layer selects one or more nodes as Directory Authorization (DA) nodes based on node health and network topology, and issues DA fingerprints and torrc templates via HS gRPC (see Appendix for example torrc snippets).

[0054] Hidden Service (HS gRPC) Setup: The VPS creates a Hidden Service using a local Tor instance and registers its hidden domain name (onion address) and fingerprint with the management layer via HS gRPC. Subsequent sensitive communications will use this channel.

[0055] (III) Operation Phase Persistent connections and heartbeats: Each VPS maintains a persistent connection with the management layer via bidirectional streaming gRPC, periodically reporting heartbeats, traffic statistics, and port connectivity (e.g., adjustable every 10 seconds).

[0056] Situational awareness: The management-level situational awareness module aggregates all metrics reported by nodes (such as CPU / memory load, inbound / outbound traffic rates, port anomaly rate, connection failure rate, etc.) and calculates node health scores according to preset strategies.

[0057] Configuration distribution and dynamic updates: When the management layer determines that a configuration change is needed (e.g., adding a route, replacing a DA, or adjusting a rate limiting policy), the update is distributed to the target node via basic gRPC or HS gRPC and the new torrc configuration is dynamically loaded via the Stem interface.

[0058] Command and Execution Result Feedback: The management plane receives commands from the front end, converts them into Protobuf format commands through the gRPC Server, and sends them to one or more specified VPS cloud nodes. The gRPC clients on these nodes parse the commands and execute them locally, such as modifying node roles, obtaining node information, establishing internal circuits, etc. After execution, the execution results are sent back to the management plane for further processing, such as saving to the database or generating alarms.

[0059] (iv) Abnormal Handling (1) After receiving an anomaly report, the situation awareness module triggers the rule: if the port anomaly rate is greater than the threshold and the connection failure rate exceeds the corresponding threshold within a short period of time (within 60 seconds), it is marked as "high risk".

[0060] (2) The management system sends an isolation command to the node via gRPC (suspend receiving new connections and stop accepting forwarding tasks) and simultaneously triggers the local self-destruct procedure (erase the local DA key and clear the sensitive routing table).

[0061] (3) The management team selects a healthy node from the candidate node set as a backup, executes the same role torrc template distribution, updates the DA information record, and broadcasts fingerprint information to add the new node to the network and complete the role migration.

[0062] (4) If the node is a DA node, after the DA node is taken offline, the DA migration initiated by the management layer includes: allocating a new DA, issuing the fingerprint of the new DA via HS gRPC, loading the torrc on the new DA using Stem, and broadcasting the new DA list to the entire network. After receiving the new DA list, the client updates its ingress configuration and continues to establish anonymous links.

[0063] (v) Test methods and expected results To verify the resilience and anonymity continuity of this invention, the following tests can be performed: Functional correctness test: Complete node registration and DA assignment according to the above steps, and verify that the client can successfully access the anonymous network and access the target service.

[0064] Failure Recovery Test: During operation, a controlled shutdown (simulating blocking or intrusion) is performed on a Relay or DA node. Observe whether the management layer can complete the isolation and DA migration under the preset policy and ensure client link recovery. Expectation: After the DA shutdown is triggered, the system should complete the new DA online and client reconnection logic within a "reasonable time window" (the example gives a range of 30-120 seconds as a reference; the specific value can be optimized based on deployment scale and network conditions), ensuring the availability of the primary anonymous network.

[0065] Security processing verification: After the self-destruct process is triggered, check whether the sensitive files / keys of the self-destructed node have been securely deleted and confirm that the node no longer provides sensitive services to the outside world.

[0066] (vi) Results are shown in Table 1. Table 1. Overview of Test Results .

[0067] (1) Functional correctness test Test steps: Follow the “Initial Configuration Process” in the document to complete the management layer preparation, VPS registration, DA specification, and HS gRPC establishment. Then, the client initiates authentication and accesses the sample target service according to the access process.

[0068] Observation results: All 200 concurrent clients completed registration and access; the average initial connection establishment latency was approximately 1.2 seconds (a composite value of network round trip latency and Tor connection establishment latency); no abnormal failures were recorded.

[0069] Conclusion: Under normal operating conditions, this solution can reliably complete node registration, DA assignment, and client access, proving that the system's functional flow is correct.

[0070] (2) Failure recovery test Test Scenario 1: Perform a "controlled shutdown" on a VPS currently acting as a Data Controller (DA). Simulated reasons include being blocked by the ISP or being proactively shut down due to node intrusion. Monitor the management layer's situational awareness triggering, isolation, self-destruct issuance, and new DA assignment process, and record the migration completion time.

[0071] Number of trials: 10.

[0072] Table 2 Failure Recovery DA Migration Test .

[0073] Experimental results: The average migration time was 39.55 seconds. In most cases, the system was able to complete the DA switch in a short time.

[0074] Test Scenario 2: Perform "controlled disconnection" on a VPS that a client is currently connected to, causing one or more links that the client is connected to to disconnect, and record the client's reconnection status.

[0075] Number of trials: 50.

[0076] Test results: 48 / 50 of the clients were able to reconnect to the link on their own within the specified time (120s), thus ensuring that the function could operate normally.

[0077] (3) Safety processing verification Test procedure: A "self-destruct" command is issued to the target node, triggering the node to execute a self-destruct script (stopping processes, erasing hidden service directories, and destroying keys and sensitive configurations). The erasure results are then verified in the node's file system layer and management layer logs.

[0078] Example observation: All 12 files / keys marked as sensitive were removed from the node's file system and overwritten immediately after the self-destruct script was executed, and the management layer received a self-destruction completion confirmation event. Attempts to read 3 randomly selected filenames after self-destruction returned either "file does not exist" or "permission denied".

[0079] Conclusion: The self-destruct process can effectively remove sensitive files / keys declared on the node after receiving the instruction, reducing the risk of compromised nodes leaking critical information.

Claims

1. A resilient network system in a non-cooperative environment, characterized in that, The architecture consists of three layers: management layer, VPS cloud node layer, and user layer. The management layer is deployed on a local trusted server and serves as the control center of the system, undertaking identity authentication, node scheduling, directory authorization management, and situational awareness analysis functions. The cloud node layer consists of multiple VPSs, each of which deploys an anonymous communication subsystem and a node management subsystem, which are respectively responsible for Tor anonymous protocol communication and node status reporting and command execution. The user layer completes initial authentication and accesses the main network through the access private network mechanism. This three-tier architecture emphasizes clear division of labor and anti-interference capabilities within the system, including: The management layer, as the centralized control unit of the system, includes the following internal functional modules: a front-end interface for human-computer interaction and status display, which links with the back-end data to access configuration and log data; the core Management gRPCServer, which achieves comprehensive management and control through internal gRPC services, regular node gRPC services for VPS nodes, and hidden service gRPC based on the Tor service; in addition, it also includes a RESTful API interface for external integration, an internal client group for the management layer, and alarm daily and traffic sub-modules for situational awareness, and visualizes network status through node topology and links; The VPS cloud node layer consists of multiple geographically distributed VPSs, and each VPS node is equipped with a node management subsystem and an anonymous communication subsystem. The node management subsystem includes gRPC API and HS gRPC API interfaces, as well as node clients and Stem controllers, and is responsible for responding to management layer instructions and reporting node status. The anonymous communication subsystem is deployed in a containerized form, including access containers and anonymous containers. Each container integrates the Stem API for the Stem controller to call and runs a specific Onion instance and its related management components, including Stem node management, Stem link manager and Stem flow management. Meanwhile, the user space of the VPS node is equipped with a traffic detection and intrusion module, and uses eBPF hooks and system call interfaces in the kernel space to perform low-level data processing and monitoring, supplemented by a self-destruct mechanism to deal with serious security threats. Through the collaborative work of the above components, a high degree of resilience and communication security of the network system can be achieved in non-cooperative environments.

2. The resilient network system in a non-cooperative environment according to claim 1, characterized in that, A minimum functional onion network is built by designating some nodes as directory authorization (DA) nodes through the management layer; the management layer uses the gRPC protocol as a communication bridge to broadcast DA information to all relay nodes, realizing automatic expansion and contraction of the network topology; in addition, to ensure the anonymity and anti-tracking capability of the system, all sensitive communications are completed based on the onion service gRPC. In addition, by setting up a situational awareness module, the system can acquire real-time node behavior status and external traffic characteristics, including node load, inflow / outflow rate, port anomalies, and connection request failure rate. When an attack occurs or an anomaly is detected, a self-destruct mechanism is triggered, causing the node to be immediately removed from the system and local sensitive data to be cleared.

3. The resilient network system in a non-cooperative environment according to claim 1, characterized in that, This includes inter-layer communication design, specifically: The gRPC protocol used is divided into two types of interfaces: basic gRPC and Hidden Service gRPC (HS gRPC) based on the Onion domain. Basic gRPC is used for basic control communication such as node initialization, state synchronization, and configuration distribution. HS gRPC is used for the exchange of sensitive information, including DA fingerprint distribution, situational awareness data uploading, user authentication, and distribution of recommended routing information. VPS nodes maintain a continuous connection with the management layer through bidirectional streaming gRPC, forming a blocking command flow channel, which enables the management layer to control node behavior in real time. The entire communication process runs asynchronously and is scheduled based on a coroutine concurrency framework to improve system response efficiency in scenarios with massive VPS node deployments. VPS nodes maintain a continuous connection with the management layer via bidirectional streaming gRPC, forming a blocking command flow channel. This allows the management layer to push control commands to the nodes in real time, including role setting, configuration changes, or execution of destruction operations. The entire communication process is designed to run asynchronously and is scheduled and managed based on a coroutine concurrency framework, improving system response efficiency and concurrent processing capabilities in scenarios with massive VPS node deployments. The basic gRPC interface is mainly used for routine interactive activities during system initialization, including node registration, DA fingerprint reporting, receiving and executing role assignment commands, submitting health check reports, and issuing and confirming self-destruct commands. In contrast, the Hidden Service gRPC interface uses the Onion domain name of the Onion service to build a secure communication channel, which is specifically used to transmit link information, situational awareness data, user authentication credentials, DA update information, and highly sensitive data such as route recommendations.

4. The resilient network system in a non-cooperative environment according to claim 3, characterized in that, This includes implementing onion network management based on Stem, specifically: The Stem library, written in Python, is used as the tool for managing and monitoring the Onion network. This tool allows for dynamic remote control of Onion instances deployed on various VPS nodes. System administrators can use the interfaces provided by Stem to start, stop, reload, or reconfigure the Onion instance processes, thereby supporting role switching, dynamic registration and deregistration of hidden services, verification of communication path validity, and on-demand reconstruction of network topology throughout the node's lifecycle. The torrc configuration file for each node is automatically generated by the node management subsystem on that node based on the latest policies issued by the management plane; the generated configuration is then dynamically loaded and updated during the runtime of the Onion instance via the Stem interface; Directory Authorization (DA) nodes act as DA servers in the network, and their operational status is monitored in real time by the management layer. The management layer obtains detailed current status information of DA nodes through the Stem library, including service fingerprints, port connectivity test results, and network consensus file generation progress indicators. When the management detects that a DA node has unexpectedly disconnected from the network, exhibits abnormal behavior, or has been identified as compromised, the management will issue a command to the abnormal node via the gRPC interface to instruct it to execute a self-destruct procedure. At the same time, the management will select a new, healthy node to complete the migration of the DA role according to a preset strategy or real-time evaluation results. The entire role migration process is completed in real time and with precise control through the Stem library, thereby ensuring the resilience of the Onion Network's core structure and the continuous availability of anonymous communication links in the face of node failures or attacks.

5. The resilient network system in a non-cooperative environment according to claim 1, characterized in that, This includes user verification, specifically: Provide effective identity isolation and network buffering mechanisms before users access the system, i.e., adopt an access network mechanism; this mechanism includes building a small dedicated access network independent of the main onion network, which consists of a group of controlled relay nodes whose function is to handle the verification of user authentication credentials and the issuance of permissions. Users who successfully pass the access network authentication will obtain the DA information of the main network, and thus be able to access the true anonymous communication network; This dedicated access network consists of a group of strictly controlled VPS nodes whose specific functions are to handle access authentication requests from new users, perform fine-grained permission configuration, and distribute initial network information to users after successful authentication. The anonymous communication subsystem deployed within the access network nodes is also built on the Onion routing protocol or the Snowflake anonymous protocol, but its use is strictly limited to the authentication process, and it ensures that these nodes do not participate in the transmission of Relay or Bridge data in the main anonymous network.

6. The resilient network system in a non-cooperative environment according to claim 5, characterized in that, When a user requests authentication, the user first establishes a connection with the management plane through the access network. The user client integrates a node management module, an access network container, and an anonymous network container. The specific access process is as follows: (1) Authentication request initiation: The user client submits an authentication request and related credentials to the management layer through the hidden service gRPC interface provided by the access network container; each access network container is configured to run a dedicated Onion instance to handle such requests; (2) Management layer verification and authorization: After receiving the user's authentication request, the gRPC server of the management layer will verify the user's legitimacy based on the user policy preset in the system and the user credential information stored in the database; (3) Authorization information distribution: Once the user authentication is successful, the management layer will return the current valid DA information of the main anonymous network, a list of recommended available relay nodes, and the access policy specially configured for the user account to the user client that initiated the request through the access network HS gRPC API channel. (4) Accessing the main anonymous network: After the user client successfully receives and parses the authorization information issued by the management layer, its internal anonymous network container uses this information to initialize and successfully establish an anonymous communication link to the anonymous network containers of other VPS nodes, thereby starting real anonymous data transmission.

7. The resilient network system in a non-cooperative environment according to claim 1, characterized in that, It also includes a resilient resource pool design, which refers to a collection of VPS nodes with anonymous communication capabilities, dynamic role configuration capabilities, and self-destruct capabilities. The nodes in this collection are flexibly assigned different roles according to actual needs and network policies, including being configured as Relay nodes, DA directory authorization nodes, and Bridge nodes. Furthermore, they are dynamically and automatically reconfigured based on real-time network situation awareness results. The overall structure and operational status of the resilience resource pool are uniformly coordinated and scheduled by the management layer. Specifically, based on the current overall network load status, quality assessment of each communication link, link stability indicators, and overall security situation analysis results, intelligent grouping and role reassignment operations are performed on the nodes in the resource pool. The anonymous communication subsystem deployed on each node adopts a pluggable modular architecture, which allows different nodes to select and load the most suitable anonymous communication protocol according to the policies issued by the management plane or their own network environment, thereby achieving diversity at the protocol level and adaptability to dynamically changing network environments.

8. The resilient network system in a non-cooperative environment according to claim 7, characterized in that, The resilient resource pool further incorporates health checks and node lifecycle management mechanisms. These mechanisms are closely integrated with the situational awareness module to continuously assess the availability and potential risk status of each node in the resource pool. Once the system detects abnormal behavior in a node, the management layer immediately takes measures to temporarily remove or permanently remove the problematic node from the resource pool, thereby ensuring the continuous stability and security of the entire network operation.