A distributed AI security testing method and system based on a routing agent
Patent Information
- Application Number
- CN202610858904.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-15
Smart Images

Figure CN122764585A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence security technology, and more specifically, to a distributed AI security testing method and system based on routing proxies. Background Technology
[0002] With the widespread application of artificial intelligence models in fields such as image recognition, natural language processing, and multimodal analysis, adversarial attacks against AI models are becoming increasingly diverse, highlighting the growing importance of model security testing. Current AI security testing often employs single-point testing architectures or centralized testing platforms, where test tasks are executed serially on the same node or in the same environment. This results in low testing efficiency and limited coverage when facing large-scale models and diverse attack methods.
[0003] In distributed testing scenarios, existing task scheduling technologies generally allocate tasks based on computing power load or node resource utilization, failing to consider the attack surface coverage requirements unique to AI security testing. This leads to the same attack type being repeatedly assigned to the same nodes, while some attack dimensions remain untested for extended periods, resulting in incomplete attack surface coverage and the potential to miss critical security vulnerabilities. Furthermore, existing adversarial sample verification is mostly conducted in a single model architecture or single attack environment, lacking a mechanism for systematically migrating and verifying adversarial samples into heterogeneous model runtime environments. This makes it difficult to discover general threat samples with robustness across model architectures and attack strategies.
[0004] When adversarial examples are transferred between distributed test nodes, existing data management does not isolate and clean up the output information of the intermediate layers of the model. This results in sensitive information such as gradient tensors and feature maps generated during white-box attacks being transmitted with the samples, posing a risk of model privacy leakage. At the same time, each test node uses a heterogeneous model inference framework and a complementary attack toolset. The current task sharding granularity is relatively coarse, making it difficult to achieve fine-grained matching of attack types, sample features, and perturbation constraints, thus restricting the full utilization of the complementary advantages of heterogeneous nodes.
[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0006] The purpose of this application is to provide a distributed AI security testing method and system based on routing agents, which has the advantages of achieving comprehensive attack surface coverage and heterogeneous node collaborative verification, thereby improving the comprehensiveness of AI model security testing and the threat detection rate.
[0007] Firstly, this application provides a distributed AI security testing method based on a routing proxy, the technical solution of which is as follows: Receive a test task message, which includes the identifier of the model to be tested, the test sample set, attack parameters, and the coverage threshold. Based on the attack type, sample characteristics, and perturbation constraint fragmentation, multiple attack subtasks are generated, each attack subtask carrying an attack surface identifier and perturbation constraint parameters. Query the node capability mapping table and attack surface coverage record table to obtain the attack capability information of each test node and the number of tests for each attack surface identifier; The current coverage is obtained based on the number of tests corresponding to the attack surface identifier. Based on the current coverage, the coverage threshold, and the attack capability information, each attack subtask is distributed to the target test node. Receive adversarial samples returned by the target test node, which are generated by the target test node by executing the attack sub-task based on the test model identifier, the attack parameters, and the perturbation constraint parameters; input the adversarial samples generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculate the cross-attack surface attack success rate; A security test report is generated based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
[0008] Furthermore, the node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node; the attack surface coverage record table records the number of tests conducted on each attack surface identifier within a time window. The step of distributing each attack subtask to the target test node based on the current coverage, the coverage threshold, and the attack capability information includes: When the current coverage is less than the coverage threshold, the attack subtask is marked as a coverage task and distributed to the test node that matches the attack surface identifier; When the current coverage is greater than or equal to the coverage threshold, the attack subtask is marked as an enhancement task and distributed to the test node with the highest historical success rate.
[0009] Furthermore, when marking the attack subtask as an overlay task and distributing it to the test node matching the attack surface identifier, the following steps are also performed: Calculate the architectural differences between the model architecture deployed on each candidate test node and the test nodes that have covered the attack surface identifier; Candidate test nodes are sorted according to the architecture difference, and the attack subtask is distributed to the test node with the largest architecture difference; the architecture difference is calculated based on the differences in the number of layers, activation function types, and parameter sizes of the model architecture.
[0010] Furthermore, after distributing each attack subtask to the target test node, the process also includes: Update the number of tests corresponding to the attack surface identifier in the attack surface coverage record table; When the number of tests for all attack surface identifiers reaches the coverage threshold, the attack surface coverage record table is reset.
[0011] Furthermore, the sample features include the input data modality; when the test sample set is text data, it is divided into multiple text segments according to the text paragraph boundaries, and each text segment is used as a different attack subtask.
[0012] Furthermore, the calculation of the success rate of cross-attack surface attacks includes: The adversarial sample is input into at least two heterogeneous model runtime environments that are different from the first attack surface identifier; Compare the output results of each heterogeneous model runtime environment with the target labels, and record the successful attack indicators across the attack surface; Calculate the cross-attack surface attack success rate based on the cross-attack surface attack success identifier; When the success rate of the cross-attack surface attack is greater than a preset migration threshold, the adversarial sample is marked as a general threat sample.
[0013] Furthermore, the method also includes the following data management steps: Each test node maintains an independent adversarial sample cache area, and the generated adversarial samples are stored in partitions according to the attack surface identifier. When reading adversarial samples, verify the attack surface identifier of the adversarial sample against the allowed access identifier of the target test node; When the adversarial sample is marked as a general threat sample, the general threat sample is copied to the shared threat library, and the corresponding model intermediate layer output information in the adversarial sample cache is erased.
[0014] Furthermore, before inputting the adversarial sample generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, the process also includes: The adversarial sample generated by the first attack surface identifier is input into the test node corresponding to the third attack surface identifier, wherein the third attack surface identifier is different from the first attack surface identifier and also different from the second attack surface identifier. The test node corresponding to the third attack surface identifier performs secondary perturbation optimization on the adversarial sample to generate an enhanced adversarial sample; The enhanced adversarial sample is input into the heterogeneous model runtime environment corresponding to the second attack surface identifier to calculate the success rate of the secondary cross-attack surface attack.
[0015] Secondly, this application also proposes a distributed AI security testing system based on a routing proxy, comprising: The task interface unit receives a test task message, which includes the identifier of the model to be tested, the test sample set, attack parameters, and the coverage threshold. The attack dimension sharding unit shards according to the attack type, sample characteristics and perturbation constraints, generating multiple attack subtasks. Each attack subtask carries an attack surface identifier and perturbation constraint parameters. The routing proxy unit stores a node capability mapping table and an attack surface coverage record table. The node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node. The attack surface coverage record table records the number of tests performed on each attack surface identifier within a time window. The routing proxy unit queries the node capability mapping table and the attack surface coverage record table to obtain the attack capability information of each test node and the number of tests performed on each attack surface identifier. Based on the number of tests performed on the attack surface identifier, the current coverage is obtained. Based on the current coverage, the coverage threshold, and the attack capability information, each attack subtask is distributed to the target test node. A distributed test node cluster includes at least two test nodes, each of which is deployed with different model inference frameworks and attack toolsets. The target test node executes the attack sub-task and generates adversarial samples based on the test model identifier, the attack parameters, and the perturbation constraint parameters. The cross-attack surface cross-validation unit inputs the adversarial sample generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculates the cross-attack surface attack success rate. The report generation unit generates a security test report based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
[0016] Furthermore, the routing proxy unit is also used to mark the attack subtask as a coverage task and distribute it when the current coverage is less than the coverage threshold, and to mark the attack subtask as an enhancement task and distribute it to the test node with the highest historical success rate when the current coverage is greater than or equal to the coverage threshold. The cross-attack surface cross-validation unit is also used to mark the adversarial sample as a general threat sample when the cross-attack surface attack success rate is greater than a preset migration threshold. The system also includes an adversarial sample caching and isolation module, which is deployed on each test node. It stores the generated adversarial samples in partitions according to the attack surface identifier, and erases the corresponding model intermediate layer output information when the adversarial sample is marked as a general threat sample.
[0017] As can be seen from the above, the distributed AI security testing method and system based on routing proxy provided in this application solves the problems of incomplete attack surface coverage, low cross-environment threat detection rate, and easy leakage of model information in existing distributed AI security testing by constructing a node capability mapping table and an attack surface coverage record table, adopting a routing proxy mechanism based on current coverage and attack capability information, and a cross-attack surface heterogeneous verification and adversarial sample data isolation mechanism. It has the advantages of being able to achieve comprehensive attack surface coverage and heterogeneous node collaborative verification, and improving the comprehensiveness and threat detection rate of AI model security testing. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the steps of the distributed AI security testing method based on routing proxy disclosed in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the steps for calculating the success rate of cross-attack surface attacks as disclosed in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the relay enhancement steps disclosed in an embodiment of the present invention; Figure 4 This is a flowchart illustrating the data management steps disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of the distributed AI security testing system based on routing proxy disclosed in an embodiment of the present invention. Detailed Implementation
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments belong; the terminology used herein and in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit these embodiments; the terms "comprising" and "having," and any variations thereof, in the specification of these embodiments and the foregoing drawings, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification of these embodiments and the foregoing drawings are used to distinguish different objects, not to describe a particular order.
[0021] The implementation details of the technical solution in this embodiment are described in detail below: This application proposes a distributed AI security testing method based on routing proxies, such as... Figure 1As shown, the method includes: S101, Receive a test task message, the test task message includes the identifier of the model to be tested, the test sample set, attack parameters and coverage threshold.
[0022] Specifically, the system receives test task messages, obtaining basic description information of the distributed security test task from the upstream test management platform through the task interface unit. This message is transmitted in a structured data format, where the test model identifier uniquely identifies the AI model being tested; for example, the identifier field is denoted as `modelid=ResNet50ImageNet`, and the system uses this to retrieve the corresponding model file and weight parameters from the model repository. The test sample set is the set of input data used in this security test; for example, the sample set field is denoted as `samplesetid=VAL1K`, containing 1000 224×224 pixel images and their corresponding ground truth labels. Attack parameters describe the execution conditions of adversarial attacks; for example, the perturbation budget `epsilon` is set to 8 / 255, the attack iteration count `iterations` is set to 20, and the attack type is a PGD white-box attack. The coverage threshold sets the initial judgment benchmark for the attack surface coverage record table; for example, a threshold `covthreshold=3` means that each attack surface identifier must complete at least 3 tests in a single test round.
[0023] The test model identifier represents the unique identity of the tested model. Based on this identifier, the system loads the corresponding network architecture file and pre-trained weights from the model repository of the distributed test node cluster. The test sample set serves as the input basis for adversarial attacks, and its data size and distribution characteristics directly affect the granularity of the attack subtasks in subsequent steps S102. The perturbation budget epsilon in the attack parameters constrains the maximum modification range of the adversarial sample relative to the original sample, and the iterations limit the optimization depth of a single-step attack. Together, they determine the balance between attack strength and computational cost. The coverage threshold covthreshold provides the initial criterion for task distribution for the routing agent unit. When the current test count of an attack surface identifier is lower than this threshold, the routing agent unit marks it as a coverage task and schedules it to a matching node; otherwise, it marks it as an enhancement task and schedules it to the node with the highest historical success rate.
[0024] In practical applications, after receiving the message, the system first verifies the integrity of the fields. For example, it checks whether the identifier of the model under test is in the list of registered models, whether the data format of the test sample set matches the model input modality, whether the perturbation budget in the attack parameters exceeds the preset security range, and whether the coverage threshold is a positive integer. After the verification is successful, the system writes the message content to the task queue and triggers the subsequent processing flow of the S102 fragmentation unit.
[0025] The above-described scheme in this embodiment standardizes test task messages into structured data containing the four core fields mentioned above, enabling the input end of the distributed AI security testing system to have a unified task description specification. The test model identifier ensures the traceability of the test object, the test sample set provides the input basis for attacks, the attack parameters limit the boundary conditions of attack behavior, and the coverage threshold provides the initial scheduling judgment basis for the routing agent unit. It is precisely because these four fields play a key role in model location, data supply, attack control, and coverage judgment in the subsequent S102 sharding and S103 routing steps that the entire distributed security testing process can obtain clear and quantifiable execution instructions from the task reception stage, avoiding test interruptions or resource waste caused by missing or ambiguous input information.
[0026] S102 generates multiple attack subtasks based on attack type, sample characteristics, and perturbation constraint fragmentation. Each attack subtask carries an attack surface identifier and perturbation constraint parameters.
[0027] The sample features include the input data modality; when the test sample set is text data, it is divided into multiple text segments according to the text paragraph boundaries, and each text segment is used as a different attack subtask.
[0028] Specifically, based on attack type, sample features, and perturbation constraints, the system segments the test sample set received by S101 in three dimensions, forming a fine-grained set of attack subtasks. The attack type dimension includes three main categories: white-box attacks, black-box attacks, and physical world attacks. White-box attacks include gradient-based PGD attacks and optimization-based C&W attacks; black-box attacks include query-based boundary attacks and search-based genetic algorithm attacks; and physical world attacks include geometric transformation attacks and color transformation attacks. The sample feature dimension includes input data modalities, such as image modalities, text modalities, and audio modalities. The system automatically identifies the modal category based on the input interface type corresponding to the model under test in S101. The perturbation constraint dimension includes L2 norm constraints and L∞ norm constraints, corresponding to pixel-level overall perturbation and single-pixel maximum perturbation, respectively.
[0029] The attack surface identifier is generated by combining attack type encoding, sample feature encoding, and perturbation constraint encoding. For example, the attack surface identifier is denoted as AID=PGDIMGLinf, indicating that the subtask uses the PGD attack type, image modality sample features, and L∞ norm perturbation constraints. The perturbation constraint parameters include the perturbation budget and the number of iterations, for example, the perturbation budget epsilon=8 / 255 and the number of iterations=20. Each attack subtask is executed as an independent unit in the subsequent S103 routing step.
[0030] In some of the above implementations, when the test sample set is text data, the attack dimension sharding unit also performs the following paragraph boundary sharding operation: The test sample set is divided into multiple text segments based on the text paragraph boundaries, and each text segment is used as a different attack subtask.
[0031] Specifically, the test sample set is divided into multiple text segments based on text paragraph boundaries. That is, when the test sample set received by S101 is in text modality, the attack dimension segmentation unit divides the long text sample into multiple semantically independent text segments based on natural paragraph markers or line breaks. For example, a news text containing 5 natural paragraphs is divided into 5 text segments, each segment retaining the complete semantics of the original paragraph. Each text segment inherits the attack surface identifier and perturbation constraint parameters of the original attack subtask, but carries an independent segment number, for example, segment numbers segid=1 to segid=5.
[0032] In practical applications, for image modal samples, the attack dimension sharding unit directly shards the data according to the attack type and perturbation constraints. For example, 1000 image samples generate 1000 attack subtasks under PGD attack and L∞ constraints. For text modal samples, in addition to sharding according to attack type and perturbation constraints, paragraph boundary sharding is also required. For example, the above 5 news text segments generate 5 attack subtasks under PGD attack and character-level perturbation constraints. Each subtask independently performs character replacement or insertion operations on a text segment.
[0033] This application's solution segments attack subtasks according to three dimensions: attack type, sample characteristics, and perturbation constraints. This allows the distributed test node cluster to receive test units with controllable granularity. The uniqueness of the attack surface identifier ensures the traceability of each subtask in subsequent routing and verification phases, while the explicitness of the perturbation constraint parameters guarantees that different nodes follow a unified perturbation boundary when executing attacks. When the test sample set is text data, the paragraph boundary segmentation mechanism avoids semantic confusion and computational redundancy caused by long text-wide attacks, enabling each text segment to execute attack tests in parallel on different test nodes. It is precisely due to the synergistic effect of the three-dimensional segmentation mechanism and paragraph boundary segmentation that the subsequent S103 routing step can achieve accurate task distribution based on the attack surface identifier, and provides a fine-grained comparison basis for cross-attack surface verification.
[0034] This embodiment further proposes the above-mentioned query node capability mapping table and attack surface coverage record table to obtain the test node capability status and attack surface coverage status.
[0035] S103, query the node capability mapping table and attack surface coverage record table to obtain the attack capability information of each test node and the number of tests for each attack surface identifier.
[0036] The node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node; the attack surface coverage record table records the number of tests conducted on each attack surface identifier within a time window.
[0037] Specifically, the node capability mapping table is queried, meaning the routing proxy unit reads the capability registration information of each test node from its local storage. The node capability mapping table uses the node identifier as the primary key, and each record contains five fields: model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate. For example, in the record with node identifier NodeA, the model architecture field is ResNet50, the input modality field is image, the attack toolchain field is PGDCW, the perturbation norm type field is Linf, and the historical success rate field is 0.85. In the record with node identifier NodeB, the model architecture field is BERT, the input modality field is text, the attack toolchain field is TextFooler, the perturbation norm type field is Levenshtein, and the historical success rate field is 0.72. The combined record of these five fields constitutes the attack capability information of each test node, and the routing proxy unit uses this information to determine whether each node can undertake the attack subtask generated by S102.
[0038] Among them, the model architecture field characterizes the type of neural network structure deployed on the test node, such as convolutional networks, recurrent networks, or Transformer networks; the input modality field characterizes the raw data types that the test node can process, such as images, text, or audio; the attack toolchain field characterizes the set of attack algorithms integrated by the test node; the perturbation norm type field characterizes the perturbation measurement method supported by the test node; and the historical success rate field characterizes the success rate of the node in executing similar attack subtasks within the most recent time window. These fields together form a five-dimensional description of the node's capabilities, enabling the routing agent unit to filter target test nodes based on multi-dimensional matching conditions in step S104.
[0039] The attack surface coverage record table is queried, meaning the routing agent unit reads the test execution records for each attack surface identifier within the most recent time window. The attack surface coverage record table uses the attack surface identifier as the primary key, and each record contains two fields: the start time of the time window and the number of tests. For example, in the record corresponding to attack surface identifier AID=PGDIMGLinf, the start time of the time window is T0, and the number of tests is 2. In the record corresponding to attack surface identifier AID=CWTXTLev, the start time of the time window is T0, and the number of tests is 0. These test counts are directly used as the current coverage for each attack surface identifier. For example, a test count of 2 means that the attack surface has been executed 2 times within the current time window, and a test count of 0 means that the attack surface has not been executed yet.
[0040] In practical applications, after receiving the attack subtasks output by S102, the routing agent unit first extracts the attack surface identifier of each attack subtask, and then queries the node capability mapping table and the attack surface coverage record table in parallel. For example, for the attack surface identifier AID=PGDIMGLinf, the routing agent unit queries the node capability mapping table and finds that both NodeA and NodeC support this attack surface. It then queries the attack surface coverage record table and finds that the current test count for this attack surface is 2. If the coverage threshold covthreshold in S101 is 3, then the current coverage of 2 is less than the coverage threshold of 3, and this attack subtask is determined to be a coverage task. For the attack surface identifier AID=CWTXTLev, only NodeB supports this attack surface, and the current test count is 0, so it is also determined to be a coverage task.
[0041] This application's solution maintains two core data structures: a node capability mapping table and an attack surface coverage record table. This allows the routing agent unit to simultaneously grasp the capability boundaries of test nodes and the attack surface coverage status before task distribution. The five-dimensional information in the node capability mapping table ensures accurate matching between tasks and nodes, preventing image attack subtasks from being distributed to nodes that only support text modality. The number of time-window tests in the attack surface coverage record table provides a quantitative basis for coverage calculation, upgrading routing decisions from simple load balancing to strategy distribution based on attack surface coverage. Because these two tables provide dual information on node capability and coverage status in step S103, the subsequent step S104 can accurately mark attack subtasks as coverage tasks or enhancement tasks based on the comparison between the current coverage and the coverage threshold, and distribute them to target test nodes with execution capabilities.
[0042] S104. Obtain the current coverage based on the number of tests corresponding to the attack surface identifier. Distribute each attack subtask to the target test node based on the current coverage, the coverage threshold, and the attack capability information.
[0043] Specifically, the current coverage is obtained based on the number of tests corresponding to the attack surface identifier. That is, the routing proxy unit directly reads the number of tests queried in S103 as the current coverage value. For example, if the attack surface identifier AID=PGDIMGLinf has 2 tests, then the current coverage cov=2. Based on the current coverage, coverage threshold, and attack capability information, each attack subtask is distributed to the target test node. That is, the routing proxy unit executes the following binary routing logic: comparing the current coverage with the coverage threshold received in S101, and combining this with the attack capability information obtained in S103, determining the task marker and target node for each attack subtask.
[0044] Among them, the current coverage represents the frequency of tests that have been performed on the attack surface identifier within the time window, the coverage threshold represents the minimum coverage standard required by the test task, and the attack capability information represents the adaptability of each test node to execute the attack sub-task. The three together constitute the three-dimensional input of the distribution decision, enabling the routing agent unit to simultaneously consider the comprehensiveness of attack surface coverage and the execution efficiency of nodes.
[0045] Furthermore, in S104, the step of distributing each attack subtask to the target test node based on the current coverage, the coverage threshold, and the attack capability information includes: when the current coverage is less than the coverage threshold, marking the attack subtask as a coverage task and distributing it to the test node that matches the attack surface identifier.
[0046] Specifically, when the current coverage is less than the coverage threshold, the routing agent unit determines that the attack surface has not yet met the sufficient testing criteria and the testing coverage needs to be expanded. For example, if the current coverage cov=2 and the coverage threshold covthreshold=3, since 2 is less than 3, the routing agent unit marks the attack subtask as a coverage task. It then distributes the attack to a test node matching the attack surface identifier. Specifically, the routing agent unit queries the node capability mapping table in S103 and selects candidate test nodes whose attack toolchain and perturbation norm type are compatible with the attack surface identifier. For example, if the attack surface identifier AID=PGDIMGLinf requires image modality, Linf norm, and PGD toolchain, and NodeA and NodeC in the node capability mapping table both meet these conditions, then both are included in the candidate set.
[0047] Furthermore, when the current coverage is greater than or equal to the coverage threshold, the attack subtask is marked as an enhancement task and distributed to the test node with the highest historical success rate.
[0048] Specifically, when the current coverage is greater than or equal to the coverage threshold, the routing agent unit determines that the attack surface has reached the basic coverage standard and requires intensified testing. For example, if the current coverage cov=4 and the coverage threshold covthreshold=3, since 4≥3, the routing agent unit marks this attack subtask as an enhancement task. It is then distributed to the test node with the highest historical success rate. That is, after filtering compatible nodes from the node capability mapping table, the routing agent unit directly compares the historical success rate fields of each candidate node and selects the node with the largest value as the target test node. For example, if NodeA has a historical success rate of 0.85 and NodeC has a historical success rate of 0.78, then the enhancement task is distributed to NodeA.
[0049] In some of the above embodiments, when marking the attack subtask as a coverage task and distributing it to the test node that matches the attack surface identifier, the following steps are also performed: calculating the architecture difference degree between the model architecture deployed on each candidate test node and the test node that has covered the attack surface identifier; sorting the candidate test nodes according to the architecture difference degree, and distributing the attack subtask to the test node with the largest architecture difference degree; the architecture difference degree is calculated based on the difference in the number of layers, the difference in activation function type, and the difference in parameter size of the model architecture.
[0050] Specifically, the architecture difference is calculated. During the coverage task distribution phase, the routing agent unit not only requires candidate nodes to match the attack surface identifier, but also requires that the candidate node have the greatest difference in model architecture compared to nodes that have already executed that attack surface. For example, if the test node with the attack surface identifier AID=PGDIMGLinf is NodeA, and its deployed model architecture is ResNet50, and the candidate test node NodeC has a deployed model architecture of VGG16, then the architecture difference between NodeC and NodeA is calculated.
[0051] Candidate test nodes are sorted based on their architectural differences. Specifically, the routing agent unit ranks the candidate nodes from largest to smallest difference value and selects the first node as the target test node. For example, if the architectural difference value of candidate node NodeC is greater than that of NodeD, the coverage task will be distributed to NodeD.
[0052] The architectural difference is calculated based on the differences in the number of layers, activation function types, and parameter sizes of the model architectures. Specifically, the routing agent unit quantifies the degree of difference between two model architectures using the following formula: diff = w1 × |L1 - L2| + w2 × |P1 - P2| / Pmax + w3 × A Where, diff represents the architectural difference, L1 and L2 represent the number of layers in the two model architectures, |L1 - L2| represents the absolute value of the difference in the number of layers, P1 and P2 represent the parameter sizes in the two model architectures, |P1 - P2| represents the absolute value of the difference in parameter sizes, Pmax represents the maximum parameter size in the system, used for normalization, and A represents the difference in activation function type. When the two model architectures use the same activation function, A=0; when they use different activation functions, A=1. w1, w2, and w3 represent the weighting coefficients for the differences in the number of layers, parameter sizes, and activation function types, respectively. For example, w1=0.4, w2=0.4, and w3=0.2.
[0053] In practical applications, for example, Node A deploys ResNet50 with 50 layers (L1=50), a parameter size (P1=25.6M), and ReLU activation. Node C deploys VGG16 with 16 layers (L2=16), a parameter size (P2=138M), and ReLU activation. The maximum parameter size in the system is Pmax=138M. Therefore, the difference in layer count is |50-16|=34, and the difference in parameter size is |25.6-138|=112.4M. After normalization, 112.4 / 138≈0.815. Since the activation function is ReLU in both cases, A=0. Substituting these values into the formula, we get diff=0.4×34+0.4×0.815+0.2×0=13.6+0.326=13.926. If NodeD deploys DenseNet121 with 121 layers (L3), 8M parameters (P3), and ReLU activation, then the layer difference |50-121|=71, the parameter difference |25.6-8|=17.6M, and after normalization, 17.6 / 138≈0.128, A=0, and diff=0.4×71+0.4×0.128=28.4+0.051=28.451. Since 28.451>13.926, the routing agent unit will distribute the coverage task to NodeD instead of NodeC.
[0054] This application's solution introduces a binary comparison mechanism between current coverage and coverage threshold, enabling the routing agent unit to distinguish between two distribution modes: coverage tasks and enhancement tasks. During the coverage task phase, a strategy maximizing architectural differences ensures that adversarial samples are generated between model architectures with the greatest differences, thereby increasing the detection probability of cross-architecture migration attacks. During the enhancement task phase, a strategy maximizing historical success rates ensures that the most experienced nodes deepen testing on covered nodes. It is precisely this phased collaboration between coverage and enhancement tasks, along with the quantitative calculation of architectural differences, that allows the distributed test node cluster to achieve systematic security verification in a heterogeneous model environment while ensuring comprehensive attack surface coverage. This provides a highly transferable adversarial sample foundation for subsequent cross-attack surface cross-verification of S105.
[0055] Furthermore, after distributing each attack subtask to the target test node, the method further includes: updating the test count of the corresponding attack surface identifier in the attack surface coverage record table; and resetting the attack surface coverage record table when the test count of all attack surface identifiers reaches the coverage threshold.
[0056] Specifically, the test count for the corresponding attack surface identifier in the attack surface coverage record table is updated. That is, after the routing agent unit distributes the attack subtask to the target test node in step S104, it immediately increments the test count field of the attack surface identifier in the attack surface coverage record table. For example, if the attack surface identifier AID=PGDIMGLinf had a test count of 2 before distribution, the routing agent unit updates this record to 3 after distribution. This update operation is performed synchronously with the distribution action to ensure that the test count in the attack surface coverage record table matches the number of distributed tasks.
[0057] The real-time updating of the number of tests ensures that the calculation of current coverage in step S104 is always up-to-date. If the update operation lags behind the distribution action, subsequent attack subtasks may be incorrectly routed based on outdated coverage values, causing the same attack surface identifier to be repeatedly distributed to the same node within a short period of time, resulting in a waste of test resources.
[0058] When the number of tests for all attack surface identifiers reaches the preset coverage threshold, the routing proxy unit iterates through all records in the attack surface coverage record table after each update operation, checking whether the number of tests for each attack surface identifier is greater than or equal to the coverage threshold set in S101. For example, if the system currently has 3 attack surface identifiers, AID1, AID2, and AID3, with test counts of 3, 3, and 3 respectively, and the preset coverage threshold covthreshold=3, since 3≥3 and all conditions are met, the routing proxy unit determines that the current test round is complete and triggers the reset condition.
[0059] The attack surface coverage record table is reset. This means the routing agent unit clears the test count field for each attack surface identifier in the table to zero and updates the start time of the time window to the current system time. For example, after the reset, the test count for each attack surface identifier becomes 0, the start time of the time window is updated from T0 to T1, and the system enters a new round of attack surface coverage statistics.
[0060] In practical applications, update and reset operations constitute a closed-loop maintenance mechanism for the attack surface coverage record table. For example, in the first round, the routing agent unit distributes tasks sequentially to AID1, AID2, and AID3, and the number of tests for each attack surface identifier gradually accumulates. When AID1 reaches 3 tests, AID2 reaches 3 tests, and AID3 reaches 3 tests, the system performs a reset and starts the second round. If in a certain round AID1 has reached 3 tests while AID2 has only reached 2 tests, the system pauses the reset operation, continues to receive new tasks, and distributes them to the test nodes corresponding to AID2 until the number of tests for all attack surface identifiers reaches the coverage threshold.
[0061] Based on this, by updating the number of tests immediately after distribution and resetting the record table after all targets are met, the attack surface coverage record table can accurately reflect the test status within the current time window. The update operation ensures that the calculation basis for the current coverage in step S104 is valid in real time, while the reset operation avoids misjudgment of coverage caused by the accumulation of historical data, ensuring that the binary routing determination of coverage tasks and enhancement tasks is always based on the actual test progress of the current round.
[0062] S105, receive the adversarial sample returned by the target test node. The adversarial sample is generated by the target test node by executing the attack sub-task according to the test model identifier, the attack parameters and the perturbation constraint parameters. Input the adversarial sample generated by the first attack surface identifier into the heterogeneous model running environment corresponding to the second attack surface identifier, and calculate the cross-attack surface attack success rate.
[0063] Specifically, the cross-attack surface cross-validation unit receives the adversarial sample returned by the target test node, i.e., it obtains the attack execution result from the target test node distributed in step S104. For example, the target test node NodeA corresponding to the first attack surface identifier AID=PGDIMGLinf, based on the test model identifier ResNet50 received in S101, the attack parameters epsilon=8 / 255 and the number of iterations 20, and the perturbation constraint parameters generated in S102, performs a white-box PGD attack and generates an adversarial sample advsample1, and returns this sample and its attack success identifier to the cross-attack surface cross-validation unit. The adversarial sample generated by the first attack surface identifier is input into the heterogeneous model runtime environment corresponding to the second attack surface identifier, i.e., the cross-attack surface cross-validation unit inputs advsample1 into the test node runtime environment with a different architecture than NodeA. For example, the target test node NodeD corresponding to the second attack surface identifier AID=FGSMIMGL2 deploys a VGG16 architecture, whose input modality is also an image but whose perturbation norm is L2. This heterogeneous model runtime environment uses the same test model classification task to infer advsample1.
[0064] The cross-attack surface attack success rate characterizes the ability of an adversarial sample to perform attacks across different attack surface identifiers, corresponding model architectures, attack toolchains, and combinations of perturbation constraints. A higher success rate indicates greater robustness of the adversarial sample to differences in model architecture and defense strategies, and a higher potential threat level.
[0065] Furthermore, in some of the above embodiments, in S105, the calculation of the cross-attack surface attack success rate, such as... Figure 2 As shown, it includes the following steps: S2001, input the adversarial sample into at least two heterogeneous model runtime environments that are different from the first attack surface identifier; Specifically, the adversarial sample is input into at least two heterogeneous model runtime environments that differ from the first attack surface identifier. That is, the cross-attack surface cross-validation unit simultaneously inputs `advsample1` into NodeD corresponding to the second attack surface identifier and NodeE corresponding to the fourth attack surface identifier. NodeD deploys the VGG16 architecture and the FGSM attack toolchain with an L2 perturbation norm; NodeE deploys the MobileNet architecture and the PGD attack toolchain with an L2 perturbation norm. These two heterogeneous model runtime environments differ from NodeA corresponding to the first attack surface identifier in terms of model architecture, attack toolchain, and perturbation norm, forming a multi-dimensional heterogeneous validation environment.
[0066] S2002, compare the output results of each heterogeneous model's runtime environment with the target label, and record the successful attack markers across the attack surface; Specifically, the output results of each heterogeneous model's operating environment are compared with the target label. That is, the cross-attack surface cross-validation unit obtains the inference output categories of NodeD and NodeE for advsample1 and compares them with the corresponding real target labels in the test sample set in S101. If the output category does not match the target label, the attack in this heterogeneous environment is considered successful, and the cross-attack surface attack success flag is recorded as 1; if the output category matches the target label, the attack is considered unsuccessful, and the cross-attack surface attack success flag is recorded as 0. For example, if the output category of NodeD is an incorrect category, it is recorded as 1; the output category of NodeE is also an incorrect category, and it is recorded as 1.
[0067] S2003, Calculate the cross-attack surface attack success rate based on the cross-attack surface attack success identifier;
[0068] Among them, P cross This represents the success rate of cross-attack surface attacks, where N represents the total number of heterogeneous model runtime environments, and s i This represents the successful cross-attack surface attack record of the i-th heterogeneous model runtime environment. This represents the cumulative sum of attack success flags across all heterogeneous environments. For example, if N=2, s1=1, s2=1, then sum=2, P cross =2 / 2=1.0, meaning that the adversarial sample successfully attacked in all heterogeneous environments.
[0069] S2004, when the success rate of the cross-attack surface attack is greater than a preset migration threshold, the adversarial sample is marked as a general threat sample.
[0070] Specifically, when the success rate of cross-attack surface attacks exceeds a preset migration threshold, i.e., the cross-attack surface cross-validation unit will P crossThe sample is compared with the system's preset migration threshold Tmig. For example, if Tmig = 0.6, since 1.0 > 0.6, the cross-attack surface cross-validation unit marks advsample1 as a general threat sample. This marking indicates that the adversarial sample has strong migration capabilities across model architectures and attack strategies. It can not only successfully attack in the generated environment NodeA, but also maintain attack effectiveness in heterogeneous environments NodeD and NodeE, and belongs to a high-priority security threat sample.
[0071] In some of the above embodiments, before inputting the adversarial sample generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, such as Figure 3 As shown, the following relay enhancement steps are also included: S3001, input the adversarial sample generated by the first attack surface identifier into the test node corresponding to the third attack surface identifier, wherein the third attack surface identifier is different from the first attack surface identifier and different from the second attack surface identifier; Specifically, the adversarial sample generated by the first attack surface identifier is input into the test node corresponding to the third attack surface identifier. That is, before inputting advsample1 into NodeD, the cross-attack surface cross-validation unit first inputs advsample1 into NodeC corresponding to the third attack surface identifier. The third attack surface identifier AID=CWIMGLinf is different from the first attack surface identifier AID=PGDIMGLinf and the second attack surface identifier AID=FGSMIMGL2. NodeC deploys the DenseNet121 architecture and the CW attack toolchain, with a perturbation norm of Linf.
[0072] S3002, the test node corresponding to the third attack surface identifier performs secondary perturbation optimization on the adversarial sample to generate an enhanced adversarial sample; Specifically, the test node adversarial sample corresponding to the third attack surface identifier undergoes secondary perturbation optimization. That is, NodeC utilizes its deployed CW attack toolchain to further perform optimization-based perturbation search on top of advsample1. Because the CW attack employs a different loss function construction method and optimization iteration strategy, it can further compress the distance between the adversarial sample and the decision boundary based on the perturbation already generated by the PGD attack, generating the enhanced adversarial sample advsample1enhanced. This enhanced sample inherits the initial perturbation direction of advsample1 and simultaneously incorporates gradient optimization corrections unique to the CW attack.
[0073] S3003, input the enhanced adversarial sample into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculate the success rate of the secondary cross-attack surface attack.
[0074] Specifically, the enhanced adversarial sample is input into the heterogeneous model runtime environment corresponding to the second attack surface identifier. That is, the cross-attack surface cross-validation unit replaces `advsample1enhanced` with `advsample1` and inputs it into the heterogeneous model runtime environments of NodeD and NodeE. The secondary cross-attack surface attack success rate is calculated, i.e., following the same process from S2001 to S2003, the attack success identifiers of NodeD and NodeE on `advsample1enhanced` are recorded, and the secondary success rate `Pcrossenhanced` is calculated. For example, if both NodeD and NodeE output an error category for the enhanced sample, then `Pcrossenhanced` = 1.0. If `Pcrossenhanced` is greater than `Tmig`, then `advsample1enhanced` is also marked as a general threat sample, and its enhanced attributes generated by the PGD attack through secondary CW optimization are recorded. This secondary cross-attack surface attack success rate and the cross-attack surface attack success rate of the original adversarial sample in S105 are calculated independently and included together in the security test report.
[0075] In practical applications, the relay enhancement mechanism enables the system to leverage the different attack algorithm advantages of the test nodes corresponding to the third attack surface identifier to perform secondary optimization of the initial adversarial sample across attack strategies. Since the perturbation generation mechanisms of different attack algorithms are fundamentally different, the enhanced adversarial sample after secondary perturbation optimization often has stronger cross-environment migration capabilities than the original sample, thus giving the general threat sample library marked in step S2004 a higher threat coverage density.
[0076] Based on this, this embodiment achieves systematic migration and verification of adversarial samples across distributed heterogeneous nodes by inputting the adversarial samples generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier. The quantitative calculation of the success rate of cross-attack surface attacks provides an objective numerical basis for threat level determination, while the general threat sample label transforms the verification results into reusable security assets. The relay enhancement mechanism further explores the potential threat ceiling of the initial adversarial samples through secondary perturbation optimization of the third attack surface identifier.
[0077] S106 generates a security test report based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
[0078] Specifically, a security test report is generated based on the coverage, attack success rate, and cross-attack surface attack success rate of each attack surface identifier. That is, after completing cross-attack surface cross-validation in step S105, the report generation unit extracts three types of data from the attack surface coverage record table, the node capability mapping table, and the cross-attack surface cross-validation unit, respectively, and summarizes them to form a structured security test report. The report generation unit first reads the number of tests for each attack surface identifier in the attack surface coverage record table and compares it with the coverage threshold set in S101 to obtain the coverage of each attack surface identifier. For example, if the number of tests for attack surface identifier AID1 is 3 and the coverage threshold is 3, then the coverage is 100%; if the number of tests for AID2 is 2 and the coverage threshold is 3, then the coverage is 66.7%. Secondly, it reads the historical success rate of each test node under each attack surface identifier from the node capability mapping table as the attack success rate. For example, NodeA's historical success rate under AID1 is 0.85, and NodeB's historical success rate under AID2 is 0.72. Finally, the cross-attack surface attack success rate recorded by the cross-attack surface cross-validation unit is read. For example, the cross-attack surface attack success rate of the adversarial sample generated by AID1 in the AID2 environment is 0.9.
[0079] Among them, coverage represents the proportion of tasks completed under the attack surface identifier in the current test round; attack success rate represents the ability of the target test node to generate effective adversarial samples under the attack surface identifier; and cross-attack surface attack success rate represents the strength of the adversarial sample's ability to migrate attacks across heterogeneous environments. These three elements together constitute the three-dimensional data foundation of the security test report, enabling the report to simultaneously reflect the comprehensiveness of test coverage, the effectiveness of single-node attacks, and the migratory nature of threats across environments.
[0080] In practical applications, the report generation unit calculates the comprehensive security index using the following formula: Stotal = α × Rcov + β × Rsuc + γ × Rcross Wherein, Total represents the overall security index, Rcov represents the attack surface coverage, which is the arithmetic mean of the coverage of each attack surface identifier, Rsuc represents the average attack success rate, which is the arithmetic mean of the historical success rates of each test node under each attack surface identifier, and Rcross represents the average cross-attack surface attack success rate, which is the arithmetic mean of the success rates of all cross-attack surface verification pairs. α, β, and γ represent the weighting coefficients of coverage, attack success rate, and cross-attack surface success rate, respectively. For example, α is 0.3, β is 0.3, and γ is 0.4.
[0081] For example, if the system has two attack surface identifiers, AID1 and AID2, with AID1 having 100% coverage and AID2 having 66.7% coverage, then Rcov is 83.35%. Node A has a historical success rate of 0.85 under AID1, and Node B has a historical success rate of 0.72 under AID2, so Rsuc is 78.5%. The success rate of cross-attack surface attacks from AID1 to AID2 is 0.9, and the success rate of cross-attack surface attacks from AID2 to AID1 is 0.6, so Rcross is 75%. Substituting these values into the formula, Total equals 0.3 multiplied by 83.35 plus 0.3 multiplied by 78.5 plus 0.4 multiplied by 75, which is 25.0 plus 23.55 plus 30, resulting in 78.55.
[0082] The report generation unit organizes the above data into a security test report. The report includes an attack surface coverage detail table, an attack success rate statistics table for each test node, a cross-attack surface attack success rate matrix, a general threat sample list, and a comprehensive security index. The attack surface coverage detail table lists the number of tests, coverage thresholds, and coverage percentages for each attack surface identifier; the attack success rate statistics table lists the historical success rate and the actual success rate of each test node under each attack surface identifier; the cross-attack surface attack success rate matrix uses the attack surface identifier as the row and column index, recording the cross-attack surface attack success rate values between each pair of attack surface identifiers; the general threat sample list lists the general threat samples marked in step S105, along with their source attack surface identifier, target attack surface identifier, and cross-attack surface attack success rate.
[0083] Furthermore, in this embodiment, as Figure 4 As shown, the method further includes the following data management steps: S4001 maintains an independent adversarial sample cache area at each test node, and stores the generated adversarial samples in partitions according to the attack surface identifier. Specifically, each test node maintains an independent adversarial sample cache. This means the system deploys a separate storage area locally on each test node to temporarily store adversarial samples generated after that node performs its attack subtasks. For example, after NodeA completes the attack subtask on AID=PGDIMGLinf in step S105, it stores the generated adversarial sample advsample1 in the local cache partition identified as PGDIMGLinf. After NodeD completes the attack subtask on AID=FGSMIMGL2, it stores the generated adversarial sample advsample2 in the local cache partition identified as FGSMIMGL2. The caches on each test node are physically isolated and can only be read and written by the node itself.
[0084] Specifically, the cache is partitioned by attack surface identifier. This means that subdirectories or index tables are created within the cache based on attack surface identifiers, allowing all adversarial samples generated under the same attack surface identifier to be stored together. For example, in NodeA's cache, the AID=PGDIMGLinf partition stores all adversarial samples generated by that node under that attack surface, while the AID=CWIMGLinf partition stores samples from another attack surface. This partitioning mechanism enables subsequent read operations to quickly locate the target sample based on the attack surface identifier, avoiding a full cache traversal.
[0085] S4002, When reading the adversarial sample, verify the attack surface identifier of the adversarial sample against the allowed access identifier of the target test node; Specifically, when reading adversarial samples, the attack surface identifier is verified against the allowed access identifier of the target test node. That is, when a test node needs to read an adversarial sample from the buffer of another test node, the system first compares the attack surface identifier of the adversarial sample to be read with the allowed access identifier list of the target test node. For example, in step S105, if NodeD needs to read advsample1 from NodeA for cross-attack surface verification, the system extracts the attack surface identifier PGDIMGLinf of advsample1 and simultaneously queries NodeD's allowed access identifier list. If PGDIMGLinf is included in NodeD's allowed access identifier list, the verification passes and reading is allowed; otherwise, reading is rejected and an access exception is recorded.
[0086] The list of allowed access identifiers represents the set of adversarial sample attack surface identifiers that each test node is authorized to access. This list is configured synchronously by the routing agent unit when distributing attack subtasks in step S104. For example, if NodeD is configured to allow access to PGDIMGLinf and CWIMGLinf, but not to FGSMIMGL2, then NodeD can only read adversarial samples from the first two attack surface identifier partitions and cannot obtain sample data from other partitions.
[0087] In practical applications, verification mechanisms prevent unauthorized data flow between test nodes. For example, if a test node deploys a white-box attack toolchain, the adversarial samples it generates may carry gradient information from intermediate layers of the model. If this sample is directly transmitted to a black-box attack node, it could lead to the leakage of white-box model information. By verifying access permissions, the system ensures that only nodes with the corresponding permissions can read adversarial samples on a specific attack surface.
[0088] S4003, When the adversarial sample is marked as a general threat sample, the general threat sample is copied to the shared threat library, and the corresponding model intermediate layer output information in the adversarial sample cache is erased.
[0089] Specifically, when an adversarial sample is marked as a general threat sample—that is, after the system determines in step S105 via step S2004 that an adversarial sample is marked as a general threat sample—a data management operation is triggered. For example, if advsample1 generated by NodeA is marked as a general threat sample in S105, the system copies this sample from NodeA's local cache to the shared threat library. The shared threat library is a centralized storage area independent of each test node, used to store high-threat-level adversarial samples long-term for subsequent model hardening and defense training.
[0090] Specifically, the system erases the corresponding intermediate layer output information of the model in the adversarial sample cache. This means that after copying, the system cleans up the model inference process data associated with the general threat sample in the original test node cache. For example, during the generation of advsample1, NodeA called the ResNet50 model under test for forward and backward propagation, generating intermediate layer feature maps and gradient tensors as output information. This information is temporarily stored in the cache along with the adversarial sample for local verification. When advsample1 is marked as a general threat sample and copied to the shared threat library, the system deletes all intermediate layer feature maps and gradient tensors associated with this sample from the NodeA cache, retaining only the input data of the adversarial sample itself.
[0091] In practical applications, the output information of the model's intermediate layers contains the activation states of neurons and the directions of weight gradients, which are sensitive model information. If this information is stored locally on the test node along with general threat samples for an extended period, it may leak model structural details during subsequent node access or system maintenance. By copying the information to a shared threat library and erasing the local intermediate layer information, the system achieves centralized management of high-threat samples and local cleanup of sensitive model information, balancing the reusability of threat samples with the security protection of model privacy.
[0092] This embodiment utilizes a three-dimensional sharding mechanism based on attack type, sample characteristics, and perturbation constraints to decompose the test task into independently executable attack subtasks. Dynamic routing and distribution are then performed based on an attack surface coverage record table and a node capability mapping table, enabling the distributed test node cluster to balance comprehensive attack surface coverage with the execution efficiency of heterogeneous nodes. Through cross-attack surface cross-validation and relay enhancement mechanisms, adversarial samples are systematically migrated and tested using the architectural differences between heterogeneous model runtime environments, thereby discovering general threat samples with cross-environment robustness. Simultaneously, by leveraging the partitioning isolation, access verification, and intermediate layer information erasure mechanisms of the adversarial sample cache, the secure preservation and reuse of threat samples are achieved while effectively preventing the leakage of sensitive information within the model during the testing process. Finally, a security test report reflecting a three-dimensional view of the model's security status is generated through multi-dimensional quantitative summarization of coverage, attack success rate, and cross-attack surface attack success rate, providing an actionable priority basis for subsequent model hardening.
[0093] In summary, this embodiment forms a closed-loop testing system from task sharding, intelligent routing, cross-face verification to security accumulation. While ensuring comprehensive test coverage, it achieves a dual improvement in the collaborative efficiency of heterogeneous nodes and the privacy and security of models, providing a highly reliable, efficient, and secure technical solution for large-scale distributed AI security testing.
[0094] Secondly, this embodiment also proposes a distributed AI security testing system based on a routing proxy, such as... Figure 5 As shown, it includes a task interface unit, an attack dimension sharding unit, a routing proxy unit, a distributed test node cluster, a cross-attack surface cross-verification unit, and a report generation unit. Figure 5 As shown, the system includes: The task interface unit receives a test task message, which includes the identifier of the model to be tested, the test sample set, attack parameters, and the coverage threshold. The attack dimension sharding unit shards according to the attack type, sample characteristics and perturbation constraints, generating multiple attack subtasks. Each attack subtask carries an attack surface identifier and perturbation constraint parameters. The routing proxy unit stores a node capability mapping table and an attack surface coverage record table. The node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node. The attack surface coverage record table records the number of tests performed on each attack surface identifier within a time window. The routing proxy unit queries the node capability mapping table and the attack surface coverage record table to obtain the attack capability information of each test node and the number of tests performed on each attack surface identifier. Based on the number of tests performed on the attack surface identifier, the current coverage is obtained. Based on the current coverage, the coverage threshold, and the attack capability information, each attack subtask is distributed to the target test node. A distributed test node cluster includes at least two test nodes, each of which is deployed with different model inference frameworks and attack toolsets. The target test node executes the attack sub-task and generates adversarial samples based on the test model identifier, the attack parameters, and the perturbation constraint parameters. The cross-attack surface cross-validation unit inputs the adversarial sample generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculates the cross-attack surface attack success rate. The report generation unit generates a security test report based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
[0095] Specifically, in this embodiment, the task interface unit is deployed at the system front end and connects to the upstream test management platform through a standard API interface to receive test task messages. This unit has a built-in message parser that performs integrity checks on the message fields. After the check passes, it extracts the identifier of the model under test, the test sample set, attack parameters, and coverage threshold, and writes the parsing results into the task queue for downstream units to read.
[0096] The attack dimension sharding unit is deployed downstream of the task interface unit and connects to it via an internal bus. This unit loads the test sample set from the task queue and calls the sharding engine to decompose the task according to three dimensions: attack type, sample characteristics, and perturbation constraints. For example, for an image modality sample set, the sharding engine generates an attack subtask carrying AID=PGDIMGLinf; for a text modality sample set, when the sample is long text, the sharding engine divides the sample into multiple text segments based on text paragraph boundaries, generating attack subtasks carrying independent segment numbers. Each attack subtask is sent to the routing broker unit via a message middleware.
[0097] The routing proxy unit is deployed at the system control layer, connecting to the attack dimension sharding unit via a high-speed network and to the distributed test node cluster via the management network. This unit locally stores two core data structures: a node capability mapping table and an attack surface coverage record table. The node capability mapping table uses the node identifier as the primary key, recording the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node. The attack surface coverage record table uses the attack surface identifier as the primary key, recording the number of tests performed on each attack surface identifier within a time window. The routing proxy unit has a built-in query engine and a distribution engine. The query engine reads both tables in parallel to obtain attack capability information and the number of tests. The distribution engine calculates the current coverage based on the number of tests and performs binary routing based on the current coverage, coverage threshold, and attack capability information. For example, when the current coverage 2 is less than the coverage threshold 3, the distribution engine marks the attack subtask as a coverage task and distributes it to the test node matching the attack surface identifier; when the current coverage 4 is greater than or equal to the coverage threshold 3, the distribution engine marks the attack subtask as an enhancement task and distributes it to the test node with the highest historical success rate. During the task distribution phase, the distribution engine also calls the architecture difference calculation module to calculate the architecture difference between candidate nodes and covered nodes based on the differences in the number of layers, activation function types, and parameter sizes of the model architecture, and then distributes the task to the node with the greatest architecture difference.
[0098] The distributed test node cluster is deployed at the system execution layer and includes at least two physically isolated test nodes. For example, the first test node NodeA deploys the TensorFlow framework and the gradient-based PGD attack toolkit, the second test node NodeB deploys the PyTorch framework and the search-based TextFooler attack toolkit, and the third test node NodeC deploys the ONNX runtime and the optimized CW attack toolkit. Each test node is connected to the routing proxy unit through a computing network. After receiving the attack subtask, it loads the corresponding model from its local model repository according to the identifier of the model under test, executes the attack subtask according to the attack parameters and perturbation constraint parameters, generates adversarial samples, and returns them to the cross-attack surface cross-validation unit.
[0099] The cross-attack surface cross-validation unit is deployed in the system validation layer and connects to the distributed test node cluster via a data bus. This unit receives adversarial samples returned by the test nodes corresponding to the first attack surface identifier and inputs them into the heterogeneous model runtime environment corresponding to the second attack surface identifier. For example, the adversarial sample AID=PGDIMGLinf generated by NodeA is input into the VGG16 heterogeneous model runtime environment deployed by NodeD. The output result is compared with the target label, the cross-attack surface attack success identifier is recorded, and the cross-attack surface attack success rate is calculated. This unit also deploys a relay enhancement module. Before inputting the adversarial sample into the second attack surface identifier, the sample is first input into the test node corresponding to the third attack surface identifier to perform secondary perturbation optimization, generating enhanced adversarial samples before cross-attack surface validation. When the cross-attack surface attack success rate is greater than a preset migration threshold, this unit marks the adversarial sample as a general threat sample.
[0100] The report generation unit is deployed at the system output layer and connects to the routing proxy unit and the cross-attack surface cross-verification unit via an analysis interface. This unit reads coverage data from the attack surface coverage record table, attack success rate data from the node capability mapping table, and cross-attack surface attack success rate data recorded by the cross-attack surface cross-verification unit, and generates a security test report through weighted aggregation. The report includes detailed attack surface coverage, attack success rate statistics, a cross-attack surface attack success rate matrix, a general threat sample list, and a comprehensive security index.
[0101] Furthermore, the routing proxy unit is also used to mark the attack subtask as a coverage task and distribute it when the current coverage is less than the coverage threshold, and to mark the attack subtask as an enhancement task and distribute it to the test node with the highest historical success rate when the current coverage is greater than or equal to the coverage threshold; the cross-attack surface cross-validation unit is also used to mark the adversarial sample as a general threat sample when the cross-attack surface attack success rate is greater than a preset migration threshold; the system also includes an adversarial sample cache isolation module, which is deployed on each test node, stores the generated adversarial samples in partitions according to the attack surface identifier, and erases the corresponding model intermediate layer output information when the adversarial sample is marked as a general threat sample.
[0102] Specifically, when the routing agent unit performs distribution based on the current coverage, coverage threshold, and attack capability information, its built-in binary tagging module first performs a coverage comparison. For example, for an attack surface identifier AID=PGDIMGLinf, with a current coverage of 2 and a coverage threshold of 3, since 2 is less than 3, the binary tagging module marks this attack subtask as a coverage task. The goal of distributing coverage tasks is to expand the attack surface coverage. The routing agent unit filters candidate nodes that match the attack surface identifier from the node capability mapping table and distributes the task to the test node with the greatest architectural difference based on the architecture difference calculation results.
[0103] When the current coverage is greater than or equal to the coverage threshold, for example, if the current coverage of attack surface identifier AID=FGSMIMGL2 is 4 and the coverage threshold is 3, the binary labeling module marks the attack subtask as an enhancement task because 4 is greater than or equal to 3. The goal of distributing enhancement tasks is to deepen the test intensity. After filtering compatible nodes from the node capability mapping table, the routing agent unit directly reads the historical success rate field of each candidate node and selects the node with the largest value as the target test node. For example, if the historical success rate of candidate node NodeA is 0.85 and the historical success rate of NodeC is 0.78, then the routing agent unit will distribute the enhancement task to NodeA.
[0104] The binary labeling mechanism for coverage and enhancement tasks enables the routing agent unit to balance the comprehensiveness of attack surface coverage and the depth of test execution within the same test round. Coverage tasks ensure that insufficiently tested attack surfaces receive adequate node resources, while enhancement tasks further improve the attack quality on covered attack surfaces using the most experienced nodes.
[0105] The cross-attack surface cross-validation unit is also used to mark adversarial samples as general threat samples when the cross-attack surface attack success rate is greater than a preset migration threshold.
[0106] Specifically, in step S105, the cross-attack surface cross-validation unit calculates the cross-attack surface success rate Pcross and compares this value with the system's preset migration threshold Tmig. For example, for the adversarial sample advsample1, its cross-attack surface success rate Pcross in two heterogeneous model operating environments is 1.0, and the preset migration threshold Tmig is 0.6. Since 1.0 is greater than 0.6, the cross-attack surface cross-validation unit marks this adversarial sample as a general threat sample.
[0107] The general threat sample label indicates that the adversarial sample has strong transferability across model architectures and attack strategies, meaning it can not only successfully attack in the generated environment NodeA, but also maintain attack effectiveness in heterogeneous environments NodeD and NodeE. The labeled adversarial samples are added to the general threat sample list for subsequent model hardening and defense training.
[0108] The system also includes an adversarial sample caching and isolation module, which is deployed on each test node. It stores the generated adversarial samples in partitions according to the attack surface identifier, and erases the corresponding intermediate layer output information of the model when the adversarial sample is marked as a general threat sample.
[0109] Specifically, the adversarial sample caching isolation module is deployed locally on each test node as an independent process, maintaining an independent cache storage area for each test node. For example, the cache isolation module of NodeA creates a partition identified as PGDIMGLinf to store all adversarial samples generated by that node; the cache isolation module of NodeB creates a partition identified as CWTXTLev to store text attack samples. The partitions are logically isolated from each other and can only be accessed by the process on that node.
[0110] When an adversarial sample is marked as a general threat sample, for example, after advsample1 generated by NodeA is marked as a general threat sample by the cross-attack surface cross-validation unit, the cache isolation module performs two operations. First, it copies the sample from the local cache to the shared threat library, achieving centralized storage of high-threat samples. Second, it erases the intermediate layer output information of the model associated with the sample in the local cache, including the intermediate layer feature maps generated by forward propagation and the gradient tensors generated by backpropagation. After the erasure operation is completed, the local cache only retains the input data of the adversarial sample itself and no longer contains any information about the internal inference process of the model.
[0111] The intermediate layer outputs of the model contain neuron activation states and weight gradient directions, which are sensitive model information. The cache isolation module prevents cross-attack surface data mixing through partitioned storage and prevents sensitive information from being leaked due to the long-term retention of common threat samples through erasure operations.
[0112] In practical applications, the cache isolation module also includes a built-in access verification submodule. When a test node needs to read adversarial samples from other nodes, the access verification submodule compares the attack surface identifier of the sample with the target node's list of allowed access identifiers. Reading permission is granted only if the list contains the identifier. For example, when NodeD reads NodeA's advsample1, the access verification submodule checks NodeD's list of allowed access identifiers. If the list contains PGDIMGLinf, reading is allowed; otherwise, it is denied.
[0113] This system can be used to perform the distributed AI security testing method based on routing proxies described in the first aspect, which will not be elaborated further here.
[0114] This system uses a binary tagging mechanism in the routing agent unit to distinguish between coverage tasks and enhancement tasks and perform differentiated distribution; it uses a general threat sample tagging mechanism across the attack surface to enable highly mobile adversarial samples to obtain reusable threat asset identities; and it uses partitioned storage and information erasure in the adversarial sample cache isolation module to ensure that data flow in the distributed testing environment has a secure boundary.
[0115] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A distributed AI security testing method based on routing proxies, characterized in that, include: Receive a test task message, which includes the identifier of the model to be tested, the test sample set, attack parameters, and the coverage threshold. Based on the attack type, sample characteristics, and perturbation constraint fragmentation, multiple attack subtasks are generated, each attack subtask carrying an attack surface identifier and perturbation constraint parameters. Query the node capability mapping table and attack surface coverage record table to obtain the attack capability information of each test node and the number of tests for each attack surface identifier; The current coverage is obtained based on the number of tests corresponding to the attack surface identifier. Based on the current coverage, the coverage threshold, and the attack capability information, each attack subtask is distributed to the target test node. Receive adversarial samples returned by the target test node, which are generated by the target test node by executing the attack sub-task based on the test model identifier, the attack parameters, and the perturbation constraint parameters; input the adversarial samples generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculate the cross-attack surface attack success rate; A security test report is generated based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
2. The distributed AI security testing method based on routing proxy according to claim 1, characterized in that, The node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node; the attack surface coverage record table records the number of tests conducted on each attack surface identifier within a time window. The step of distributing each attack subtask to the target test node based on the current coverage, the coverage threshold, and the attack capability information includes: When the current coverage is less than the coverage threshold, the attack subtask is marked as a coverage task and distributed to the test node that matches the attack surface identifier; When the current coverage is greater than or equal to the coverage threshold, the attack subtask is marked as an enhancement task and distributed to the test node with the highest historical success rate.
3. The distributed AI security testing method based on routing proxy according to claim 2, characterized in that, When marking the attack subtask as an overlay task and distributing it to the test node matching the attack surface identifier, the following steps are also performed: Calculate the architectural differences between the model architecture deployed on each candidate test node and the test nodes that have covered the attack surface identifier; Candidate test nodes are sorted according to the architecture difference, and the attack subtask is distributed to the test node with the largest architecture difference; the architecture difference is calculated based on the differences in the number of layers, activation function types, and parameter sizes of the model architecture.
4. The distributed AI security testing method based on routing proxy according to claim 2, characterized in that, After distributing each attack subtask to the target test node, the process also includes: Update the number of tests corresponding to the attack surface identifier in the attack surface coverage record table; When the number of tests for all attack surface identifiers reaches the coverage threshold, the attack surface coverage record table is reset.
5. The distributed AI security testing method based on routing proxy according to claim 1, characterized in that, The sample features include the input data modality; when the test sample set is text data, it is divided into multiple text segments according to the text paragraph boundaries, and each text segment is used as a different attack subtask.
6. The distributed AI security testing method based on routing proxy according to claim 1, characterized in that, The calculation of the success rate of cross-attack surface attacks includes: The adversarial sample is input into at least two heterogeneous model runtime environments that are different from the first attack surface identifier; Compare the output results of each heterogeneous model runtime environment with the target labels, and record the successful attack indicators across the attack surface; Calculate the cross-attack surface attack success rate based on the cross-attack surface attack success identifier; When the success rate of the cross-attack surface attack is greater than a preset migration threshold, the adversarial sample is marked as a general threat sample.
7. The distributed AI security testing method based on routing proxy according to claim 6, characterized in that, The method also includes the following data management steps: Each test node maintains an independent adversarial sample cache area, and the generated adversarial samples are stored in partitions according to the attack surface identifier. When reading adversarial samples, verify the attack surface identifier of the adversarial sample against the allowed access identifier of the target test node; When the adversarial sample is marked as a general threat sample, the general threat sample is copied to the shared threat library, and the corresponding model intermediate layer output information in the adversarial sample cache is erased.
8. The distributed AI security testing method based on routing proxy according to claim 1, characterized in that, Before inputting the adversarial samples generated from the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, the following steps are also included: The adversarial sample generated by the first attack surface identifier is input into the test node corresponding to the third attack surface identifier, wherein the third attack surface identifier is different from the first attack surface identifier and also different from the second attack surface identifier. The test node corresponding to the third attack surface identifier performs secondary perturbation optimization on the adversarial sample to generate an enhanced adversarial sample; The enhanced adversarial sample is input into the heterogeneous model runtime environment corresponding to the second attack surface identifier to calculate the success rate of the secondary cross-attack surface attack.
9. A distributed AI security testing system based on a routing proxy, characterized in that, include: The task interface unit receives a test task message, which includes the identifier of the model to be tested, the test sample set, attack parameters, and the coverage threshold. The attack dimension sharding unit shards according to the attack type, sample characteristics and perturbation constraints, generating multiple attack subtasks. Each attack subtask carries an attack surface identifier and perturbation constraint parameters. The routing proxy unit stores a node capability mapping table and an attack surface coverage record table. The node capability mapping table records the model architecture, input modality, attack toolchain, perturbation norm type, and historical success rate supported by each test node. The attack surface coverage record table records the number of tests performed on each attack surface identifier within a time window. The routing proxy unit queries the node capability mapping table and the attack surface coverage record table to obtain the attack capability information of each test node and the number of tests performed on each attack surface identifier. Based on the number of tests performed on the attack surface identifier, the current coverage is obtained. Based on the current coverage, the coverage threshold, and the attack capability information, each attack subtask is distributed to the target test node. A distributed test node cluster includes at least two test nodes, each of which is deployed with different model inference frameworks and attack toolsets. The target test node executes the attack sub-task and generates adversarial samples based on the test model identifier, the attack parameters, and the perturbation constraint parameters. The cross-attack surface cross-validation unit inputs the adversarial sample generated by the first attack surface identifier into the heterogeneous model runtime environment corresponding to the second attack surface identifier, and calculates the cross-attack surface attack success rate. The report generation unit generates a security test report based on the coverage of each attack surface identifier, the attack success rate, and the cross-attack surface attack success rate.
10. The system according to claim 9, characterized in that, The routing proxy unit is also used to mark the attack subtask as a coverage task and distribute it when the current coverage is less than the coverage threshold, and to mark the attack subtask as an enhancement task and distribute it to the test node with the highest historical success rate when the current coverage is greater than or equal to the coverage threshold. The cross-attack surface cross-validation unit is also used to mark the adversarial sample as a general threat sample when the cross-attack surface attack success rate is greater than a preset migration threshold. The system also includes an adversarial sample caching and isolation module, which is deployed on each test node. It stores the generated adversarial samples in partitions according to the attack surface identifier, and erases the corresponding model intermediate layer output information when the adversarial sample is marked as a general threat sample.