Data detection method and apparatus, device, and storage medium
Patent Information
- Application Number
- CN202611386176.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-08
- Publication Date
- 2026-10-09
AI Technical Summary
[0004]本申请的主要目的在于提供一种数据检测方法、装置、设备及存储介质,旨在解决现有的依赖链路层指标的数据投毒检测方法的准确性不高的技术问题
[0015]本申请提供了一种数据检测方法,本申请公开了获取当前训练任务在RDMA传输过程中数据分片的传输层元数据,并根据传输层元数据确定数据分片的传输层风险评分,传输层风险评分用于表征数据分片在传输层面的行为偏离程度;提取数据分片的内容层安全特征,并根据内容层安全特征生成数据分片的内容层风险评分,内容层风险评分用于表征数据分片在内容层面的投毒风险程度;对传输层风险评分和内容层风险评分进行跨平面聚合,生成数据分片的跨平面关联推理指标,跨平面关联推理指标用于表征传输层风险评分与内容层风险评分之间的关联关系;基于跨平面关联推理指标对数据分片进行异常检测,获得异常检测结果;相较于现有的仅依赖链路层指标的检测方法,无法区分链路扰动与恶意投毒,也无法识别隐蔽投毒样本,由于本申请可以在获取传输层行为偏离程度的同时,进一步提取数据分片在内容层面的投毒风险程度,并通过跨平面聚合构建传输层风险评分与内容层风险评分之间的关联关系,再基于该关联关系进行异常检测,从而能够在检测中区分链路扰动与恶意攻击、识别隐蔽投毒样本,解决了现有的依赖链路层指标的数据投毒检测方法的准确性不高的技术问题。
Smart Images

Figure CN122885984A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network communication technology, and in particular to data detection methods, apparatus, devices and storage media. Background Technology
[0002] As the parameter scale of Large Language Models (LLMs) continues to grow, compute-in-memory (CIM) architectures are gradually becoming a crucial infrastructure supporting large-scale distributed model training. In this architecture, Remote Direct Memory Access (RDMA) technology, with its high bandwidth, low latency, and kernel bypass capabilities, is widely used in the data plane to enable direct writing of data from remote storage to the memory of compute nodes. However, in wide-area RDMA compute-in-memory training scenarios, storage nodes and compute nodes are often distributed across different geographical regions or even administrative domains, requiring training data to be transmitted via complex wide area network (WAN) links. WAN links inherently suffer from physical disturbances such as jitter, packet loss, and latency fluctuations. Furthermore, training data faces the risk of malicious tampering, injection, or replacement during transmission (data poisoning), posing a serious threat to the security and reliability of AI systems.
[0003] Currently, existing RDMA network anomaly detection methods mainly rely on link-layer metrics such as performance counters or completion queue entry (CQE) status. These metrics analyze retransmission counts, congestion events, and CQE error codes to pinpoint performance bottlenecks and link anomalies at the network or transport layer. However, while these methods can effectively identify transmission quality degradation caused by link failures, buffer overflows, or traffic conflicts, link-layer metrics only reflect anomalies at the transmission behavior level and cannot detect whether data content has been poisoned. Furthermore, WAN link noise and malicious attacks exhibit similar behavior at the link layer, making it impossible to distinguish between normal link disturbances and malicious poisoning. It also fails to identify covertly poisoned samples where transmission behavior appears normal but content has been tampered with, resulting in low accuracy in data poisoning detection. Summary of the Invention
[0004] The main objective of this application is to provide a data detection method, apparatus, device, and storage medium, which aims to solve the technical problem of low accuracy in existing data poisoning detection methods that rely on link layer indicators.
[0005] To achieve the above objectives, this application proposes a data detection method, the method comprising: Obtain transport layer metadata of data fragments during RDMA transmission for the current training task, and determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer. Extract the content layer security features of the data slices, and generate a content layer risk score for the data slices based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data slices at the content level. The transport layer risk score and the content layer risk score are aggregated across planes to generate a cross-plane association inference index for the data shards. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score. Anomaly detection is performed on the data shards based on the cross-plane association reasoning index to obtain anomaly detection results.
[0006] In one embodiment, the step of determining the transport layer risk score of the data fragment based on the transport layer metadata includes: Construct the transport layer feature vector corresponding to the data fragment based on the transport layer metadata; Based on the constraints in the RDMA transmission process, a directed dependency relationship is determined between several transport layer security feature nodes. Construct the transport layer semantic probability graph of the current training task based on each of the transport layer security feature nodes and the directed dependency relationships; Based on the set of historical transport layer feature vectors generated by the current training task in the trusted transport phase and the transport layer semantic probability graph, a transport layer trusted benchmark for the current training task is constructed. The transport layer trusted benchmark is used to characterize the conditional dependencies and normal value ranges between the security features of each transport layer under the trusted transport state. The transport layer risk score of the data fragment is determined based on the transport layer feature vector and the transport layer trust benchmark.
[0007] In one embodiment, the step of constructing a transport layer trust benchmark for the current training task based on the set of historical transport layer feature vectors generated during the trusted transport phase of the current training task and the transport layer semantic probability graph includes: Based on the set of historical transport layer feature vectors generated by the current training task in the trusted transmission phase, the conditional probability of each transport layer security feature node in the transport layer semantic probability graph under the corresponding parent node value combination is calculated. Based on the conditional probabilities, construct a conditional probability table for each of the transport layer security feature nodes; The transport layer reliability benchmark for the current training task is determined based on the conditional probability table.
[0008] In one embodiment, the step of determining the transport layer risk score of the data fragment based on the transport layer feature vector and the transport layer trust benchmark includes: The joint logarithmic matching value of the data fragment is determined based on the transport layer feature vector and the transport layer trust benchmark; Determine the average matching value and low matching boundary of historical reliable samples; The transport layer risk score of the data fragment is determined based on the joint logarithmic matching value, the average matching value, and the low matching boundary.
[0009] In one embodiment, the step of extracting the content layer security features of the data fragment and generating a content layer risk score for the data fragment based on the content layer security features includes: The data segments are decoded and segmented to identify decoded abnormal bytes in the data segments and map the decoded abnormal bytes to abnormal character identifiers. A word segmentation sequence is constructed based on the normal bytes and the abnormal character identifiers in the data segment, and the content layer security features of the data segment are extracted based on the word segmentation sequence. The content layer security features include word segmentation distribution skewness, proportion of long-tail rare word segmentation, proportion of abnormal characters, density features of abnormal repeated segments, and semantic distribution drift features. The content security summary value of the data segment is determined by a preset content security summary model based on the word segmentation distribution skewness, the proportion of long-tail rare word segmentation, the proportion of abnormal characters, the density feature of abnormal repeated segments, and the semantic distribution drift feature. The content layer risk score of the data fragment is determined based on the content security summary value using a preset risk assessment function.
[0010] In one embodiment, the step of cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane correlation inference index for the data fragment includes: The data sharding is generated based on the transport layer risk score and the content layer risk score. The joint increase index is used to characterize the degree to which the transport layer risk score and the content layer risk score increase synchronously, and the contradiction index is used to characterize the degree of deviation between the transport layer risk score and the content layer risk score. The transmission-side anomaly factor score of the data fragment is determined based on the transport layer metadata; Based on the transmission-side anomaly factor score and the content-layer risk score, the correlation risk indicators between each transmission-side anomaly factor and content anomaly are determined. The correlation risk indicators include identity correlation risk indicators, source correlation risk indicators, memory correlation risk indicators, and queue correlation risk indicators. The joint elevation index, the contradiction index, and the associated risk index are aggregated into the cross-plane association inference index.
[0011] In one embodiment, the step of performing anomaly detection on the data shards based on the cross-plane association inference index and obtaining anomaly detection results includes: Determine the maximum associated risk value among all associated risk indicators in the cross-plane association inference index; The indicator type corresponding to the maximum associated risk value is determined as the dominant associated type; The security status level of the data shard is determined based on the joint escalation index, the contradiction index, and the maximum associated risk value. The target correction strategy is determined from the candidate correction strategies based on the dominant correlation type, and the candidate correction strategies correspond to different transmission-side anomaly factor types. Anomaly detection results are generated based on the security status level and the target correction strategy.
[0012] Furthermore, to achieve the above objectives, this application also proposes a data detection device, the device comprising: The transport layer scoring module is used to obtain transport layer metadata of data fragments in the current training task during RDMA transmission, and to determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer. The content layer scoring module is used to extract the content layer security features of the data slice and generate a content layer risk score for the data slice based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data slice at the content level. The scoring aggregation module is used to perform cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane association inference index for the data shard. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score. An anomaly detection module is used to perform anomaly detection on the data shards based on the cross-plane association reasoning index and obtain anomaly detection results.
[0013] In addition, to achieve the above objectives, this application also proposes a data detection device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data detection method as described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the data detection method described above.
[0015] This application provides a data detection method. It discloses obtaining transport layer metadata of data fragments during RDMA transmission for the current training task, determining a transport layer risk score for each data fragment based on the transport layer metadata (the transport layer risk score characterizes the degree of behavioral deviation of the data fragment at the transport layer), extracting content layer security features of the data fragments, and generating a content layer risk score for each data fragment based on these features (the content layer risk score characterizes the degree of poisoning risk of the data fragment at the content layer), and performing cross-plane aggregation of the transport layer risk score and content layer risk score to generate a cross-plane correlation inference index for the data fragments (the cross-plane correlation inference index characterizes the correlation between the transport layer risk score and the content layer risk score). The application establishes a correlation between data fragments and performs anomaly detection based on cross-plane correlation inference indicators. Compared to existing detection methods that rely solely on link-layer indicators, which cannot distinguish between link disturbances and malicious poisoning, nor can they identify concealed poisoning samples, this application can further extract the poisoning risk level of data fragments at the content level while obtaining the degree of deviation of transport layer behavior. It also constructs a correlation between transport layer risk scores and content layer risk scores through cross-plane aggregation, and performs anomaly detection based on this correlation. This allows the application to distinguish between link disturbances and malicious attacks and identify concealed poisoning samples during detection, thus solving the technical problem of low accuracy in existing data poisoning detection methods that rely on link-layer indicators. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1This is a flowchart illustrating the data detection method of this application in Embodiment 1. Figure 2 The system architecture diagram provided for the data detection method of this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the data detection method of this application; Figure 4 This is a schematic diagram of the transmission risk scoring process in the data detection method of this application; Figure 5 This is a flowchart illustrating the content risk scoring process in the data detection method of this application; Figure 6 This is a flowchart illustrating Embodiment 3 of the data detection method of this application; Figure 7 This is a flowchart illustrating the joint inference process in the data detection method of this application; Figure 8 This is a schematic diagram of the module structure of the data detection device according to an embodiment of this application; Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the data detection method in the embodiments of this application.
[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solution of this application embodiment is as follows: Obtain the transport layer metadata of data fragments during RDMA transmission for the current training task, and determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of behavioral deviation of the data fragments at the transport layer. Extract the content layer security features of the data fragments, and generate the content layer risk score of the data fragments based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data fragments at the content layer. Perform cross-plane aggregation on the transport layer risk score and the content layer risk score to generate a cross-plane correlation inference index for the data fragments. The cross-plane correlation inference index is used to characterize the correlation between the transport layer risk score and the content layer risk score. Perform anomaly detection on the data fragments based on the cross-plane correlation inference index to obtain anomaly detection results.
[0023] Because existing methods rely solely on link-layer metrics for security assessment, these metrics can only reflect anomalies at the transmission behavior level and cannot detect whether data content has been poisoned. Furthermore, WAN link noise and malicious attacks exhibit similar behavior at the link layer, making it impossible to distinguish between normal link disturbances and malicious poisoning, or to identify covertly poisoned samples with normal transmission behavior but altered content.
[0024] This application provides a solution that can extract the degree of data fragmentation poisoning risk at the content level while obtaining the degree of deviation of transport layer behavior. It can also construct the correlation between transport layer risk score and content layer risk score through cross-plane aggregation, and then perform anomaly detection based on the correlation. This enables the detection to distinguish between link disturbances and malicious attacks, and identify hidden poisoning samples, thus solving the technical problem of low accuracy of existing data poisoning detection methods that rely on link layer indicators.
[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or data detection device capable of performing the above functions. The following description uses a data detection device (hereinafter referred to as the device) as an example to illustrate this embodiment and the subsequent embodiments.
[0026] Based on this, embodiments of this application provide a data detection method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the data detection method of this application in Embodiment 1.
[0027] In this embodiment, the data detection method includes steps S10 to S40: Step S10: Obtain the transport layer metadata of the data fragments in the current training task during RDMA transmission, and determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer.
[0028] Understandably, the current training task can be the training task of a large-scale deep learning model that is currently being executed. Large-scale deep learning models, such as large language models, typically require processing massive amounts of training data, which may be distributed across multiple remote storage nodes. In a compute-storage separation architecture, compute nodes (such as GPU server clusters) are physically separated from storage nodes. Compute nodes can obtain training data from remote storage nodes on demand through high-performance networks. In practical applications, the current training task represents the model training job that the compute node is currently executing. Each training task has a unique task identifier used to distinguish data flows between different training jobs.
[0029] It is also understood that the RDMA transfer process can be a process of data transfer between a remote storage node and a local computing node via Remote Direct Memory Access (RDMA). RDMA is a high-speed data transfer technology that allows remote devices to bypass the local central processing unit and operating system kernel and directly write data to local memory. In this embodiment, the RDMA transfer process can specifically refer to the process of transferring training data from a remote storage node to a computing node in a wide area network environment.
[0030] It should be understood that a data fragment can be the smallest unit of transmission into which training data is divided during RDMA transmission. In practical applications, since the amount of data transmitted in a single RDMA transmission is usually limited by the network card buffer and transmission protocol, large training data (such as a complete training sample or a batch of training samples) needs to be divided into multiple data fragments at the sending end, transmitted separately via RDMA, and then reassembled into complete data at the receiving end according to the logical position of the fragments. For example, a training text sample of several megabytes in size will be divided into multiple data fragments of several kilobytes to tens of kilobytes during RDMA transmission and sent sequentially. In this embodiment, the device can use data fragmentation as the basic granularity of processing, and detect each data fragment when it arrives at the receiving end memory.
[0031] It should also be understood that transport layer metadata can be descriptive data generated by the transport layer during RDMA transmission that is related to data fragmentation transmission behavior. Specifically, it can include connection identity data, memory access data, timing traffic data, queue event data, and source authentication data.
[0032] The connection identity data describes the entity establishing the RDMA transmission connection, the connection object, and the operational behavior. It includes the source node identifier, queue pair number and status, and RDMA operation type. The source node identifier is the identifier of the node initiating the RDMA transmission, which can be used to determine whether the data comes from an authorized source. The queue pair number and status are the RDMA communication queue pairs and their running status, which can be used to determine whether the connection belongs to the current task and to identify abnormal connections. The RDMA operation type includes WRITE, READ, SEND / RECV, etc., which can be used to determine whether there is unauthorized access or abnormal operation.
[0033] Memory access data describes the access location and write method of remote nodes to local memory regions, including memory region identifier, write address range, fragment length, and write offset. Among them, the memory region identifier is the identifier of the memory region that has been registered and allowed to be accessed by RDMA, which can be used to determine whether data is written to the authorized memory region; the write address range is the start and end address of the fragment write to the training side memory, which can be used to determine whether the data fragment has out-of-bounds write, overwrite write, or write to the wrong region; the fragment length is the size of the data written in the current fragment, which can be used to determine whether the data fragment has abnormal truncation, splicing, or insertion; the write offset represents the logical position of the fragment in the data object or buffer, which can be used to determine whether the fragment order is abnormal and to identify out-of-order or replay.
[0034] Time-series traffic data is used to describe the arrival time and transmission rhythm of data fragments during the transmission process, including fragment arrival timestamps, which indicate the time when the fragment arrives at the training side. These timestamps can be used to determine whether the data fragments are transmitted within the valid synchronization window.
[0035] Queue event data is used to describe the operational status of RDMA queue pairs and completion queues, including completion queue event status and error codes, retransmission count, etc. Among them, completion queue event status and error codes indicate the status and error type of RDMA completion events, which can be used to determine whether the operation was completed normally and to distinguish between permission abnormalities, address abnormalities, or link abnormalities. The retransmission count indicates the number of times the RDMA operation was retried, which can be used to determine whether the link is abnormal and to help distinguish between attacks and network disturbances.
[0036] Source authentication data is used to describe the task, dataset, and object version to which the data shard belongs, including task identifier and metadata manifest signature. The task identifier is the training task, dataset, and synchronization batch identifier to which the shard belongs, which can be used to determine whether the data belongs to the current training task. The metadata manifest signature is the data object version and manifest signature, which can be used to identify old version replay, manifest tampering, or unauthorized version mixing.
[0037] In this embodiment, the device can acquire transport layer metadata while data fragments are being written to memory by listening to the completion event of RDMA transmission, without blocking the data transmission path.
[0038] It should be noted that the transport layer risk score can be a numerical indicator used to reflect the degree of deviation of the behavior of data fragments at the transport layer from normal trusted transport behavior. Its value ranges from 0 to 1. The closer the score is to 1, the more the behavior pattern of data fragments at the transport layer deviates from the normal trusted transport pattern, and there may be risks of unauthorized write, replay, overwrite, abnormal insertion or poisoning transmission.
[0039] In a specific implementation, the device can acquire the transport layer metadata corresponding to the data fragments received during the RDMA transmission of the current training task, and convert each metadata item into a numerical feature based on the acquired transport layer metadata. Then, the numerical feature is compared with a pre-built transport layer trust benchmark to calculate the transport layer risk score.
[0040] Step S20: Extract the content layer security features of the data shard, and generate a content layer risk score for the data shard based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data shard at the content level.
[0041] It should be noted that content-layer security features can be statistical or semantic features extracted from the content payload of data segments to characterize the security attributes (such as the presence of poisoning risks) of the data segment at the content level. Specifically, these can include word segmentation distribution skewness, the proportion of long-tailed rare words, the proportion of anomalous characters, the density of anomalous repeated fragments, and semantic distribution drift features. Specifically, word segmentation distribution skewness measures the degree of deviation of the frequency distribution of each word in the current data segment from the baseline of the trusted corpus; the proportion of long-tailed rare words is used to calculate the proportion of low-frequency words in the current data segment; the proportion of anomalous characters is used to calculate the proportion of anomalous characters such as undecodeable bytes, control characters, and zero-width characters in the current data segment; the density of anomalous repeated fragments is used to calculate the density of recurring consecutive word segments in the current data segment; and the semantic distribution drift feature is used to map the data segment into a semantic vector using a pre-trained text embedding model and calculate the degree of deviation of this vector from the semantic center of the trusted corpus. In practical applications, unlike traditional detection methods based on parsing complete file content, the content layer security features in this embodiment can be extracted even when data fragments are incomplete and not visible.
[0042] It should also be noted that the content layer risk score can be a numerical indicator used to reflect the degree of poisoning risk of data shards at the content level. Its value ranges from 0 to 1. The closer the score is to 1, the stronger the poisoning risk signal of the data shards at the content level, such as abnormal word distribution deviating from the trusted baseline, containing a large number of rare words or abnormal characters, and semantic representation being far away from the trusted semantic center.
[0043] In its implementation, the device can perform decoding and word segmentation on the acquired data segments to obtain decoded abnormal bytes. These abnormal bytes are then mapped to abnormal character identifiers. A word segmentation sequence is constructed based on the characters of the normal bytes and the abnormal character identifiers. Content-layer security features, such as word segmentation distribution skewness, the proportion of long-tailed rare words, the proportion of abnormal characters, the density of abnormal repeated segments, and semantic distribution drift features, are then extracted from the word segmentation sequence. Subsequently, the device can input the extracted content-layer security features into a pre-trained content security summarization model. An attention mechanism is used to weight and fuse the features to generate a content security summary value. This value is then mapped to the 0-1 range using a non-linear function to obtain a content-layer risk score.
[0044] Step S30: Perform cross-plane aggregation on the transport layer risk score and the content layer risk score to generate a cross-plane association inference index for the data shard. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score.
[0045] It should be noted that the cross-plane correlation inference index can be a comprehensive index generated by cross-plane aggregation of transport layer risk scores and content layer risk scores, which can be used to characterize the correlation between transport layer risks and content layer risks. In this embodiment, the cross-plane correlation inference index can be aggregated from joint increase index, contradiction index, and multiple correlation risk indicators. Among them, the joint increase index is used to characterize the degree of synchronous increase of transport layer risk scores and content layer risk scores; the contradiction index is used to characterize the degree of deviation between transport layer risk scores and content layer risk scores; the correlation risk indicators include identity correlation risk indicators, source correlation risk indicators, memory correlation risk indicators, and queue correlation risk indicators, which are used to characterize the coupling relationship between transport-side anomaly factors and content anomalies in different dimensions.
[0046] It should also be noted that the "plane" in the cross-plane association inference index refers to the transmission plane and the content plane. The transmission plane focuses on how the data is transmitted, while the content plane focuses on whether the data content itself is risky.
[0047] In its implementation, the device can perform cross-plane aggregation processing on transport layer risk scores and content layer risk scores. It calculates joint escalation indicators and contradiction indicators based on these scores, and determines transport-side anomaly factor scores based on various dimensions of data in the transport layer metadata, including connection identity anomaly scores, source authentication anomaly scores, memory access anomaly scores, and queue event anomaly scores. Then, the device can correlate each transport-side anomaly factor score with the content layer risk score to obtain identity-related risk indicators, source-related risk indicators, memory-related risk indicators, and queue-related risk indicators. Finally, it aggregates the joint escalation indicator, contradiction indicator, and the aforementioned multiple correlated risk indicators into a cross-plane correlation inference indicator.
[0048] Step S40: Perform anomaly detection on the data shards based on the cross-plane association reasoning index to obtain anomaly detection results.
[0049] It should be understood that the anomaly detection result can be the final detection conclusion output after performing anomaly detection on the current data shard based on cross-plane correlation inference indicators. This conclusion may include the security status level of the data shard and the corresponding target correction strategy. The security status level may include four levels: high risk, conflict, suspicious, and safe, indicating the severity of the anomaly in the current data shard. The target correction strategy instructs the device to perform specific protective actions on the current data shard, such as verifying source node authorization, re-verifying metadata manifest signatures, checking memory region access permissions, or pausing the anomaly synchronization window.
[0050] In this embodiment, the device can determine the maximum association risk value among all association risk indicators in the cross-plane association reasoning index, and identify the indicator type corresponding to the maximum association risk value as the dominant association type. Then, based on the joint escalation indicator, contradictory indicator, and maximum association risk value, the device determines the security status level of the data segment according to a preset security status judgment rule. Subsequently, the device can match a target correction strategy from candidate correction strategies based on the dominant association type, and generate anomaly detection results for the data segment based on the security status level and the target correction strategy.
[0051] In the specific implementation, refer to Figure 2 , Figure 2 A system architecture diagram provided for the data detection method of this application. (See diagram below.) Figure 2As shown, firstly, the device can perform training task initialization operations when the training task starts. During initialization, the device can assign a unique task identifier to the current training task, establish an RDMA connection with the remote storage node, register the local memory region that the remote node is allowed to access, and load the metadata list corresponding to the current training task. Next, during RDMA transmission, the remote storage node can directly write data fragments to the registered memory region of the compute node where the device resides through a one-sided write operation. This entire process bypasses the central processing unit and operating system kernel; the data fragments are written into memory directly in hardware. Afterward, the device can detect the arrival of data fragments while they are being written to memory by listening to RDMA completion queue events, and obtain the transport layer metadata associated with the data fragment. Then, the device can calculate a transport layer risk score for the arriving data fragment using module one. Specifically, the device can construct a transport layer feature vector for the data fragment based on the obtained transport layer metadata, and compare this transport layer feature vector with a pre-built transport layer trust benchmark to determine the transport layer risk score of the data fragment. Meanwhile, the device can calculate the content layer risk score through module two. Specifically, the device can decode and segment the content payload of the data fragments, extract content layer security features, and determine the content layer risk score based on the extracted content layer security features using a content security summary model. Then, the device can perform cross-plane aggregation of the transport layer risk score output by module one and the content layer risk score output by module two through module three, generate cross-plane correlation inference indicators, and comprehensively judge the data fragments based on the cross-plane correlation inference indicators to determine the security status level and dominant correlation type of the data fragments, and finally generate the anomaly detection result for the data fragments. Finally, the device can execute corresponding corrective actions based on the security status level and dominant association type in the anomaly detection results. If the security status level is "safe," the data shard is allowed to enter the training sample queue for model training. If the security status level is "high risk," "conflict," or "suspicious," the target corrective strategy is matched from the candidate corrective strategies based on the dominant association type and executed. Specific corrective actions may include verifying the source node authorization and freezing the abnormal connection, re-verifying the metadata manifest signature and rolling back to the trusted object version, checking memory region access permissions and isolating the abnormal shard, or checking the completion queue status and pausing the abnormal synchronization window. Through these corrective actions, the device can intercept and handle data poisoning samples before they enter the training queue.
[0052] This embodiment provides a data detection method. The method discloses acquiring transport layer metadata of data fragments during RDMA transmission for the current training task, determining a transport layer risk score for each data fragment based on the transport layer metadata (the transport layer risk score characterizes the degree of behavioral deviation of the data fragment at the transport layer), extracting content layer security features of the data fragments, and generating a content layer risk score for each data fragment based on these features (the content layer risk score characterizes the degree of poisoning risk of the data fragment at the content layer), and performing cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane correlation inference index for the data fragments (the cross-plane correlation inference index characterizes the correlation between the transport layer risk score and the content layer risk score). The method establishes the correlation between data fragments; it performs anomaly detection on data fragments based on cross-plane correlation inference indicators to obtain anomaly detection results; compared with existing detection methods that only rely on link layer indicators, which cannot distinguish between link disturbances and malicious poisoning, nor can they identify hidden poisoning samples, this embodiment can further extract the poisoning risk level of data fragments at the content level while obtaining the degree of deviation of transport layer behavior, and construct the correlation between transport layer risk score and content layer risk score through cross-plane aggregation, and then perform anomaly detection based on this correlation, thereby being able to distinguish between link disturbances and malicious attacks and identify hidden poisoning samples in the detection process, solving the technical problem of low accuracy of existing data poisoning detection methods that rely on link layer indicators.
[0053] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the data detection method of this application.
[0054] In this embodiment, the step of determining the transport layer risk score of the data fragment based on the transport layer metadata includes steps S11 to S15: Step S11: Construct the transport layer feature vector corresponding to the data fragment based on the transport layer metadata.
[0055] It should be understood that the transport layer feature vector can be a numerical vector constructed based on the transport layer metadata of data fragments, used to quantify and describe various security attributes of data fragments at the transport layer. In this embodiment, the transport layer feature vector can specifically be an ordered array formed by numerically converting each data item in the transport layer metadata, which can be composed of multiple transport layer security feature sub-vectors, including connection identity features. Memory access characteristics Queue event characteristics and source authentication features For example, the first Transport layer feature vectors of data fragments It can be represented as:
[0056] Among them, connecting identity features Used to describe the subject of RDMA transmission connection establishment, connection object, and operational behavior, its internal transmission semantics are as follows: ,in, Indicates the legitimacy characteristics of the source node. Indicates the queue pair authorization characteristics, This represents the state characteristics of the queue. In practical applications, the source node's validity characteristic... This is used to indicate whether the source node initiating the RDMA transfer belongs to the authorized source set of the current training task. It can be calculated based on the source node ID. For example, if the source node ID comes from the authorized source set, then... =1, otherwise The value is 0. The queue has authorization features. This indicates whether the queue pair currently used by RDMA belongs to the pre-authorized set of queue pairs for the current training task, and whether the RDMA operation type belongs to the set of operation types allowed for that queue pair. It can be calculated based on the queue pair number and status, and the RDMA operation type. For example, if the queue pair number comes from the pre-authorized set of queue pairs and the RDMA operation type belongs to the set of operation types allowed for the queue pair task, then... =1, otherwise =0. Queue pair state characteristics This indicates whether the current queue pair is in a normal, allowable transmission state, and it can be calculated based on the queue pair number and its state. The queue pair state includes RTS (Ready To Send), ERR (Error), RESET (Reset), and SQD (Paused). When the queue pair state is RTS, and the state transition path conforms to the normal transmission path INIT→RTR→RTS, then... The value is 1; when the queue is currently in RTS state, but there are unexpected jumps, repeated reconstructions, or multiple switches in a short period of time during historical state transitions, then... This is denoted as an intermediate risk value of 0.5; when the queue pair is in the ERR, RESET, or SQD state, or when data is written before a valid state transition is completed, then... Marked as 0.
[0057] Memory access characteristics This describes the access location, access permissions, and write methods of a remote node to the training-side memory region. Its internal transport semantics are as follows: ,in, Indicates the authorization characteristics of the memory region. This indicates the legality of the address range to be written. This indicates the consistency characteristic of fragment length. In practical applications, it refers to memory region authorization characteristics. This indicates whether the memory region currently being written to is within the memory region authorized by the current task. It can be calculated based on the memory region identifier and the task ID. For example, if the memory region identifier comes from an authorized memory region, then... =1, otherwise The value is 0. This indicates a valid write address range feature. This is used to indicate whether the data fragment write address range falls entirely within the authorized address range of the current memory region. It can be calculated based on the memory region identifier and the write address range. For example, if the memory region identifier and the write address range are both valid, then... The value is 1; if there is an out-of-bounds error, then... =0. Fragment length consistency feature. This indicates how close the actual write length of the current data fragment is to the expected length declared in the task security context or manifest. It can be calculated based on the fragment length and manifest signature, and the corresponding calculation formula is as follows: The closer it is to 1, the closer the actual result is to the expected result.
[0058] Queue event characteristics Used to describe the operational status of RDMA queue pairs and completion queues, its internal transmission semantics are as follows: ,in, This indicates the completion status characteristics of completed queue entries. This represents the retransmission count characteristic. In practical applications, it represents the completion status characteristic of a completion queue entry. This is used to indicate whether the RDMA completion event corresponding to the current data fragment has been completed normally. It can be calculated based on the completion queue entry status, error code, and fragment arrival timestamp. For example, when the completion queue entry status is "Success" and the fragment arrival timestamp is within the synchronization window allowed by the current training task, it indicates normal completion. The value is 1; if the status of a completed queue entry is Success, but the arrival timestamp of the fragment is not within the valid synchronization window, it indicates that there may be premature injection, delayed replay, or abnormal insertion behavior. The value is 0; when a temporary error occurs due to link congestion, then... 0.5 otherwise 0. Retransmission count characteristic. This is used to indicate whether the number of RDMA retries occurring during the current data fragmentation transmission is within a reliable range, and it can be calculated based on the number of retries.
[0059] Source authentication features Used to describe the training task to which the data shard belongs, the version of the data object, and the authenticity of the inventory, its internal transport semantics are as follows: ,in, This indicates the task ID matching feature. This indicates the version consistency characteristic of an object. This represents the manifest signature verification feature. In practical applications, it's the task ID matching feature. This is used to indicate whether the task ID of the current data shard declaration or associated task matches the current training task. It can be calculated based on the task ID; for example, if the task ID matches the current training task, then... =1, otherwise =0. Object version consistency characteristic. This indicates whether the version of the object to which the current data shard belongs is consistent with the version of the object declared in the current task list. It can be calculated based on the list signature; for example, if the list signature matches the version of the object declared in the current task list, then... =1, otherwise =0. List signature verification feature The signature used to indicate whether the current list of data objects has passed verification can be calculated based on the list signature. For example, if the list signature has passed verification, then... =1, otherwise It is 0.
[0060] In practical implementation, after obtaining the transport layer metadata of data fragments during RDMA transmission, the device first converts this transport layer metadata into a numerical feature vector. Specifically, the device can extract corresponding numerical features from each data item of the transport layer metadata according to predefined feature extraction rules. Among these features, the source node legitimacy feature is extracted first. The device can obtain the source node ID from the transport layer metadata and determine whether the source node ID belongs to the authorized source set of the current training task. If it does, the feature is set to 1; otherwise, it is set to 0. For queue pair authorization features... The device can obtain the queue pair number and RDMA operation type from the transport layer metadata, and determine whether the queue pair number belongs to the pre-authorized queue pair set for the current training task and whether the RDMA operation type belongs to the set of operation types allowed for that queue pair. If so, the value is 1; otherwise, the value is 0. Regarding queue pair status characteristics... The device can obtain the queue pair status from the transport layer metadata. The value is 1 when the queue pair status is RTS and the state transition path conforms to a normal transport path; 0.5 when the queue pair status is RTS but there are unexpected jumps in the historical state transition; and 0 when the queue pair status is in an error state or a valid state transition has not been completed. Then, the device can verify the validity of the source node. Queue for Authorization Features and queue state characteristics By splicing the data, a connection identity feature is generated. ,Right now .
[0061] For memory region authorization features The device can obtain the MR identifier from the transport layer metadata and determine whether the MR identifier belongs to the memory region authorized by the current task. If it does, the value is 1; otherwise, the value is 0. (This is related to the write address range validity feature.) The device can obtain the write address range from the transport layer metadata and determine whether the write address range falls entirely within the current MR's authorized address range. If so, the value is 1; otherwise, the value is 0. (Regarding fragment length consistency features...) The device can obtain the fragment length and Manifest signature from the transport layer metadata and calculate them according to the formula. The calculation is performed, and the closer the result is to 1, the closer the actual result is to the expected result. The device can then authorize features for memory regions. , Write address range validity features Consistency features with fragment length Concatenate the data to generate memory access characteristics. ,Right now .
[0062] CQE completion status characteristics The device can obtain the CQE status, error code, and fragment arrival timestamp from the transport layer metadata. The value is 1 when the CQE status is successful and the fragment arrival timestamp is within the valid synchronization window; 0 when the CQE status is successful but the fragment arrival timestamp is outside the valid synchronization window; 0.5 when a temporary error occurs due to link congestion; otherwise, 0. Regarding the retransmission count characteristic... The device can obtain the retransmission count from the transport layer metadata and determine its value based on whether that count is within a reliable range. Then, the device can complete the state characteristics of the CQE. Features of retransmission count Concatenate the data to generate queue event characteristics. ,Right now .
[0063] For task ID matching features The device can obtain the task ID from the transport layer metadata and determine whether the task ID is consistent with the current training task. If it is, the value is 1; otherwise, the value is 0. (Regarding object version consistency features...) The device can obtain the Manifest signature from the transport layer metadata and determine whether the version of the object to which the current fragment belongs is consistent with the version of the object declared in the current task's Manifest. If so, the value is 1; otherwise, the value is 0. This is for Manifest signature verification features. The device can obtain the Manifest signature from the transport layer metadata and determine whether the signature has passed verification. If it has, the value is 1; otherwise, the value is 0. Then, the device can match features to the task ID. Object version consistency feature and Manifest signature verification features The features are then concatenated to generate source authentication features. ,Right now .
[0064] Afterwards, the device can verify the connection's identity characteristics. Memory access characteristics Queue event characteristics and source authentication features The data is then spliced together to generate data fragments. Transport layer feature vector ,Right now
[0065] Step S12: Determine the directed dependency relationships between several transport layer security feature nodes based on the constraints in the RDMA transmission process.
[0066] It should be noted that constraints can be rules and execution order restrictions at the business logic level during RDMA transmission. These constraints determine whether there are causal relationships or conditional dependencies between security feature nodes at different transport layers. In practical applications, constraints can originate from the working principle of the RDMA transmission protocol and the business logic of the storage-computation separation training task. For example, in RDMA transmission, constraints at the access control level stipulate that only authorized source nodes can use authorized queue pairs; constraints at the queue pair state machine level stipulate that only queue pairs in a normal state (i.e., RTS state) can complete data transmission normally; constraints at the memory access permission level stipulate that only authorized memory areas can be written to by remote nodes; constraints at the training task business logic level stipulate that the task context limits the writable memory areas and valid object versions; and constraints at the data integrity verification level stipulate that there is a binding relationship between object versions and manifest signatures.
[0067] It should also be noted that transport layer security feature nodes can be various transport layer security features that function as graph nodes in the transport layer semantic probability graph, where each transport layer security feature node corresponds to a feature component in the transport layer feature vector. For example, transport layer security feature nodes may include source node legitimacy nodes, task ID matching nodes, QP authorization nodes, QP status nodes, MR authorization nodes, write address legitimacy nodes, fragment length consistency nodes, CQE completion status nodes, retransmission count nodes, object version consistency nodes, and Manifest signature verification nodes, etc.
[0068] It should be noted that directed dependency can refer to a unidirectional causal relationship or conditional dependency between transport layer security feature nodes. It can be used to indicate that the value state of one feature node affects the value state of another feature node, or that the normality or abnormality of one feature node affects whether another feature node is abnormal. In this embodiment, the directed dependency can be determined by the constraints in the RDMA transmission process, rather than an arbitrary statistical correlation.
[0069] In practical implementation, the device can establish directed dependencies between several transport layer security feature nodes based on constraints during RDMA transmission. Specifically, the device can establish dependencies between transport layer security feature nodes according to the service constraints and execution order specified in the RDMA transmission protocol. For example, a directed dependency can be established based on the constraint that "only authorized nodes can use authorized QPs." Establish a directed dependency relationship based on the constraint of "legality of state transitions in authorized QPs": Establish a directed dependency relationship based on the constraint that "authorized QPs can only access authorized MRs": Establish directed dependencies based on the constraint that "the task context limits the writable memory area": Establish a directed dependency relationship based on the constraint that "the task has determined a valid object version": Establish a directed dependency relationship based on the constraint that "object version is closely related to Manifest signature": Establish a directed dependency based on the constraint that "the object version specifies the expected fragment length": Establish a directed dependency relationship based on the constraint that "only authorized MapReduces (MRs) have valid address ranges": Based on the constraint that "different MRs may have different expected fragment length characteristics," a directed dependency relationship is established: Establish a directed dependency relationship based on the constraint that "the QP state directly affects the completion event state": Establish a directed dependency relationship based on the constraint that "abnormal completion status will be accompanied by retransmission": After this, the device can organize all established directed dependencies into a set of directed dependencies. .
[0070] Step S13: Construct the transport layer semantic probability graph of the current training task based on each of the transport layer security feature nodes and the directed dependency relationship.
[0071] It should be noted that the transport layer semantic probabilistic graph can be a probabilistic graphical model used to describe the conditional dependencies and joint probability distributions among the transport layer security feature nodes during normal RDMA data fragmentation transmission. It can be composed of a set of transport layer security feature nodes and a set of directed dependency edges, specifically represented as follows:
[0072] in, Represents the set of transport layer security feature nodes. The number of features is denoted by , where each transport layer security feature node corresponds to one transport layer security feature. This represents the set of directed dependency edges between feature nodes, where each directed dependency edge represents the conditional dependency relationship between two transport layer secure feature nodes.
[0073] In this embodiment, the device can treat all transport layer security feature nodes as a graph node set. Directed dependencies are treated as a set of directed edges. This involves constructing the transport layer semantic probability graph for the current training task. In practical applications, the structure of the transport layer semantic probability graph (i.e., the set of nodes and edges) is predetermined based on the working principle of the RDMA transport protocol and the business logic of the training task, and remains unchanged throughout the entire lifecycle of the current training task once determined.
[0074] Step S14: Construct a transport layer trust benchmark for the current training task based on the set of historical transport layer feature vectors generated during the trusted transport phase and the transport layer semantic probability graph. The transport layer trust benchmark is used to characterize the conditional dependencies and normal value ranges between the security features of each transport layer under trusted transport conditions.
[0075] It should be understood that the trusted transmission phase can be a normal transmission period in the current training task where the data source is trustworthy and there are no data poisoning attacks or transmission anomalies. During the trusted transmission phase, all data fragments received by the device are trusted training data fragments, their transmission behavior conforms to the normal RDMA transmission mode, and their content consists of untampered, trustworthy samples that can be used to build a trusted baseline for the transport layer. In practical applications, the trusted transmission phase can be determined manually or automatically through pre-set security policies, such as defining the initial stable transmission window after the training task starts as the trusted transmission phase.
[0076] It should also be understood that the historical transport layer feature vector set can be the set of transport layer feature vectors from all data fragments collected during the trusted transport phase. For example, for any training task... The device can collect data during the trusted transmission phase. The transport layer feature vectors of each data fragment constitute the historical transport layer feature vector set. The corresponding expression can be represented as:
[0077] in, Indicates the first The transport layer feature vector of a trusted data fragment. This represents the total number of data fragments collected during the trusted transmission phase, which is also the number of historical trusted samples. In practical applications, each feature vector in the historical transmission layer feature vector set is constructed according to the same rules as in the online detection phase, i.e., it includes connection identity features, memory access features, queue event features, and source authentication features.
[0078] It should be noted that the transport layer trust benchmark can be a data model used to characterize the conditional dependencies and normal value ranges between the security features of each transport layer under a trusted transmission state. It can be used to measure whether the transmission behavior of the current data fragment deviates from the normal mode.
[0079] In this embodiment, the transport layer trusted benchmark It can be constructed by the conditional probability tables of all transport layer security feature nodes in the transport layer semantic probability graph, which record the conditional probability distribution of each feature node in the transport layer semantic probability graph under the various combinations of values of its parent node.
[0080] Further, the step of constructing a transport layer trust benchmark for the current training task based on the set of historical transport layer feature vectors generated during the trusted transport phase of the current training task and the transport layer semantic probability graph includes: Step S141: Based on the set of historical transport layer feature vectors generated by the current training task in the trusted transport phase, calculate the conditional probability of each transport layer security feature node in the transport layer semantic probability graph under the corresponding parent node value combination.
[0081] Understandably, the combination of parent node values can be a specific combination of the values of all parent nodes (i.e., the preceding nodes that directly point to this node) of a transport layer security feature node with directed edges in the transport layer semantic probability graph. In the transport layer semantic probability graph, a transport layer security feature node may have one or more parent nodes. When there is only one parent node, the combination of parent node values is simply the value of that parent node; when there are multiple parent nodes, the combination of parent node values is the Cartesian product of the values of all parent nodes, i.e., a combination of all parent node values. For example, if a node... There are two parent nodes and ,and The value can be {0, 1}. If the value of is {0,1}, then the combination of values of the parent node can include: (0,0), (0,1), (1,0) and (1,1).
[0082] It should be noted that conditional probability can be the probability that a transport layer security feature node takes a specific value given a specific combination of values for its parent node. It can be used to quantify the strength of the statistical dependency between parent and child nodes under trusted transmission conditions. In this embodiment, the device can count the total number of times the parent node's value combination appears in historical trusted samples, and the number of times the child node takes a specific value under that parent node value combination, and determine the ratio of the two as the conditional probability under that condition.
[0083] In this embodiment, the formula for calculating the conditional probability of a transport layer security feature node can be:
[0084] in, Indicates the first A transport layer security feature node, This represents one of the possible discrete values of the node. Represents a node The set of parent nodes, This represents a combination of values for the set of parent nodes. Indicates the number of occurrences in historical reliable samples As a smoothing factor, Represents a node The number of discrete values.
[0085] In practical applications, the device can first obtain the set of historical transmission layer feature vectors generated during the trusted transmission phase of the current training task. The device discretizes each feature vector in the historical transport layer feature vector set. For the fragment length consistency feature, the device can divide the continuous values of this feature into several intervals; for the retransmission count feature, the device can divide the values of this feature into discrete categories such as 0 retransmissions, 1 retransmission, 2 retransmissions, and multiple retransmissions. Subsequently, for each node in the transport layer semantic probability graph... The device can determine the set of parent nodes of the node. And based on the number of discrete values of each parent node in the parent node set, calculate all possible combinations of parent node values. For each combination of parent node values... The device can count the first occurrence of the parent node's value combination in the historical transport layer feature vector set. and the values are combined in the parent node. Historical sample subset statistics of the current node Take each discrete value The second occurrence Then, the conditional probability is obtained by adding a smoothing factor to the second occurrence count and dividing by the product of the first occurrence count, the smoothing factor, and the number of discrete values of the node.
[0086] Step S142: Construct a conditional probability table for each of the transport layer security feature nodes based on the conditional probabilities.
[0087] It should be understood that a conditional probability table can be a data structure used to store all conditional probabilities of a transport layer security feature node in a transport layer semantic probability graph under all possible combinations of values of its parent nodes. For each node in the transport layer semantic probability graph, the conditional probability table can record the conditional probability distribution of the node taking each discrete value under all possible combinations of values of its parent nodes.
[0088] In practical applications, for each node The device can organize the conditional probabilities of a node under all combinations of parent node values and all discrete values into a conditional probability table according to a predetermined format, wherein the row index of the conditional probability table is the combination of parent node values. The column index is a discrete value of the current node. Each entry represents a corresponding conditional probability value. Subsequently, the device can construct a conditional probability table for each node in the transport layer semantic probability graph.
[0089] Step S143: Determine the transport layer reliability benchmark for the current training task based on the conditional probability table.
[0090] In practical applications, the device can aggregate the conditional probability tables of all nodes in the transport layer semantic probability graph to form a transport layer trust benchmark for the current training task. This transport layer trust benchmark is stored in the form of a set of conditional probability tables, with each conditional probability table corresponding to a node in the transport layer semantic probability graph.
[0091] In this embodiment, based on the historical transport layer feature vector set generated during the trusted transmission phase of the current training task, the conditional probabilities of each transport layer security feature node in the transport layer semantic probability graph under the corresponding parent node value combinations are statistically analyzed. This allows the construction of a conditional probability table for each node, which is then aggregated to form the transport layer trusted benchmark for the current training task. This ensures that the transport layer trusted benchmark accurately reflects the conditional dependencies and normal value ranges between security features under normal transmission conditions, rather than using a general static threshold or cross-task statistical model. This effectively adapts to differences in transmission behavior between different training tasks. Furthermore, since the transport layer trusted benchmark explicitly records the complete probability distribution of each feature node under each parent node value combination in the form of a conditional probability table, the subsequent calculation of the transport layer risk score for each data fragment in the online detection phase has a clear statistical benchmark and probabilistic basis as support. This ensures the interpretability and verifiability of transport layer anomaly determination, avoiding the difficulty in tracing the determination results caused by black-box models.
[0092] Step S15: Determine the transport layer risk score of the data fragment based on the transport layer feature vector and the transport layer trust benchmark.
[0093] In practical applications, after obtaining the transport layer trusted benchmark, the device can use the transport layer feature vector of the current data fragment as input, targeting each feature node in the transport layer semantic probability graph. Query the current value combination of the node in the parent node. Under the condition, the current node takes the value Conditional probability at time The system calculates the logarithm of the conditional probabilities for all nodes and sums them to obtain a joint logarithmic matching value. A higher joint logarithmic matching value indicates that the combination of transport layer features in the current data fragment closely matches historical reliable transmission patterns; a lower joint logarithmic matching value indicates that the combination of transport layer features in the current data fragment deviates significantly from historical reliable transmission patterns. Subsequently, the device calculates the average matching value based on historical reliable samples and, combined with low-match boundaries observed in the reliable samples, converts the matching value of the real-time fragment into a transport layer deviation, which is then used as a transport layer risk score.
[0094] Further, step S15 includes: Step S151: Determine the joint logarithmic matching value of the data fragment based on the transport layer feature vector and the transport layer trust benchmark.
[0095] It should be noted that the joint logarithmic matching value can be used as an indicator to quantify the degree to which the transport layer feature vector of the current data fragment conforms to the transport layer trust benchmark as a whole. In this embodiment, the device can obtain the joint logarithmic matching value by inputting the transport layer feature vector of the current data fragment into the transport layer trust benchmark, taking the logarithm of the conditional probability of each transport layer security feature node under the current value of the parent node, and then summing the results. The corresponding calculation formula can be:
[0096] in, The number of transport layer security feature nodes, For the first Each feature node The value of the current data shard on this node. For the current data shard, the th The combination of values of the parent node of each node. This represents the conditional probability obtained from the transport layer trusted benchmark. In practical applications, the joint logarithmic matching value is a negative or zero value. The smaller the absolute value, the more the combination of features of the current data fragment conforms to the historical trusted transmission pattern; the larger the absolute value, the greater the deviation of the combination of features of the current data fragment from the historical trusted transmission pattern.
[0097] Step S152: Determine the average matching value and low matching boundary of historical reliable samples.
[0098] It should be understood that historical trusted samples can be transport layer feature vectors of historical data fragments belonging to the current training task, collected during the trusted transport phase. These vectors can originate from the construction phase of the transport layer trusted baseline for the current training task and are the raw data used to statistically analyze the conditional probabilities of each node. In this embodiment, the set of historical trusted samples is represented as follows:
[0099] in, This represents the number of historically reliable samples.
[0100] Understandably, the average matching value can be the numerical result obtained by taking the arithmetic mean of the joint logarithmic matching values of each data fragment in the historical trusted samples. It can be used to reflect the average degree of conformity between the transport layer feature vector of a normal data fragment and the transport layer trusted benchmark during the trusted transmission phase. In this embodiment, the average matching value... The corresponding calculation formula can be:
[0101] in, For the first The joint logarithmic matching value of a historical reliable sample.
[0102] It is also understood that the low matching boundary can be the lowest acceptable value used to characterize the joint matching degree between the transport layer feature vector of the data fragment and the transport layer trust benchmark under the current training task in a trusted transmission state. That is, when the joint logarithmic matching value of a data fragment is lower than this boundary value, it indicates that the transmission behavior of the data fragment deviates from the historical trusted transmission pattern. In this embodiment, the device can calculate the joint logarithmic matching value for each historical trusted sample in the historical trusted sample set, and select a lower quantile or boundary value in the distribution of the joint logarithmic matching values of the historical trusted samples as the low matching boundary. For example, the low matching boundary can be the 5th quantile of the joint logarithmic matching value of the historical trusted samples, or it can be the minimum value of the joint logarithmic matching value of the historical trusted samples. This embodiment does not limit this.
[0103] Step S153: Determine the transport layer risk score of the data fragment based on the joint logarithmic matching value, the average matching value, and the low matching boundary.
[0104] In practical applications, the device can be based on joint logarithmic matching values. Average matching value and low matching boundary Calculate the transport layer deviation of data fragmentation The corresponding calculation formula can be:
[0105] in, It is a very small positive number, used to avoid the denominator being zero.
[0106] Then, the device can determine the transport layer deviation as a transport layer risk score for data fragmentation. ,Right now:
[0107] in, The value ranges from 0 to 1. The closer the score is to 0, the more the transport layer behavior of the current data fragment conforms to the historical reliable benchmark. The closer the score is to 1, the more serious the deviation of the transport layer behavior of the current data fragment from the normal mode, and there may be risks of unauthorized write, replay, overwrite, abnormal insertion or poisoning transmission.
[0108] In this embodiment, the joint logarithmic matching value of the data segment is determined by comparing the transport layer feature vector of the current data segment with the transport layer trusted benchmark. The average matching value and low matching boundary of historical trusted samples are determined, and then the transport layer risk score is determined based on the joint logarithmic matching value, average matching value and low matching boundary. This makes the transport layer risk score have a clear trusted benchmark reference and a unified quantitative scale, thereby eliminating the problem of incomparable scores caused by differences in feature distribution between different training tasks and different transport stages.
[0109] In the specific implementation, refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the transmission risk scoring process in the data detection method of this application. For example... Figure 4 As shown, the device can monitor RDMA completion queue events, detect the arrival of RDMA data fragments while they are being written to memory, and collect lightweight transport metadata upon arrival, including four dimensions: connection identity data, queue event data, memory access data, and source authentication data. Subsequently, the device can perform numerical transformation on each data item in the transport layer metadata according to predefined feature extraction rules, forming an ordered numerical feature vector, including connection identity features, memory access features, queue event features, and source authentication features. Based on these features, a transport layer feature vector is then constructed. Next, the device can use the constructed transport layer feature vector as input to query the RDMA semantic probability map pre-constructed for the current training task. Based on the structure of the RDMA semantic probability map, it queries the conditional probability of each feature node under the current value combination of its parent node in the transport layer trust baseline, calculates the joint logarithmic matching value, and then compares the joint logarithmic matching value of the transport layer feature vector of the current data segment with the historical trust transport baseline (i.e., the transport layer trust baseline) to calculate the transport layer deviation. Finally, the calculated transport layer deviation is determined as the transport layer risk score of the current data segment and output.
[0110] Further, the step of extracting the content layer security features of the data fragment and generating a content layer risk score for the data fragment based on the content layer security features includes: Step S21: Decode and segment the data fragments, identify the decoded abnormal bytes in the data fragments, and map the decoded abnormal bytes to abnormal character identifiers.
[0111] It should be understood that decoded aberration bytes can be bytes that cannot be correctly mapped to valid characters according to a preset character encoding standard (e.g., UTF-8, UTF-16, ASCII, etc.) encountered when decoding the content of a data fragment according to that standard. During RDMA transmission, due to WAN link noise, bit flipping, packet corruption, or malicious tampering by attackers, data fragments may contain bytes that do not conform to the expected encoding format; these bytes are decoded aberration bytes. In this embodiment, decoded aberration bytes can reflect that data fragments may have been corrupted or maliciously tampered with during transmission or storage; therefore, the device can retain these decoded aberration bytes as one of the sources of abnormal signals.
[0112] It should also be understood that the abnormal character identifier can be an identifier obtained by mapping the decoded abnormal byte to a special symbol. This identifier can be used to mark and count the position and number of decoded abnormal bytes during subsequent feature extraction. In practical applications, when a device encounters a decoded abnormal byte, it typically does not discard the byte but instead maps it to a preset special symbol, such as...<UNK_BYTE> This allows decoded abnormal bytes to be converted into symbols that can participate in subsequent statistical calculations, enabling the risks of garbled text and abnormal characters to be quantified and included in content layer security features.
[0113] In practical implementation, when a device needs to perform content-layer risk assessment on data fragments, it first obtains the original byte sequence of the data fragment and decodes it according to a preset character encoding standard (e.g., UTF-8). During the decoding process, if a byte sequence that does not conform to the encoding standard is encountered, it can be mapped to a preset special symbol.<UNK_BYTE> , used as an abnormal character identifier.
[0114] Step S22: Construct a word segmentation sequence based on the normal bytes and the abnormal character identifier in the data segment, and extract the content layer security features of the data segment according to the word segmentation sequence. The content layer security features include word segmentation distribution skewness, long-tail rare word segmentation ratio, abnormal character ratio, abnormal repeated fragment density features, and semantic distribution drift features.
[0115] Understandably, a word segmentation sequence can be an ordered arrangement of tokens obtained after performing word segmentation processing on the content payload of data fragments. Here, a token is the basic unit after text word segmentation, which can be a complete word, a character, a punctuation mark, or a sub-word unit. In this embodiment, the device can combine successfully decoded characters and mapped abnormal character identifiers into a character sequence, and perform word segmentation processing on this character sequence according to the word segmenter's vocabulary to divide the continuous character sequence into independent tokens, obtaining a word segmentation sequence. This word segmentation sequence can be represented as... ,in This represents the total number of tokens in the shard.
[0116] It should be noted that content-layer security features can be statistical or semantic features extracted from the content payload of data fragments to characterize whether the content of the data fragment has a risk of being poisoned. These features can be extracted even when the data fragments are not fully visible. In this embodiment, content-layer security features include word segmentation distribution skewness, the proportion of long-tail rare word segments, the proportion of abnormal characters, the density features of abnormal repeated segments, and semantic distribution drift features.
[0117] It should be noted that word segmentation distribution skewness can be a content-layer security feature used to characterize the degree of deviation of the frequency distribution of each token in the current data segment from the trusted corpus baseline. In data poisoning attacks, attackers typically introduce abnormally high-frequency specific tokens, fixed template phrases, pseudo-natural language fragments, or repetitive instruction structures into the poisoned samples, causing the token frequency distribution of the poisoned samples to deviate from the distribution of the normal corpus. This degree of deviation is then termed word segmentation distribution skewness. In this embodiment, word segmentation distribution skewness can be obtained by calculating the Jensen-Shannon distance between the token distribution of the current segment and the trusted corpus baseline distribution. The Jensen-Shannon distance is a symmetry measure of the difference between two probability distributions. Specifically, the device can first acquire a pre-built trusted corpus containing a large number of normal training text samples, and then construct a token probability baseline based on this trusted training corpus. :
[0118] in, This indicates that the word segmenter has a fixed vocabulary, i.e., a reliable corpus. |V| represents the number of times token v appears in the trusted corpus, and |V| represents the vocabulary size.
[0119] For the current data sharding The device can count the occurrences of each token in the data shard and calculate the token distribution of the data shard. The corresponding calculation formula can be:
[0120] in, This indicates that token v is in the current data shard. The number of times it appears in This indicates the total number of tokens in the current data shard.
[0121] Subsequently, the device can use the Jensen-Shannon distance to calculate the deviation of the token distribution of the current data fragment from the baseline of the trusted corpus. Specifically, the device can first calculate the average distribution of the two distributions. The corresponding calculation formula can be:
[0122] Then, the Jensen-Shannon distance is calculated as the skewness of the word segmentation distribution. The corresponding calculation formula can be:
[0123] in, Shard the current data The probability distribution of the occurrence of token v in the middle. This represents the baseline probability distribution of token v in the trusted corpus. Let these be the mean distributions of these two distributions. In practical applications, The larger the value, the more the token distribution of the current data segment deviates from the baseline of the credible corpus, meaning that the data segment is more likely to have anomalies in its word usage patterns.
[0124] It should also be noted that the long-tail rare token ratio can be a content-layer security feature used to statistically analyze the proportion of low-frequency tokens in the current data segment to the total number of tokens. In data poisoning, attackers often use rare strings, special trigger tokens, pseudo-random tags, hidden instructions, or abnormal encoded fragments. These contents usually correspond to rare tokens that appear very infrequently in the segmenter's vocabulary. In this case, the device can calculate the long-tail rare token ratio by statistically analyzing the ratio of the number of tokens belonging to the rare token set in the current segment to the total number of tokens in the segment. Specifically, the device can first define tokens that appear less than a preset threshold (e.g., less than 10 times in a trusted corpus) as rare tokens. This set of rare tokens can be represented as:
[0125] in, The rarity threshold is the probability threshold corresponding to a token appearing less than 10 times in a trusted corpus, and it is fixed as follows:
[0126] Then, the device can calculate the current data fragment. The proportion of long-tail rare word segments for:
[0127] in, This indicates that the data shard belongs to the rare token set. The number of tokens, This indicates the total number of tokens in the data shards. In practical applications, The larger the value, the more rare tokens the current data shard contains, and the higher the probability of poisoned content (such as rare strings, pseudo-random tags, hidden instructions, etc.).
[0128] Understandably, the anomalous character ratio can be a content-layer security feature used to characterize the degree of anomalous character contamination in data fragments. When attackers insert or replace training data during the synchronization phase, they may introduce encoded corruption, invisible control characters, mixed encoded content, garbled fragments, or unexpected binary content. In this case, the device can introduce the anomalous character ratio to measure the degree of contamination of the data fragments by these anomalous characters. In this embodiment, the anomalous character ratio can be calculated by statistically analyzing the ratio of the number of anomalous characters in the fragment to the total number of characters. Specifically, the device can first fragment the data... Decoded into a Unicode character sequence: And define the abnormal character set G as a character set that meets any of the following conditions: (1) Unicode category is control character Cc, such as newline, carriage return, tab, etc., but does not include normal whitespace characters; (2) Unicode category is private zone character Co, the characters in this zone are not assigned specific character meanings in the unified character encoding standard, and are usually used for private protocols or custom purposes; (3) Unicode category is surrogate zone character Cs, the characters in this zone are only used in UTF-16 encoding to represent high-bit surrogate and low-bit surrogate of auxiliary plane characters, and should not appear independently after being decoded into Unicode scalar values; (4) the abnormal character identifier (i.e., ") obtained after decoding failure.<UNK_BYTE> (5) Zero-width characters, including U+200B (zero-width space), U+200C (zero-width non-connector), U+200D (zero-width connector), and U+FEFF (zero-width non-newline space). At this time, the device can calculate the proportion of abnormal characters using the following formula. :
[0129] in, This represents the number of abnormal characters belonging to the abnormal character set G in the data shard. This represents the total number of characters after data fragment decoding.
[0130] It should be noted that the abnormal repeating fragment density feature can be a content-layer security feature used to characterize the density of repeating fragments in data shards. In practical applications, poisoned samples often contain repeated trigger phrases, repeated templates, copy-paste polluting text, or batch-generated fixed sentence patterns. This type of abnormal content will form a large number of repeating fragments in the token sequence. Therefore, this embodiment can introduce the abnormal repeating fragment density feature to identify such structural anomalies.
[0131] In this embodiment, the device can calculate the density feature of abnormal repeating segments by dividing the fragmented token sequence into continuous segments of a fixed length, counting the occurrence frequency of each segment, and calculating the density of repeating segments. Specifically, the device can first divide the token sequence into segments of a fixed length. Slicing consecutive token fragments of fixed length n (e.g., n=8) yields an 8-gram set: ,in This represents a segment consisting of 8 consecutive tokens starting from the k-th token. The device can then count the occurrences of each 8-gram segment within the data fragment. And calculate the density of abnormal repeating fragments. :
[0132] in, This means summing up the repetition counts of all repeated segments (i.e., the number of occurrences minus the first occurrence). Indicates data sharding The total number of 8-gram segments. When <8 o'clock, , The larger the value, the higher the proportion of duplicate segments in the current data segment, and the higher the possibility of duplicate templates and copy-paste text pollution.
[0133] It should also be noted that semantic distribution drift features can be content-layer security features used to characterize the degree of deviation of the semantic representation of a data shard relative to the semantic center of the trusted training data. In practical applications, even if the attack text does not contain obvious abnormal tokens or repetitive structures, it may still deviate from the distribution of the trusted training data in the semantic space. For example, it may contain a concentrated occurrence of a certain abnormal topic, behavior hijacking samples, pseudo-question-answer templates, or target category contamination samples. Therefore, this embodiment can introduce semantic distribution drift features to measure the degree of abnormality of the data shards at the semantic level.
[0134] In this embodiment, the device can map data fragments into semantic vectors using a pre-trained text embedding model (e.g., Sentence-BERT) and calculate the deviation of these vectors from the semantic center of the trusted corpus to obtain semantic distribution drift features. Specifically, the device can first use a pre-trained text embedding model (e.g., Sentence-BERT) as a fixed semantic encoder. Sentence-BERT is a sentence-level semantic encoding model based on the BERT (Bidirectional Encoder Representations from Transformers) network, capable of encoding text of arbitrary length into fixed-dimensional semantic vectors. The cosine distance between semantic vectors can effectively measure the semantic similarity between two text segments. In the offline stage, the device can process each sample in the trusted training corpus... The semantic vectors are vectorized, and each semantic vector is normalized (L2 norm normalization). Then, the average of all normalized semantic vectors is calculated to obtain the credible semantic center. The corresponding calculation formula can be:
[0135] in, The total number of credible corpus samples. It is the m-th sample in the credible corpus. Indicates the sample The semantic vector output after inputting into the Sentence-BERT model. This represents the L2 norm (i.e., the Euclidean length) of the vector.
[0136] For the current data slice, the device can input the content of the data slice into the Sentence-BERT model to obtain a semantic vector, and then perform normalization processing to obtain a normalized semantic vector:
[0137] in, This indicates the content of the current data shard. This represents the normalized semantic vector. It also shows the semantic distribution drift characteristics of the current data slice. It can be:
[0138] in, This represents the cosine similarity between the unitized semantic vector of the current data slice and the trusted semantic center. In practical applications, The larger the value, the more the semantic vector of the current data slice deviates from the credible semantic center, that is, the higher the degree to which the data slice deviates from the normal training data at the semantic level.
[0139] Furthermore, to ensure that features of different dimensions can be uniformly input into the security summary model, after extracting the five content-layer security features mentioned above, the device can perform normalization processing on each feature based on trusted baseline statistics. Specifically, for the k-th feature, the device can obtain the pre-statistical mean of each feature across the trusted corpus fragment set. and standard deviation :
[0140]
[0141] in, Then, the device can standardize each feature according to the corresponding normalization formula to eliminate differences in units and value ranges between features, enabling features of different dimensions to be fused at the same scale. The normalization formula can be:
[0142] Obtain the normalized content layer security feature vector .
[0143] Step S23: Determine the content security summary value of the data segment based on the word segmentation distribution skewness, the proportion of long-tail rare word segments, the proportion of abnormal characters, the density feature of abnormal repeated segments, and the semantic distribution drift feature using a preset content security summary model.
[0144] It should be noted that the preset content security summarization model can be a pre-trained lightweight neural network model, which can be used to non-linearly fuse content-layer security features and output shard-level security summary values. In practical applications, the device can acquire a large number of labeled training samples, each of which includes a content-layer security feature vector for a data shard and a corresponding risk label (0 indicates a trusted shard, 1 indicates a poisoning or anomalous shard). Then, the device can perform supervised training on the content security summarization model based on these labeled samples, optimizing the model parameters (including the attention weight matrix) through backpropagation and a binary cross-entropy loss function. Attention bias (etc.) to minimize the error between the model's output and the labeled results. After training, the device can permanently deploy the model parameters, which will only participate in the forward computation during the subsequent online detection phase without being updated.
[0145] In the subsequent online detection phase, the device can first normalize the five content layer security features to obtain a standardized content layer security feature vector. And the standardized content layer security feature vector Input a pre-defined content security summary model, which calculates the attention weight vector as follows: :
[0146] in, Represents the attention weight matrix. This indicates attention bias. This represents the attention weight vector for security features at each content layer. Indicates the first The contribution weight of each content layer security feature to the current data fragmentation risk assessment. is the activation function used to map the input numerical vector to a probability distribution in which the sum of all elements is 1.
[0147] It should also be noted that the content security digest value can be a comprehensive value output by weighting and fusing content layer security features through a preset content security digest model. It can be used to characterize the comprehensive degree of anomaly of the current data shard at the content level.
[0148] In this embodiment, the device can transmit the attention weight vector. Compared with the standardized content layer security feature vector Perform a weighted summation to obtain the content security digest value. :
[0149] in, Indicates the first Attention weights for each content layer security feature. Represents the standardized first Content layer security feature values. In practical applications, attention weights... This can be reflected in the judgment logic of the preset content security summary model. The content layer security features of each dimension have different degrees of importance for the risk judgment of the current data shard. Among them, when a certain feature value contributes significantly to the judgment result, the attention weight corresponding to that feature will also increase accordingly.
[0150] Step S24: Determine the content layer risk score of the data fragment based on the content security summary value using a preset risk assessment function.
[0151] It should be noted that the preset risk assessment function can be a non-linear mapping function used to map content security summary values to preset risk ranges. In this embodiment, the preset risk assessment function can be in the form of a Sigmoid function. The Sigmoid function is a common S-shaped activation function that can map any real value to a probability value between 0 and 1. Specifically, it can be expressed as: ,in, As the risk score weight, This is for risk scoring bias.
[0152] In practical applications, after calculating the content security digest value of the data fragment, the device can input the content security digest value into a preset risk assessment function to obtain a content risk score:
[0153] in, Indicates the risk score weight. This indicates risk score bias. Indicates the first The content risk score for each data segment is such that the closer the value is to 1, the higher the risk of content poisoning.
[0154] It should be noted that the parameters of the preset content security digest model and the preset risk assessment function in this embodiment are... The training samples are pre-trained using labeled samples during the offline training phase. The training samples are represented as follows: ,in This is the normalized content security feature vector. This represents a risk label for sharding (0 for trusted shards, 1 for poisoning or abnormal shards). In practical applications, the device can use a binary cross-entropy loss function to train the model parameters, and after training, fix all model parameters and deploy the device. During the online detection phase, the device only participates in the forward computation and does not update the parameters. The binary cross-entropy loss function can be expressed as:
[0155] In this embodiment, data fragments are decoded and segmented, and abnormal bytes are retained and mapped to abnormal character identifiers. Then, based on the segmentation sequence and abnormal character identifiers, content layer security features including segmentation distribution skewness, long-tail rare segmentation ratio, abnormal character ratio, abnormal repeated fragment density features, and semantic distribution drift features are extracted. Subsequently, a content security summary value is determined based on these features using a preset content security summary model, and a content layer risk score is determined based on the content security summary value using a preset risk assessment function. This enables low-latency, non-blocking quantitative assessment of content layer poisoning risk under the condition that the RDMA data content is not fully visible. This avoids the problem of detection failure caused by the inability to capture complete data content in the RDMA direct write path, which is a problem in traditional solutions. At the same time, it avoids the transmission blockage caused by introducing a heavy detection model on the high-speed transmission path. Thus, it can achieve pre-identification of training data poisoning risk while ensuring RDMA transmission performance.
[0156] In the specific implementation, refer to Figure 5 , Figure 5 This is a flowchart illustrating the content risk scoring process in the data detection method of this application. Figure 5 As shown, firstly, the device can acquire RDMA data fragments and perform decoding and word segmentation on the content of the data fragments. During the decoding process, if a byte sequence that does not conform to the encoding standard is encountered, the byte can be mapped to a preset special symbol as an abnormal character identifier. Then, the successfully decoded characters and the mapped abnormal character identifiers are combined to form a character sequence, and word segmentation is performed on the character sequence according to the word segmenter's vocabulary, splitting the continuous character sequence into independent tokens to obtain a word segmentation sequence. Subsequently, the device can extract content layer security features based on the results of decoding and word segmentation, including five dimensions: token distribution bias, proportion of long-tail rare tokens, proportion of abnormal characters / garbled characters, density features of repeated fragments, and semantic distribution drift features. The device also obtains the mean and standard deviation of each feature in the trusted corpus fragment set, and normalizes each content layer security feature to unify features of different dimensions to a comparable scale, resulting in a normalized content layer security feature vector. Then, the device can input the normalized content layer security feature vector into the security prediction calculation model to achieve lightweight attention fusion, generate a segmented security summary, and input the segmented security summary into a preset risk assessment function. Through nonlinear mapping, the segmented security summary value is mapped to the range of 0 to 1, and finally the content layer risk score is obtained and output.
[0157] In this embodiment, by collecting transport layer metadata from the RDMA transmission process and constructing transport layer feature vectors corresponding to data fragments, and determining the directed dependencies between each transport layer security feature node based on the inherent service constraints and execution order of the RDMA transmission protocol, a transport layer semantic probability graph for the current training task is constructed. Then, a transport layer trust benchmark is constructed based on the historical transport layer feature vector set and the transport layer semantic probability graph of the trusted transmission stage. Finally, the transport layer feature vector of the current data fragment is compared with the transport layer trust benchmark to determine the transport layer risk score. Thus, without blocking the high-speed RDMA transmission path, a precise quantitative assessment of the deviation of the transport layer behavior of data fragments can be achieved using lightweight transport layer metadata. This effectively avoids the problem of a large number of false alarms generated by static threshold strategies under physical disturbances such as wide area network link jitter and packet loss retransmission. Meanwhile, since the transport layer credible benchmark is generated by statistical analysis of credible historical samples from the current training task itself, and the conditional probability distribution of each feature node under each combination of values of the parent node is clearly recorded in the form of a conditional probability table, the transport layer risk score of each data segment has a traceable statistical benchmark and probability basis as support, thereby ensuring the interpretability and verifiability of transport layer anomaly judgment.
[0158] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6 , Figure 6 This is a flowchart illustrating the data detection method of this application in Embodiment 3.
[0159] In this embodiment, step S30 further includes steps S31 to S34: Step S31: Generate a joint increase index and a contradiction index for the data fragment based on the transport layer risk score and the content layer risk score. The joint increase index is used to characterize the degree to which the transport layer risk score and the content layer risk score increase synchronously, and the contradiction index is used to characterize the degree of deviation between the transport layer risk score and the content layer risk score.
[0160] It should be noted that the joint increase index can be a quantitative indicator used to characterize the degree to which the transport layer risk score and the content layer risk score increase simultaneously. In this embodiment, the device can multiply the transport layer risk score and the content layer risk score to obtain the joint increase index. The corresponding calculation formula can be:
[0161] in, Indicates the first Transport layer risk scoring for each data fragment. Indicates the first Risk scoring of content layer for each data shard; In practical applications, when and When both rise, A significant increase indicates that the data fragment exhibits both behavioral deviations at the transmission level and the risk of poisoning at the content level, demonstrating cross-plane collaborative anomalies and representing a relatively clear indication of an attack.
[0162] It should also be noted that the contradiction index can be a quantitative indicator used to characterize the degree of deviation between the transport layer risk score and the content layer risk score. In this embodiment, the device can use the absolute value of the difference between the transport layer risk score and the content layer risk score as the contradiction index, and the corresponding calculation formula can be:
[0163] in, The larger this value, the greater the difference in risk assessment between the two planes, requiring a more stringent access review process.
[0164] In this embodiment, the contradiction index can be used to identify two types of covert situations: (1) the transmission behavior is credible but the content is abnormal, that is, the transmission layer risk score is low while the content layer risk score is high, indicating that the attacker may use a legitimate transmission channel to smuggle poisoned content; (2) the transmission behavior is abnormal but the content has not yet shown obvious abnormalities, that is, the transmission layer risk score is high while the content layer risk score is low, indicating that it may be a link disturbance or a transmission layer attack attempt but the content has not yet been tampered with.
[0165] Step S32: Determine the transmission-side anomaly factor score of the data fragment based on the transport layer metadata.
[0166] It should be noted that the transport-side anomaly factor score can be a quantitative score used to describe the degree of anomaly of data fragments in various transport-side dimensions. It can be calculated separately based on the data of each dimension in the transport layer metadata, and may specifically include connection identity anomaly scores. Source authentication anomaly score Memory access exception scoring And queue event exception scoring .
[0167] Among them, identity risk scoring Used to characterize the degree of anomaly in the connection identity dimension of data sharding, it can be based on connection identity features (including source node legitimacy features) in the transport layer feature vector. Queue for Authorization Features and queue state characteristics The calculation is obtained by means of the formula:
[0168] Source certification risk score Used to characterize the degree of anomaly in the source authentication dimension of data sharding, it can be based on the source authentication features (including task ID matching features) in the transport layer feature vector. Object version consistency feature and list signature verification features The calculation is obtained by means of the formula:
[0169] Memory access risk score Used to characterize the degree of anomaly in the memory access dimension of data sharding, it can be based on memory access features (including memory region authorization features) in the transport layer feature vector. , Write address range validity features Consistency features with fragment length The calculation is obtained by means of the formula:
[0170] Queue event risk scoring Used to characterize the degree of anomaly in data sharding at the queue event dimension, it can be based on queue event features in the transport layer feature vector (including completion status features of completed queue entries). Features of retransmission count The calculation yields the result, and the corresponding formula is:
[0171] Step S33: Based on the transmission-side anomaly factor score and the content layer risk score, determine the correlation risk indicators between each transmission-side anomaly factor and content anomaly. The correlation risk indicators include identity correlation risk indicators, source correlation risk indicators, memory correlation risk indicators, and queue correlation risk indicators.
[0172] Understandably, transport-side anomaly factors can be various dimensions of factors in the transport layer that may cause anomalies. These can include identity anomaly factors, source authentication anomaly factors, memory access anomaly factors, and queue event anomaly factors, corresponding to the dimensions of the four transport-side anomaly factor scoring methods mentioned above. Specifically, identity anomaly factors involve the legitimacy of the connection establishment subject, connection object, and operational behavior, such as whether the source node is authorized, whether the queue pair is authorized, and whether the queue pair status is normal. Source authentication anomaly factors involve the legitimacy of the business context to which the data shard belongs, such as whether the task ID matches, whether the object version is consistent, and whether the list signature passes verification. Memory access anomaly factors involve the legitimacy of the remote node's access to the memory region, such as whether the memory region is authorized, whether the write address range is out of bounds, and whether the shard length meets expectations. Queue event anomaly factors involve the normal operation status of the queue pair and the completion queue, such as whether the completion queue entries are completed normally and whether the number of retransmissions is within a reliable range. These four transport-side anomaly factors characterize abnormal situations at the transport layer from four dimensions: transport connection, data source, memory operation, and queue event, providing dimensional quantitative values of the degree of anomaly for the subsequent calculation of associated risk indicators.
[0173] It should also be noted that the associated risk indicator can be a quantitative indicator used to characterize the coupling relationship between various transmission-side anomaly factors and content anomalies. It can be obtained by correlating the scores of each transmission-side anomaly factor with the content-layer risk score, and is used to determine whether transmission-side anomalies and content anomalies occur simultaneously, thereby identifying the cross-plane collaborative characteristics of data poisoning attacks. In this embodiment, the associated risk indicator... It can be based on identity-related risk indicators Source-related risk indicators Memory-related risk indicators Risk indicators associated with queues It is formed by aggregation, that is:
[0174] It should be noted that the identity association risk indicator can be used to determine whether connection identity anomalies and content anomalies occur simultaneously. In this embodiment, the device can score connection identity anomalies. Content layer risk scoring Multiplying them together yields an identity-related risk indicator. The corresponding calculation formula can be:
[0175] in, In practical applications, when both connection identity risk and content risk increase simultaneously, this value rises, indicating that the current shard may originate from an illegal connection, an abnormal queue pair, or a forged transmission channel, and carries abnormal training content. Therefore, it is more likely to be a data poisoning or illegal injection behavior during the synchronization phase.
[0176] It should also be noted that the source association risk indicator can be a correlation risk indicator used to characterize the degree of correlation risk between source authentication anomalies and content anomalies. In this embodiment, the device can score the source authentication risk. Content layer risk scoring Multiplying them together yields the source-related risk indicator. The corresponding calculation formula can be:
[0177] in, In practical applications, when both the source authentication risk score and the content layer risk score rise simultaneously, the value of this indicator increases significantly, indicating that the current data shards may originate from unauthorized data source injection, replay of old version objects, or abnormal content mixed in with a tampered list, and are therefore more likely to be a data poisoning attack.
[0178] It should be noted that the memory association risk indicator can be a correlation risk indicator used to characterize the degree of correlation risk between abnormal memory access behavior and content anomalies. In this embodiment, the device can score memory access risks. Content layer risk scoring Multiplying them together yields a memory-related risk indicator. The corresponding calculation formula can be:
[0179] in, In practical applications, when both the memory access risk score and the content layer risk score increase simultaneously, the value of this indicator increases significantly, indicating that the current data sharding may be subject to attacks such as overwrite, out-of-bounds write, replacement write, or shard splicing injection, and is therefore more likely to be a data tampering or poisoning injection.
[0180] It should also be noted that the queue-related risk index can be a correlation risk index used to characterize the degree of correlation risk between queue event anomalies and content anomalies. In this embodiment, the device can score queue event risks. Content layer risk scoring Multiplying them together yields the queue-related risk index. The corresponding calculation formula can be:
[0181] in, In practical applications, when both the queue event risk score and the content layer risk score increase simultaneously, the value of this indicator increases significantly, indicating that the current data sharding may be due to abnormal timing insertion, delayed replay, out-of-synchronization window writing, or data pollution under the cover of link anomalies.
[0182] Step S34: Aggregate the joint elevation index, the contradiction index, and the correlation risk index into the cross-plane correlation inference index.
[0183] In practical applications, the equipment can combine the elevation index Contradictory Indicators and identity-related risk indicators Source-related risk indicators Memory-related risk indicators Risk indicators associated with queues These are aggregated into cross-plane correlation inference metrics for the current data shard, which are used as the basis for determining the security status of the current data shard in subsequent anomaly detection steps.
[0184] Further, the step of performing anomaly detection on the data shards based on the cross-plane association inference index and obtaining anomaly detection results includes: Step S41: Determine the maximum association risk value among all association risk indicators in the cross-plane association inference index.
[0185] Understandably, the maximum association risk value can be the largest value selected from the association risk indicators (i.e., identity association risk indicator, source association risk indicator, memory association risk indicator, and queue association risk indicator) included in the cross-plane association inference indicator set. It can be used to represent the most significant cross-plane association anomaly strength in the current data shard. In this embodiment, the maximum association risk value can be expressed as:
[0186] Step S42: Determine the indicator type corresponding to the maximum associated risk value as the dominant associated type.
[0187] It should be understood that the indicator type can be the category to which each associated risk indicator belongs in the cross-plane association inference indicator set. This can include identity association type (corresponding to identity association risk indicators), source association type (corresponding to source association risk indicators), memory association type (corresponding to memory association risk indicators), and queue association type (corresponding to queue association risk indicators). In practical applications, each indicator type corresponds to a relationship between a transmission-side anomaly factor and a content anomaly. It can be used to identify which associated risk indicator the maximum associated risk value specifically originates from, thereby determining which transmission-side factor is more likely to dominate the anomaly of the current data shard.
[0188] It should also be understood that the dominant association type can be the indicator type of the association risk index corresponding to the maximum association risk value. It can be used to indicate which type of transport-side factor is more likely to dominate cross-plane anomalies in the current data shard, including four types: connection identity dominant, source authentication dominant, memory access dominant, and queue event dominant. In this embodiment, the dominant association type can be represented as:
[0189] in, Used to represent the strength of the most significant cross-plane correlation anomaly in the current partition. This is used to indicate which type of transport-side factor is more likely to dominate the anomaly.
[0190] Step S43: Determine the security status level of the data shard based on the joint elevation index, the contradiction index, and the maximum associated risk value.
[0191] Understandably, the security status level can be the security level output after comprehensively judging the current data fragment based on joint escalation indicators, contradictory indicators, and the maximum associated risk value. It can include four levels: high risk, conflict, suspicious, and safe. Among them, high risk means that the transmission risk and content risk are increased at the same time, and this joint anomaly can be explained by at least one type of transmission-side factor, indicating that the current data fragment has strong cross-plane collaborative poisoning characteristics; conflict means that the risk judgments of the transmission plane and the content plane are significantly inconsistent, and there is a correlation between a certain type of transmission-side factor and the content anomaly, indicating that the current data fragment may be a case of covert poisoning or abnormal content hidden under the guise of legitimate transmission; suspicious means that the current data fragment has shown any of the joint anomalies, contradictory anomalies, or local associated anomalies, but the evidence is not yet sufficient to directly determine it as a high-risk or conflict state; safe means that none of the above anomaly conditions have been triggered, and the current data fragment can enter the training sample queue.
[0192] In this embodiment, the security status level inference rule can be:
[0193] in, Indicates a high-risk level. Indicates the conflict level. Indicates a suspicious level. Indicates the security level.
[0194] In practical applications, the device can obtain the joint elevation index of the current data shard. Contradictory Indicators and maximum associated risk value And obtain the preset joint anomaly threshold. Contradictory anomaly threshold and associated anomaly threshold Then, the security status level corresponding to the data shard is determined according to the security status level inference rules.
[0195] Step S44: Determine the target correction strategy from the candidate correction strategies based on the dominant association type, wherein the candidate correction strategies correspond to different transmission-side anomaly factor types.
[0196] It should be understood that candidate correction strategies can be pre-set multiple correction action schemes corresponding to different types of transmission-side anomaly factors. In this embodiment, candidate correction strategies can include four types, corresponding to four dominant association types: connection identity anomaly, source authentication anomaly, memory access anomaly, and queue event anomaly. Specifically, the correction strategy for connection identity anomaly is to review the source node authorization, QP number, and QP state migration, and freeze the abnormal connection or reject the corresponding transmission channel; the correction strategy for source authentication anomaly is to re-verify the task identifier, object version, source signature, and Manifest binding relationship, and roll back to the trusted object version if necessary; the correction strategy for memory access anomaly is to check the MR identifier, write address, write length, and fragment boundaries, and isolate fragments with risks of overwrite, out-of-bounds write, or replacement write; the correction strategy for queue event anomaly is to check the completion queue status, retransmission count, fragment arrival order, and synchronization window boundaries, and suspend synchronization windows with risks of abnormal timing insertion or delayed replay.
[0197] It should also be understood that the target correction strategy can be a specific correction action plan determined by matching the candidate correction strategies based on the dominant association type. In practical applications, after determining the dominant association type of the current data fragment, the device can select the correction strategy corresponding to that dominant association type from the four candidate correction strategies as the target correction strategy.
[0198] In practical applications, the device can identify a dominant association type and select the corresponding correction strategy from a pre-stored pool of candidate correction strategies as the target correction strategy. If the dominant association type is connection identity-driven, the target correction strategy is to verify the source node authorization, QP number, and QP state migration, and freeze abnormal connections or reject the corresponding transmission channel. If the dominant association type is source authentication-driven, the target correction strategy is to re-verify the task identifier, object version, source signature, and Manifest binding relationship, and roll back to the trusted object version if necessary. If the dominant association type is memory access-driven, the target correction strategy is to check the MR identifier, write address, write length, and fragment boundaries, and isolate fragments with risks of overwrite, out-of-bounds writes, or replacement writes. If the dominant association type is queue event-driven, the target correction strategy is to check the completion queue status, retransmission count, fragment arrival order, and synchronization window boundaries, and pause synchronization windows with risks of abnormal timing insertions or delayed replays.
[0199] Step S45: Generate anomaly detection results based on the security status level and the target correction strategy.
[0200] In this embodiment, the device can combine the determined security status level, the determined target correction strategy, and the dominant correlation type into the anomaly detection result of the current data shard. The corresponding expression can be represented as:
[0201] in, Indicates the safe state of the fragment. Indicates the dominant association type, This indicates the corresponding corrective action.
[0202] In this embodiment, by determining the maximum associated risk value and its corresponding dominant association type among the associated risk indicators in the cross-plane association inference index, and then comprehensively judging the security status level of the data fragment based on the joint escalation index, contradiction index, and maximum associated risk value, and matching the target correction strategy based on the dominant association type, an anomaly detection result containing the security status level and the target correction strategy is finally generated. This allows the cross-plane association inference results between the transport layer and the content layer to be transformed into executable security decisions. Since both the security status level and the correction strategy are directly related to the dominant association type, the interpretability of the anomaly detection results and the pertinence of the correction actions can be guaranteed.
[0203] In the specific implementation, refer to Figure 7 , Figure 7 This is a flowchart illustrating the joint inference process in the data detection method of this application. For example... Figure 7As shown, firstly, the device can obtain the transport layer risk score output by module one and the content layer risk score output by module two, and construct cross-plane security features based on the transport layer risk score and the content layer risk score. Specifically, the device can combine the transport layer risk score, the content layer risk score, and the transport-side anomaly factor scores of various dimensions calculated based on transport layer metadata into a cross-plane security feature vector for the current data fragment, which includes multiple dimensions such as transport layer risk score, content layer risk score, connection identity anomaly score, source authentication anomaly score, memory access anomaly score, and queue event anomaly score. In addition, the device can calculate joint escalation indicators and contradiction indicators based on the transport layer risk score and the content layer risk score, and then calculate identity association risk indicators, source association risk indicators, memory association risk indicators, and queue association risk indicators based on the product of the connection identity anomaly score, source authentication anomaly score, memory access anomaly score, and queue event anomaly score with the content layer risk score, respectively. Finally, the joint escalation indicator, contradiction indicator, and the four association risk indicators are used together as cross-plane association inference indicators. Then, the device can determine the maximum associated risk value from four associated risk indicators and identify the indicator type corresponding to this maximum associated risk value as the dominant association type, indicating which type of transmission-side factor is most likely to dominate the current data fragmentation anomaly. Finally, the device can determine the security status level of the current data fragmentation according to preset judgment rules based on the joint escalation indicator, contradictory indicator, and maximum associated risk value, and match the target correction strategy from candidate correction strategies according to the dominant association type to generate corresponding correction actions. Finally, it outputs anomaly detection results containing the security status level and the target correction strategy.
[0204] In this embodiment, the overall relationship between the risks of the transport layer and the content layer is characterized by generating joint elevation indicators and contradiction indicators. At the same time, the transport-side anomaly factor score is determined based on the transport layer metadata, and the scores of each transport-side anomaly factor are further correlated with the content layer risk score to obtain multiple types of associated risk indicators. Finally, all indicators are aggregated into a cross-plane correlation inference indicator, which enables the anomaly judgment of data fragmentation to comprehensively consider the synchronous elevation relationship between transport and content, the contradictory relationship between transport and content, and the coupling relationship between each transport-side anomaly factor and content anomaly. This can effectively distinguish between occasional fluctuations in the transport layer caused by WAN link disturbances and cross-plane coordinated anomalies caused by data poisoning attacks, thereby improving the accuracy of data poisoning detection.
[0205] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the data detection method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0206] This application also provides a data detection device, please refer to... Figure 8 The data detection device includes: The transport layer scoring module 10 is used to obtain transport layer metadata of data fragments in the current training task during RDMA transmission, and to determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer. The content layer scoring module 20 is used to extract the content layer security features of the data slice and generate a content layer risk score for the data slice based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data slice at the content level. The scoring aggregation module 30 is used to perform cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane association inference index for the data shard. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score. The anomaly detection module 40 is used to perform anomaly detection on the data shards based on the cross-plane association reasoning index and obtain anomaly detection results.
[0207] The data detection device provided in this application, employing the data detection method described in the above embodiments, can solve the technical problem of low accuracy in existing data poisoning detection methods that rely on link layer indicators. Compared with the prior art, the beneficial effects of the data detection device provided in this application are the same as those of the data detection method provided in the above embodiments, and other technical features in the data detection device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0208] This application provides a data detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data detection method in Embodiment 1 above.
[0209] The following is for reference. Figure 9 The diagram illustrates a structural schematic of a data detection device suitable for implementing embodiments of this application. The data detection device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The data detection device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0210] like Figure 9 As shown, the data detection device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the data detection device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the data detection device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show data detection devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0211] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0212] The data detection device provided in this application, employing the data detection method described in the above embodiments, can solve the technical problem of data detection. Compared with the prior art, the beneficial effects of the data detection device provided in this application are the same as those of the data detection method provided in the above embodiments, and other technical features of the data detection device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0213] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0214] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0215] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the data detection method in the above embodiments.
[0216] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0217] The aforementioned computer-readable storage medium may be included in the data detection device; or it may exist independently and not be assembled into the data detection device.
[0218] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the data detection device, the data detection device performs the following actions: acquires transport layer metadata of data fragments during RDMA transmission of the current training task, and determines a transport layer risk score for the data fragments based on the transport layer metadata. The transport layer risk score characterizes the degree of behavioral deviation of the data fragments at the transport layer. It then extracts content layer security features of the data fragments and generates a content layer risk score based on these features. The content layer risk score characterizes the degree of poisoning risk of the data fragments at the content layer. Finally, it performs cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane correlation inference index for the data fragments. This cross-plane correlation inference index characterizes the correlation between the transport layer risk score and the content layer risk score. Finally, it performs anomaly detection on the data fragments based on the cross-plane correlation inference index to obtain anomaly detection results.
[0219] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0220] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0221] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0222] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described data detection method. This addresses the technical problem of low accuracy in existing data poisoning detection methods that rely on link-layer indicators. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the data detection method provided in the above embodiments, and will not be elaborated upon here.
[0223] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A data detection method, characterized in that, The method includes: Obtain transport layer metadata of data fragments during RDMA transmission for the current training task, and determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer. Extract the content layer security features of the data slices, and generate a content layer risk score for the data slices based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data slices at the content level. The transport layer risk score and the content layer risk score are aggregated across planes to generate a cross-plane association inference index for the data shards. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score. Anomaly detection is performed on the data shards based on the cross-plane association reasoning index to obtain anomaly detection results.
2. The method as described in claim 1, characterized in that, The step of determining the transport layer risk score of the data fragment based on the transport layer metadata includes: Construct the transport layer feature vector corresponding to the data fragment based on the transport layer metadata; Based on the constraints in the RDMA transmission process, a directed dependency relationship is determined between several transport layer security feature nodes. Construct the transport layer semantic probability graph of the current training task based on each of the transport layer security feature nodes and the directed dependency relationships; Based on the set of historical transport layer feature vectors generated by the current training task in the trusted transport phase and the transport layer semantic probability graph, a transport layer trusted benchmark for the current training task is constructed. The transport layer trusted benchmark is used to characterize the conditional dependencies and normal value ranges between the security features of each transport layer under the trusted transport state. The transport layer risk score of the data fragment is determined based on the transport layer feature vector and the transport layer trust benchmark.
3. The method as described in claim 2, characterized in that, The step of constructing a transport layer trust benchmark for the current training task based on the set of historical transport layer feature vectors generated during the trusted transport phase of the current training task and the transport layer semantic probability graph includes: Based on the set of historical transport layer feature vectors generated by the current training task in the trusted transmission phase, the conditional probability of each transport layer security feature node in the transport layer semantic probability graph under the corresponding parent node value combination is calculated. Based on the conditional probabilities, construct a conditional probability table for each of the transport layer security feature nodes; The transport layer reliability benchmark for the current training task is determined based on the conditional probability table.
4. The method as described in claim 2, characterized in that, The step of determining the transport layer risk score of the data fragment based on the transport layer feature vector and the transport layer trust benchmark includes: The joint logarithmic matching value of the data fragment is determined based on the transport layer feature vector and the transport layer trust benchmark; Determine the average matching value and low matching boundary of historical reliable samples; The transport layer risk score of the data fragment is determined based on the joint logarithmic matching value, the average matching value, and the low matching boundary.
5. The method as described in claim 1, characterized in that, The step of extracting the content layer security features of the data fragments and generating a content layer risk score for the data fragments based on the content layer security features includes: The data segments are decoded and segmented to identify decoded abnormal bytes in the data segments and map the decoded abnormal bytes to abnormal character identifiers. A word segmentation sequence is constructed based on the normal bytes and the abnormal character identifiers in the data segment, and the content layer security features of the data segment are extracted based on the word segmentation sequence. The content layer security features include word segmentation distribution skewness, proportion of long-tail rare word segmentation, proportion of abnormal characters, density features of abnormal repeated segments, and semantic distribution drift features. The content security summary value of the data segment is determined by a preset content security summary model based on the word segmentation distribution skewness, the proportion of long-tail rare word segmentation, the proportion of abnormal characters, the density feature of abnormal repeated segments, and the semantic distribution drift feature. The content layer risk score of the data fragment is determined based on the content security summary value using a preset risk assessment function.
6. The method according to any one of claims 1 to 5, characterized in that, The step of cross-plane aggregation of the transport layer risk score and the content layer risk score to generate the cross-plane correlation inference index of the data fragment includes: The data sharding is generated based on the transport layer risk score and the content layer risk score. The joint increase index is used to characterize the degree to which the transport layer risk score and the content layer risk score increase synchronously, and the contradiction index is used to characterize the degree of deviation between the transport layer risk score and the content layer risk score. The transmission-side anomaly factor score of the data fragment is determined based on the transport layer metadata; Based on the transmission-side anomaly factor score and the content-layer risk score, the correlation risk indicators between each transmission-side anomaly factor and content anomaly are determined. The correlation risk indicators include identity correlation risk indicators, source correlation risk indicators, memory correlation risk indicators, and queue correlation risk indicators. The joint elevation index, the contradiction index, and the associated risk index are aggregated into the cross-plane association inference index.
7. The method as described in claim 6, characterized in that, The step of performing anomaly detection on the data shards based on the cross-plane association inference index and obtaining anomaly detection results includes: Determine the maximum associated risk value among all associated risk indicators in the cross-plane association inference index; The indicator type corresponding to the maximum associated risk value is determined as the dominant associated type; The security status level of the data shard is determined based on the joint escalation index, the contradiction index, and the maximum associated risk value. The target correction strategy is determined from the candidate correction strategies based on the dominant correlation type, and the candidate correction strategies correspond to different transmission-side anomaly factor types. Anomaly detection results are generated based on the security status level and the target correction strategy.
8. A data detection device, characterized in that, The device includes: The transport layer scoring module is used to obtain transport layer metadata of data fragments in the current training task during RDMA transmission, and to determine the transport layer risk score of the data fragments based on the transport layer metadata. The transport layer risk score is used to characterize the degree of deviation of the behavior of the data fragments at the transport layer. The content layer scoring module is used to extract the content layer security features of the data slice and generate a content layer risk score for the data slice based on the content layer security features. The content layer risk score is used to characterize the degree of poisoning risk of the data slice at the content level. The scoring aggregation module is used to perform cross-plane aggregation of the transport layer risk score and the content layer risk score to generate a cross-plane association inference index for the data shard. The cross-plane association inference index is used to characterize the association relationship between the transport layer risk score and the content layer risk score. An anomaly detection module is used to perform anomaly detection on the data shards based on the cross-plane association reasoning index and obtain anomaly detection results.
9. A data detection device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data detection method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the data detection method as described in any one of claims 1 to 7.