Data detection method and device, equipment, storage medium and program product
By allocating data between the central node and the distributed node and using the privacy data detection model for stain and semantic analysis, the problem of privacy data leakage and excessive resource consumption in the TEE environment is solved, efficient and accurate privacy data detection is achieved, and system performance and security are improved.
Patent Information
- Application Number
- CN202510669056.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
AI Technical Summary
Prior Art In the TEE environment, the design defects of the application lead to the leakage of privacy data, and the stain analysis technology consumes too much resources when processing large-scale data sets, affecting system performance.
By allocating the data to be detected between the central node and the distributed node, using the privacy data detection model to perform stain analysis and semantic analysis of non-text and text data, combining data compression and zero-knowledge proof, reducing communication and computing overhead, and optimizing resource allocation, using trusted hardware and data change trend prediction models to improve detection accuracy.
It improves the accuracy of privacy data leakage detection, reduces communication and computing overhead, improves system resource utilization and performance, and ensures the security of user privacy data.
Smart Images

Figure CN120493311A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data security technology, and in particular relates to a data detection method, apparatus, device, storage medium and program product. Background Art
[0002] With the rapid development of Internet technology, the need to protect private data is becoming increasingly urgent. To ensure the security of private data, Trusted Execution Environment (TEE) has gradually become one of the core technologies for protecting private data.
[0003] TEE allows applications to run in an isolated environment, preventing external attackers or unauthorized third parties from accessing the data, thus improving data security. However, despite the strong data security provided by TEE, the design of the application itself can still lead to the leakage of private data, reducing data security. Summary of the Invention
[0004] The embodiments of the present application provide a data detection method, apparatus, device, storage medium, and program product, which can improve the accuracy of detecting privacy data leakage, thereby improving data security.
[0005] In a first aspect, an embodiment of the present application provides a data detection method, applied to a distributed node, comprising:
[0006] Receiving data to be detected from a central node in real time, wherein the data to be detected includes non-text data and text data;
[0007] Performing a taint analysis on the non-text data to obtain a first detection result, where the first detection result indicates whether tainted data in the non-text data has been leaked;
[0008] Inputting the text data into a privacy data detection model, performing semantic analysis on the text data using the privacy data detection model, and obtaining a second detection result output by the privacy data detection model, wherein the second detection result indicates whether private data in the text data has been leaked;
[0009] Feedback the first detection result and the second detection result to the central node.
[0010] In one possible implementation, inputting the text data into a privacy-related data detection model, performing semantic analysis on the text data using the privacy-related data detection model, and obtaining a second detection result output by the privacy-related data detection model includes:
[0011] Inputting the text data into the privacy data detection model, performing feature extraction on the text data through a feature extraction network in the privacy data detection model to obtain a text feature vector;
[0012] The text feature vector is semantically analyzed by the context analysis network in the privacy data detection model to determine whether the privacy data in the text data has been leaked, and the second detection result is output.
[0013] In a possible implementation, the second detection result includes a probability score of leakage of the private data in the text data; and before feeding back the first detection result and the second detection result to the central node, the method further includes:
[0014] When the likelihood score is greater than a first preset threshold, extracting preset indicator feature data;
[0015] Calculating the distance between the text data and a decision boundary of a normal behavior model using a support vector machine to obtain an anomaly score, wherein the normal behavior model is trained using preset indicator feature data indicating no privacy data leakage behavior;
[0016] When the abnormality score is greater than a second preset threshold, it is determined that privacy data leakage exists in the text data.
[0017] In a second aspect, an embodiment of the present application provides a data detection method, applied to a central node, the method comprising:
[0018] Sending the data to be detected to each distributed node, so that the distributed node performs a taint analysis on non-text data in the data to be detected to obtain a first detection result, and performing a semantic analysis on the text data in the data to be detected using a privacy data detection model to obtain a second detection result output by the privacy data detection model;
[0019] Receive the first detection result and the second detection result sent by each distributed node.
[0020] In a possible implementation, before sending the data to be detected to each distributed node, the method further includes:
[0021] Using a data change trend prediction model to predict the indicators to be detected that have data changes within a preset time window, the data change trend prediction model is trained using historical data of the preset indicators;
[0022] Data corresponding to the indicator to be detected is acquired from a target distributed node to obtain the data to be detected, where a preset container set pod is deployed on the target distributed node.
[0023] In a possible implementation, before using the data change trend prediction model to predict the indicator to be detected that has undergone data changes within a preset time window, the method further includes:
[0024] Obtaining resource utilization of each distributed node, affinity between the preset pod and each distributed node, and preset trusted hardware supported by each distributed node;
[0025] Calculate a security score for each distributed node based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node;
[0026] Allocate the preset pod to the trusted distributed node with the highest security score;
[0027] Acquiring historical resource data and historical central processing unit (CPU) load data of the trusted distributed node;
[0028] Inputting the historical resource data of the trusted distributed node and the historical CPU load data into a load prediction model to obtain a first prediction result, wherein the second prediction result includes CPU load data within a preset time period in the future;
[0029] Allocate computing resources and storage resources to the trusted distributed node based on the first prediction result.
[0030] In one possible implementation, calculating the security score of each distributed node based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node includes:
[0031] For each computing node, calculating a resource utilization score of the computing node according to the resource utilization of the computing node;
[0032] For each computing node, calculate the affinity score of the computing node based on the affinity between the preset pod and the computing node;
[0033] For each computing node, calculating a trusted hardware support score of the computing node according to whether the computing node supports preset trusted hardware;
[0034] For each computing node, performing a weighted summation of the resource utilization score, the affinity score, and the trusted hardware support score to obtain a security score of the computing node;
[0035] In a third aspect, an embodiment of the present application provides a data detection device, applied to a distributed node, comprising:
[0036] A receiving module, configured to receive data to be detected sent by a central node in real time, wherein the data to be detected includes non-text data and text data;
[0037] an analysis module, configured to perform a taint analysis on the non-text data to obtain a first detection result, wherein the first detection result indicates whether taint data in the non-text data has been leaked;
[0038] a detection module, configured to input the text data into a privacy data detection model, perform semantic analysis on the text data using the privacy data detection model, and obtain a second detection result output by the privacy data detection model, wherein the second detection result indicates whether private data in the text data has been leaked;
[0039] A feedback module is used to feed back the first detection result and the second detection result to the central node.
[0040] In a fourth aspect, an embodiment of the present application provides a data detection device, applied to a central node, comprising:
[0041] a sending module, configured to send the data to be detected to each distributed node, so that the distributed node performs a taint analysis on the non-text data in the data to be detected to obtain a first detection result, and performs a semantic analysis on the text data in the data to be detected using a privacy data detection model to obtain a second detection result output by the privacy data detection model;
[0042] A receiving module is used to receive the first detection result and the second detection result sent by each distributed node.
[0043] In a fifth aspect, an embodiment of the present application provides a terminal device, the device comprising: a processor and a memory storing computer program instructions;
[0044] When the processor executes the computer program instructions, the data detection method of the first aspect is implemented, or the data detection method of the second aspect is implemented.
[0045] In a sixth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for data detection as in the first aspect is implemented, or the method for data detection as in the second aspect is implemented.
[0046] In the seventh aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by the processor of an electronic device, the electronic device executes the data detection method as in the first aspect, or implements the data detection method as in the second aspect.
[0047] According to a method, apparatus, device, storage medium and program product for data detection in an embodiment of the present application, after a distributed node receives the data to be detected sent by a central node, the distributed node can directly perform privacy data leakage detection for the non-text data in the data to be detected, and determine the first detection result of the tainted data in the non-text data. For text data, a privacy data detection model can be used to perform semantic analysis on the text data, so as to determine whether there is tainted data in the text data and whether the tainted data has been leaked. In this way, using the privacy data detection model, it is possible to accurately determine whether there is tainted data in the text data in each computing node and whether the tainted data has been leaked, thereby accurately determining whether there is privacy data leakage in each distributed node, thereby improving the accuracy of the detection results and ensuring the security of the user's privacy data. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0049] Figure 1 1 is a flow chart of a method for data detection applied to a central node provided in an embodiment of the present application;
[0050] Figure 2 This is a flow chart of a resource allocation method provided in an embodiment of the present application;
[0051] Figure 3 1 is a flow chart of a method for data detection applied to distributed nodes provided in an embodiment of the present application;
[0052] Figure 4 is an exemplary schematic diagram of a data monitoring method provided in an embodiment of the present application;
[0053] Figure 5 Schematic diagram of the structure of a device for data detection applied to distributed nodes provided in an embodiment of the present application;
[0054] Figure 6 Schematic diagram of the structure of a device for data detection applied to a central node provided in an embodiment of the present application;
[0055] Figure 7It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0057] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0058] With the development of the e-commerce industry, the demand for privacy protection is also increasing. To protect users' privacy data from attacks when using e-commerce platforms, the deployment of a trusted execution environment (TEE) can ensure the security of applications. TEEs allow applications to run in an isolated environment, preventing external attackers or unauthorized third parties from accessing the data. However, due to inherent design flaws in applications, user privacy data can still be leaked.
[0059] Currently, to prevent the leakage of user private data, various input sources of private data can be identified for detection. Specifically, these input sources can include application input parameters, file read data, or network data. The system uses the points where the application outputs external data as privacy leakage detection points. The system then analyzes the processing statements in these input sources, such as assignment statements, copy statements, and return statements, to determine whether private data is likely to be transmitted to the privacy leakage detection point through these statements. When the system detects that private data has propagated to the privacy leakage detection point, it triggers an alarm mechanism.
[0060] When the above-mentioned system detects privacy data leakage, it tracks the propagation path of the privacy data in the application through taint analysis technology. After triggering the alarm mechanism, the system can provide a detection report that shows the propagation path of the leaked privacy data.
[0061] After the above system completes the privacy data detection, the system performs a remote attestation process, proving that the current node is a trusted node by interacting with the client multiple times through encrypted messages.
[0062] In addition, in the application of distributed systems, when applications can be deployed on multiple nodes, the hash value of each node will also be verified to ensure data consistency. By verifying the hash value of the node, it ensures that even if the program runs in a malicious environment, the system can still ensure its security and prevent privacy data leakage.
[0063] However, during the aforementioned privacy data detection process, since applications are constantly running, they can dynamically generate data in real time. Therefore, the system needs to perform both static and dynamic data analysis on the applications. Static data analysis requires scanning the entire application codebase, while dynamic data analysis requires real-time monitoring of the flow of private data, resulting in excessive consumption of node computing and memory resources. Furthermore, taint analysis techniques, when processing large datasets, consume significant resources, thus impacting system performance. In summary, large-scale and complex applications can lead to low resource utilization and reduced system performance.
[0064] Furthermore, because the system uses taint analysis to track the propagation paths of private data, dynamic data is generated in real time during application execution. Therefore, in high-concurrency scenarios, the system requires more computing resources to track the propagation paths of private data within static data and within dynamic data, further increasing the computing load.
[0065] In addition, the remote attestation and hash verification steps also bring additional overhead. As the number of nodes increases and the frequency of container deployment increases, the communication load and delay between nodes and clients may further affect system performance.
[0066] The distributed nature of blockchain places higher demands on network resources. Performing privacy data leak detection on on-chain nodes can significantly increase computing and storage loads, leading to increased network bandwidth usage, potentially causing performance degradation and network congestion, and impacting data processing and practical application effectiveness.
[0067] In order to solve the problems in the prior art, the embodiments of the present application provide a method, apparatus, device, storage medium and program product for data detection. The following first describes the method for data detection provided by the embodiments of the present application.
[0068] like Figure 1 As shown, the method is applied to a central node, and the method includes:
[0069] S101. Send the data to be detected to each distributed node, so that the distributed node performs taint analysis on the non-text data in the data to be detected, and obtains a first detection result of the taint data in the non-text data. Use the privacy data detection model to perform semantic analysis on the text data in the data to be detected, and obtain a second detection result output by the privacy data detection model.
[0070] The central node establishes a communication connection with multiple distributed nodes. After obtaining the data to be tested from the target distributed nodes, the central node distributes the data to the multiple distributed nodes for private data detection. The target distributed nodes are nodes where applications are deployed. Using multiple distributed nodes to perform private data detection on the data to be tested improves data processing efficiency and resource utilization.
[0071] The data to be detected can represent the resource usage of the target distributed node and the interaction between the target distributed nodes. In one example, the data to be detected includes the CPU usage, memory usage, and system log information of the target distributed node.
[0072] The process of performing taint analysis on non-text data by the above distributed nodes can be specifically implemented as follows:
[0073] The system identifies private data within non-text data based on pre-defined pollution sources, then marks the private data and determines its propagation path based on its flow within the application. It then determines whether the propagation path reaches a private data detection point. If so, it determines that there is behavioral data leakage within the non-text data.
[0074] The process of the above-mentioned distributed nodes performing privacy data detection on text data is described in detail in subsequent embodiments.
[0075] S102: Receive a first detection result and a second detection result sent by each distributed node.
[0076] Using the above method, after the distributed nodes receive the data to be detected sent by the central node, the distributed nodes can directly perform privacy data leakage detection for the non-text data in the data to be detected, and determine the first detection result of the tainted data in the non-text data. For text data, the privacy data detection model can be used to perform semantic analysis on the text data, so as to determine whether there is tainted data in the text data and whether the tainted data has been leaked. In this way, using the privacy data detection model, it is possible to accurately determine whether there is tainted data in the text data in each computing node and whether the tainted data has been leaked, thereby accurately determining whether there is privacy data leakage in each distributed node, thereby improving the accuracy of the detection results. After receiving the first detection result and the second detection result sent by the distributed node, the central node is convenient for taking corresponding processing measures based on the detection results to ensure the security of the user's privacy data.
[0077] It should be noted that central nodes and distributed nodes can employ data compression algorithms to reduce the amount of data to be transmitted, thereby lowering communication overhead between the central and distributed nodes. Furthermore, the central node can use a remote attestation mechanism based on zero-knowledge proofs to prove to the client that the current distributed node is trustworthy, reducing the computational and communication overhead of remote attestation and program hash verification. During data transmission, incorporating blockchain technology, the central node verifies the integrity of the received data to ensure it has not been tampered with, thereby improving data transmission security.
[0078] In addition, after receiving the detection results fed back by each distributed node, the central node merges the detection results according to the application and generates a complete detection result corresponding to the application. The above complete detection result includes the privacy data propagation path, the data type of the leaked privacy data, and the preset repair suggestions.
[0079] In some embodiments of the present application, the preset container set is allocated to a node with higher security, and corresponding computing resources and storage resources are allocated to the node, which can ensure the stable operation of the node and the security of the user's private data. Figure 2 As shown, the method includes:
[0080] S201: Obtain resource utilization of each distributed node, affinity between a preset pod and each distributed node, and preset trusted hardware supported by each distributed node.
[0081] Among them, resource utilization includes storage resource utilization, computing resource utilization and network resource utilization. The embodiment of the present application does not impose any specific restrictions on resource utilization.
[0082] The degree of affinity is determined based on characteristics such as network topology, latency, and data proximity between the preset pod and distributed nodes.
[0083] Specifically, the central node uses network monitoring tools to collect network latency, bandwidth, and topology information between nodes and calculates the real-time network quality between nodes. The data storage location is located, and the affinity between the preset pod and the distributed nodes is determined based on a preset policy. For example, the preset policy may set corresponding weights for network latency, bandwidth, topology information, real-time network quality, and the distance to the data storage location. The affinity between the preset pod and the distributed nodes is determined by taking a weighted sum of the network latency, bandwidth, topology information, real-time network quality, and the distance to the data storage location.
[0084] S202. Calculate the security score of each distributed node based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node.
[0085] Specifically, the above S202 can be implemented as follows:
[0086] Step 1: For each distributed node, calculate the resource utilization score of the distributed node according to the resource utilization of the computing node.
[0087] It should be noted that the embodiments of the present application do not impose any specific restrictions on the method of calculating resource utilization.
[0088] Step 2: For each distributed node, calculate the affinity score of the distributed node based on the affinity between the preset pod and the distributed node.
[0089] The method for calculating the affinity score is described in the above embodiment regarding the determination of the affinity degree, which will not be repeated here.
[0090] Step 3: For each distributed node, calculate the trusted hardware support score of the distributed node based on whether the distributed node supports the preset trusted hardware.
[0091] Among them, in order to prioritize resource scheduling for nodes that support TEE, the trusted hardware support score is calculated based on whether the distributed node supports the preset trusted hardware.
[0092] In one example, the preset trusted hardware may be a Trusted Platform Module (TPM), Software Guard Extensions (SGX), or hardware isolation. It should be noted that the above-mentioned preset trusted hardware is only an example and is not limited thereto in actual implementation.
[0093] The K8s (Kubernetes container orchestration system) scheduler in the central node calculates the trusted hardware support score by weighting the preset trusted hardware supported by each distributed node according to the weight corresponding to each preset trusted hardware.
[0094] Step 4: For each distributed node, perform a weighted sum of the resource utilization score, affinity score, and trusted hardware support score to obtain the security score of the distributed node.
[0095] In an example, the security score can be calculated using the following formula:
[0096] S=αR+βA+γT
[0097] Where S represents the security score, R represents the resource utilization score of the distributed node, α represents the weight of the resource utilization score, A represents the affinity score between the distributed node and the preset pod, β represents the weight of the affinity score, T represents the trusted hardware support score, and γ represents the weight of the trusted hardware support score.
[0098] S203: Allocate the preset pod to the trusted distributed node with the highest security score.
[0099] It's important to note that by adding a trusted hardware support score to the traditional Kubernetes scheduling algorithm, the security weighting for distributed nodes is increased, improving the security of selected trusted distributed nodes. Furthermore, the weightings of the aforementioned resource utilization score, affinity score, and trusted hardware support score can be flexibly adjusted to prioritize distributed nodes with higher security.
[0100] Furthermore, containers within a pre-configured pod are isolated using a shared kernel, eliminating the need for a separate operating system for each virtual environment and improving system resource utilization. The central node can launch multiple isolated TEE environments on multiple distributed nodes, allowing each distributed node to quickly respond to varying computing needs and flexibly migrate computing tasks between them, thus avoiding the waste of idle resources and improving system resource utilization. Secure container technology provides an additional layer of isolation for containers. Secure key storage and secure boot using the TPM module ensure that authorized code runs in the TEE environment, guaranteeing system security.
[0101] S204: Acquire historical resource data and historical central processing unit (CPU) load data of trusted distributed nodes.
[0102] The historical resource data may include historical computing resource data, historical storage resource data, and historical network resource data.
[0103] Specifically, the central node can use real-time monitoring tools to monitor the resource usage of each distributed node, evaluate the system performance at preset intervals based on the usage of each distributed resource, and adjust resource allocation and container configuration based on the computing tasks and the resource usage of each distributed node.
[0104] In one example, the real-time monitoring tool may be Prometheus.
[0105] S205: Input the historical resource data and historical CPU load data of the trusted distributed node into the load prediction model to obtain a first prediction result.
[0106] The first prediction result includes CPU load data within a preset time period in the future.
[0107] In one example, the load prediction model may be a Long Short-Term Memory (LSTM) model, and the prediction formula is:
[0108]
[0109] in, represents the CPU load data at time t, α and β are model parameters, and y represents the historical CPU load data.
[0110] It should be noted that for low-priority tasks such as background tasks and batch tasks, the central node can adopt a delayed allocation method to prioritize the scheduling of resources to process high-priority tasks. After the high-priority tasks are processed, resources can be called to process low-priority tasks.
[0111] S206. Allocate computing resources and storage resources to the trusted distributed nodes according to the first prediction result.
[0112] Using the method provided in the embodiment of the present application, the central node obtains the resource utilization of each distributed node, presets the pod and each distributed node, and whether each distributed node supports the preset trusted hardware, and calculates the security score of each distributed node. In this way, the security of each distributed node can be accurately judged, thereby determining the trusted distributed node with the highest security score, and deploying the preset pod on the trusted distributed node can improve the security of the container. The central node then uses the historical data of the trusted distributed node to predict the resource requirements of the trusted distributed node in the future, thereby timely scheduling the required resources for the trusted distributed node, thereby improving the system processing efficiency.
[0113] In some embodiments of the present application, since the application continuously generates new data during operation, the central node can only obtain the newly added, modified, or deleted data, thereby avoiding repeated calculations and improving computing efficiency and resource utilization. Specifically, before the above S101, before sending the data to be detected to each distributed node, the method further includes:
[0114] The data trend prediction model is used to predict the indicators to be tested that have data changes within a preset time window. Within the preset time window, data corresponding to the indicators to be tested is obtained from the target distributed node, which is deployed with a preset container set pod.
[0115] The data change trend prediction model is trained using historical data of preset indicators. The data change trend prediction model can be an LSTM model or an ARIMA model.
[0116] In one example, the preset indicator can be resource usage data. For example, using CPU usage, the LSTM model can be trained using historical resource usage data from 15 time steps ago to produce a trained data trend prediction model. When using the data trend prediction model, the historical resource data from the most recent 15 time steps is used as model data, and the data trend prediction model outputs the probability of a change in CPU usage.
[0117] Based on the load changes predicted by the data change trend prediction model, the central node can schedule and allocate resources through Kubernetes. Specifically, the Kubernetes Vertical Pod Autoscaler (VPA) can be used in combination with the predicted load to dynamically adjust the CPU and memory resources of each container in the TEE.
[0118] In this way, the data change trend prediction model can be used to more accurately predict the preset indicators of possible data changes, thereby collecting the data to be tested from the target distributed nodes within the preset time window, reducing unnecessary calculations.
[0119] It should be noted that after the central node collects the data to be detected, if the data to be detected is data corresponding to multiple preset indicators, it adopts a parallel processing method to distribute the data corresponding to each preset indicator to different distributed nodes, that is, each distributed node processes the data of one preset indicator separately. Specifically, the embodiment of the present application also provides a data detection method, which is applied to distributed nodes, such as Figure 3 As shown, the specific methods include:
[0120] S301: Receive data to be detected sent by a central node in real time.
[0121] The data to be detected includes non-text data and text data.
[0122] S302: Perform stain analysis on the non-text data to obtain a first detection result of stain data in the non-text data.
[0123] The first detection result indicates whether tainted data in the non-text data has been leaked.
[0124] Specifically, the process of taint analysis on non-text data is as follows:
[0125] First, distributed nodes identify and mark private data within non-text data, such as user passwords, as tainted data sources. Then, based on the propagation of the tainted data within the program, they analyze the flow of the tainted data between variables, function calls, and return values, ensuring that all affected variables and objects are marked as tainted data.
[0126] Then, the system monitors sensitive operations in the program, which can be data input and output operations such as network transmission or file writing, and checks whether the sensitive operations involve tainted data and tainted sources.
[0127] When tainted data or tainted sources reach the sensitive operation detection point, the distributed nodes can determine that private data leakage has occurred.
[0128] S303: Input the text data into a privacy data detection model, perform semantic analysis on the text data through the privacy data detection model, and obtain a second detection result output by the privacy data detection model.
[0129] The second detection result indicates whether the private data in the text data is leaked.
[0130] As you can understand, the text data to be detected may contain private data. For example, "My bank card password is xxxxxx," where xxxxxx is the private data contained in the text data. The private data detection model can perform semantic analysis on the context of the text data to identify the private data contained in the text data, thereby improving the accuracy of text data detection.
[0131] S304: Feedback the first detection result and the second detection result to the central node.
[0132] By adopting the method provided in the embodiment of the present application, after the distributed nodes receive the data to be detected sent by the central node, the distributed nodes can directly perform privacy data leakage detection for the non-text data in the data to be detected, and determine the first detection result of the tainted data in the non-text data. For text data, a privacy data detection model can be used to perform semantic analysis on the text data to determine whether there is tainted data in the text data and whether the tainted data has been leaked. In this way, using the privacy data detection model, it is possible to accurately determine whether there is tainted data in the text data in each computing node and whether the tainted data has been leaked, thereby accurately determining whether there is privacy data leakage in each distributed node, thereby improving the accuracy of the detection results and ensuring the security of the user's privacy data.
[0133] The following describes S303, which involves inputting text data into a privacy-preserving data detection model, performing semantic analysis on the text data using the privacy-preserving data detection model, and obtaining a second detection result output by the privacy-preserving data detection model. The privacy-preserving data detection model includes a feature extraction network and a context analysis network. Specifically, S303 can be implemented as follows:
[0134] Step A: Input text data into the privacy data detection model, and extract features from the text data through the feature extraction network in the privacy data detection model to obtain a text feature vector.
[0135] Specifically, after the text data is input into the privacy data detection model, the privacy data detection model performs word segmentation processing on the text data, performs feature extraction on the text data obtained after the word segmentation processing, and obtains a text feature vector corresponding to each text.
[0136] In one example, the feature extraction network can be a convolutional neural network (CNN) and a recurrent neural network (RNN).
[0137] Convolutional neural networks and recurrent neural networks can identify complex patterns in the data to be detected, which are defined as:
[0138] h t =f(W h h t-1 +W x x t +b)
[0139] Among them, h t is the current state, W h and W xare the weights of the hidden layer and the input layer, respectively, and f is the activation function. This model can extract important features from the data to be detected and improve the sensitivity of privacy leakage detection.
[0140] Step B: Perform semantic analysis on the text feature vector through the context analysis network in the privacy data detection model to determine whether the privacy data in the text data has been leaked, and output a second detection result.
[0141] Among them, the context analysis network can be a pre-trained language model (Bert model), which can perform semantic embedding on the above-mentioned text feature vector, so that the privacy data detection model can understand the contextual semantics in the text data, thereby enabling the privacy data detection model to more accurately detect the privacy data in the text data and improve the accuracy of privacy leakage detection.
[0142] In one example, suppose the user inputs a parameter P and the application returns a value R. These input and return values are fed into a privacy-preserving data detection model, which determines their sensitivity and outputs a probability score for the input and return values leading to privacy data leakage. Specifically, S(P, R) = NLP(P, R), where S(P, R) represents the detection result output by the privacy-preserving data detection model, and NLP(P, R) represents the detection process of the privacy-preserving data detection model.
[0143] In addition, to ensure the accuracy of the model detection results, distributed nodes can dynamically adjust the privacy data detection model by learning from historical data. Specifically, the model dynamic adjustment formula is: Among them, M t is the current model, η is the learning rate, L is the loss function, and D is the historical dataset.
[0144] By adopting the method provided in the embodiments of the present application, the privacy data detection model can be used to identify important features in text data, thereby identifying the contextual semantics of the important features, making it easier for the privacy data detection model to detect privacy data leakage based on the contextual semantics, thereby improving the accuracy of the detection results and expanding the recognition range. Distributed nodes can effectively identify text data.
[0145] In some embodiments of the present application, in order to further improve the accuracy of data detection, the distributed node can verify the second detection result. Based on this, before the above S304, feeding back the first detection result and the second detection result to the central node, the method further includes:
[0146] If the likelihood score is greater than a first preset threshold, preset indicator feature data is extracted from the text data. A support vector machine is used to calculate the distance between the preset indicator feature data and the decision boundary of the normal behavior model to obtain an anomaly score. If the anomaly score is greater than a second preset threshold, a privacy data leak is determined in the text data.
[0147] The normal behavior model is trained using pre-set indicator feature data that does not indicate privacy data leakage. In one example, the pre-set indicator feature data may include packet size, transmission frequency, and timestamp.
[0148] It should be noted that the decision boundary of the normal behavior model is a hyperplane, which is used to distinguish different types of data. If the anomaly score of the text data is less than or equal to the second preset threshold, it indicates that the text data is normal behavior data, that is, no privacy data leakage has occurred. If the anomaly score of the text data is greater than the second preset threshold, it indicates that the text data is not normal behavior data, that is, privacy data leakage has occurred.
[0149] In one example, the formula for calculating the anomaly score by a support vector machine (SVM) is:
[0150] S a =f(X)
[0151] Among them, S a represents the anomaly score, and X represents the preset indicator feature.
[0152] By using the method provided in the embodiment of the present application, when the possibility score output by the privacy data detection model is greater than the first preset threshold, in order to improve the accuracy of the detection results, an anomaly detection algorithm can be combined to obtain preset indicator feature data, calculate the distance between the preset indicator feature data and the kernel decision boundary, and obtain the anomaly score. Based on the anomaly score, it can be further judged whether there is privacy data leakage in the text data, thereby improving the accuracy of the detection results.
[0153] Based on the detection results generated by the above-mentioned distributed nodes, the distributed nodes maintain a data dependency graph. Based on the data dependency graph, the data dependency graph contains the calling relationship between the data. Based on the calling relationship and the above-mentioned determination of the leaked privacy data, the propagation path of the leaked privacy data can be determined, so that the distributed nodes add the propagation path to the detection results and update the above-mentioned data dependency graph according to the detection results.
[0154] The following combination Figure 4 The following describes the process of the data detection method provided in the embodiment of the present application. Figure 4As shown in the figure, the method is divided into seven stages, namely, system initialization and environment construction stage, data collection and preprocessing stage, taint analysis and data dependency analysis stage, privacy data leakage detection stage, data transmission and communication optimization stage, detection result processing and feedback stage, and system maintenance and upgrade stage.
[0155] Among them, the system initialization and environment construction phase includes TEE configuration, lightweight virtualization deployment, dynamic resource allocation and multi-level security protection.
[0156] For the above-mentioned TEE configuration, users can configure and start TEE for each distributed node on the central control platform of the central node.
[0157] For the above-mentioned lightweight virtualization deployment, the central node calculates the node security score based on the resource utilization score, affinity score and trusted hardware support score, and deploys the preset pod to the trusted distributed node with the highest security score.
[0158] Regarding the above-mentioned dynamic resource allocation, please refer to the above-mentioned description of the central node allocating computing resources and storage resources to the trusted distributed nodes, which will not be repeated here.
[0159] To address these multi-layered security challenges, gVisor and Kata Containers can be combined to provide an additional layer of isolation for containers, particularly suitable for processing sensitive data. Furthermore, TPM provides secure key storage and secure boot to ensure that only authorized code can run in the TEE.
[0160] The data collection and preprocessing stage includes data collection, data cleaning and preprocessing, and data localization.
[0161] For the aforementioned data collection, the central node can collect data from distributed nodes and external data sources to ensure data integrity and consistency. To ensure data consistency, a middleware service can be set up in the external data source and distributed nodes. This middleware service can be used to perform real-time format conversion on the data in the distributed nodes and external data sources, ensuring that the data received by the central node is in a consistent format.
[0162] For the above-mentioned data cleaning and preprocessing, the central node can perform data cleaning, normalization and preprocessing on the received data to remove noise and abnormal data.
[0163] For the above-mentioned data localization, users define the data structure in each distributed node through the central control platform of the central node to ensure that the data types and data formats processed by all nodes are consistent, thereby ensuring data compatibility and efficient serialization between different languages.
[0164] The taint analysis and data dependency analysis phase includes taint analysis, incremental analysis, data dependency analysis, and machine learning optimization.
[0165] Regarding the above-mentioned taint analysis, reference is made to the related description of the distributed node performing taint analysis on non-text data in the above-mentioned embodiment, which will not be repeated here.
[0166] For the above incremental analysis, the central node uses a data change trend prediction model to predict the indicators to be detected that have data changes within a preset time window. Then the central node collects data corresponding to the indicators to be detected according to the indicators to be detected and sends it to the distributed nodes, so that the distributed nodes can perform incremental analysis on the data that is newly added, modified or deleted within the preset time window, thereby improving computing efficiency.
[0167] For this data dependency analysis, distributed nodes determine data dependencies based on the call relationships between data, thereby constructing a data dependency graph. This data dependency graph is used to identify the propagation paths of tainted data within the application. This facilitates subsequent feedback from distributed nodes to the central node regarding the propagation paths of tainted data in the detection results, accurately identifying the paths of private data leakage.
[0168] For the above-mentioned machine learning optimization, the central node can use the historical data within the most recent preset number of time steps to correct the data change trend prediction model, and continuously improve the accuracy of the model prediction results.
[0169] The privacy data leakage detection stage includes building deep learning and natural language processing (NLP) models as well as semantic analysis and anomaly detection.
[0170] For the above-mentioned construction of deep learning and NLP models, the deep learning and NLP models are trained using a training data set, wherein the training data set includes privacy data leakage behavior data and normal behavior data.
[0171] For the aforementioned semantic analysis and anomaly detection, distributed nodes input the data to be tested into the privacy data detection model. The model then analyzes the contextual semantics of the data and, based on the contextual information, determines whether the text within the data has been compromised. If the model determines a privacy breach has occurred, it verifies the model's results using a support vector machine (SVM) combined with an anomaly detection algorithm, ensuring accuracy.
[0172] The data transmission and communication optimization stage includes communication protocol optimization, data compression and blockchain verification.
[0173] For the above communication protocol optimization, users can pre-set an efficient communication protocol to reduce network delay. It should be noted that the embodiment of the present application does not impose specific restrictions on the communication protocol, and users can set it according to the actual application scenario.
[0174] Regarding the above data compression, before the central node and the distributed nodes transmit data, they can compress the data to be transmitted using a preset data compression algorithm to reduce the amount of data transmitted, improve transmission efficiency and reduce communication overhead.
[0175] Regarding the above-mentioned blockchain verification, during the data transmission process between the central node and the distributed nodes, the central node and the distributed nodes use blockchain technology to verify the integrity of the received data to ensure that the data has not been tampered with, thereby improving the security of data transmission.
[0176] The test result processing and feedback stage includes storing test results, generating test reports, and optimizing test strategies.
[0177] Regarding the above-mentioned stored detection results, the distributed nodes may feed back the above-mentioned first detection result and the second detection result to the central node, and the central node further summarizes the detection results fed back by each distributed node and stores them locally.
[0178] Regarding the aforementioned detection report generation, the central node aggregates the detection results fed back by each distributed node and generates a detection report. The detection report includes the path of the private data leakage, the type of private data leaked, the severity of the incident, and corresponding remediation suggestions. The incident severity and remediation suggestions are pre-set based on the data type and leakage path. After the central node obtains the leakage path and data type, it searches for the corresponding remediation suggestions and incident severity to help users promptly address the application and ensure the security of their private data.
[0179] In response to this optimized detection strategy, based on the aforementioned incremental analysis technology, the central node can determine the data added or modified within a preset time window. It then calculates the hash value for this data and updates the corresponding Merkle tree. The Merkle tree verifies the integrity of the new or modified data, improving computational efficiency and avoiding duplicate calculations, thereby reducing the computational burden.
[0180] The system maintenance and upgrade phase includes security updates, performance monitoring, and system optimization.
[0181] In response to the above security updates, the central node and distributed nodes can perform security updates regularly, scan the system for vulnerabilities, and repair the scanned vulnerabilities.
[0182] For the above performance monitoring, the central node can monitor the efficiency of distributed nodes in privacy data detection and the performance indicators of the nodes in real time, and automatically identify and handle potential performance bottlenecks.
[0183] In one example, running independent applications in multiple environments using lightweight virtualization technology can lead to resource contention, resulting in insufficient resources on each distributed node, which in turn affects the processing performance of the distributed nodes. Therefore, by monitoring the resources of each distributed node in real time, the central node can schedule the resources and processes of each distributed node. After a distributed node completes a high-priority task, it promptly reclaims and reallocates resources, ensuring the efficient operation of each distributed node.
[0184] For the above system optimization, the central node can use the collected data to be tested to dynamically update the data change trend prediction model, and the distributed nodes can also use the collected data to be tested to dynamically update the privacy data detection model, thereby ensuring the accuracy of the detection results.
[0185] Based on the same concept, the embodiment of the present application provides a data detection device, which is applied to distributed nodes, such as Figure 5 As shown, the device includes:
[0186] A receiving module 501 is configured to receive data to be detected sent by a central node in real time, wherein the data to be detected includes non-text data and text data;
[0187] An analysis module 502 is configured to perform a taint analysis on the non-text data to obtain a first detection result of tainted data in the non-text data, wherein the first detection result indicates whether the tainted data in the non-text data has been leaked.
[0188] a detection module 503 configured to input the text data into a privacy data detection model, perform semantic analysis on the text data using the privacy data detection model, and obtain a second detection result output by the privacy data detection model, wherein the second detection result indicates whether private data in the text data has been leaked;
[0189] The feedback module 504 is configured to feed back the first detection result and the second detection result to the central node.
[0190] In a possible implementation, the detection module 503 is specifically configured to:
[0191] Inputting the text data into the privacy data detection model, performing feature extraction on the text data through a feature extraction network in the privacy data detection model to obtain a text feature vector;
[0192] The text feature vector is semantically analyzed by the context analysis network in the privacy data detection model to determine whether the privacy data in the text data has been leaked, and the second detection result is output.
[0193] In a possible implementation, the second detection result includes a probability score of leakage of the private data in the text data; and the apparatus further includes:
[0194] An extraction module, configured to extract preset indicator feature data from the text data when the likelihood score is greater than a first preset threshold;
[0195] a calculation module, configured to calculate, using a support vector machine, the distance between the preset indicator feature data and a decision boundary of a normal behavior model to obtain an anomaly score, wherein the normal behavior model is trained using the preset indicator feature data in which no privacy data leakage occurs;
[0196] The determination module is configured to determine, when the anomaly score is greater than a second preset threshold, whether privacy data leakage exists in the text data.
[0197] It should be noted that the data detection device is a device corresponding to the above-mentioned method for data detection applied to distributed nodes. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of the device and can achieve the same technical effects.
[0198] Based on the same concept, the embodiment of the present application provides a data detection device, which is applied to a central node, such as Figure 6 As shown, the device includes:
[0199] A sending module 601 is configured to send the data to be detected to each distributed node, so that the distributed node performs a taint analysis on non-text data in the data to be detected to obtain a first detection result of the taint data in the non-text data, and performs a semantic analysis on the text data in the data to be detected using a privacy data detection model to obtain a second detection result output by the privacy data detection model;
[0200] The receiving module 602 is configured to receive the first detection result and the second detection result sent by each distributed node.
[0201] In a possible implementation, the device further includes:
[0202] A prediction module is used to predict the indicators to be detected that have data changes within a preset time window using a data change trend prediction model, wherein the data change trend prediction model is trained using historical data of the preset indicators;
[0203] The acquisition module is used to obtain data corresponding to the indicator to be detected from the target distributed node within the preset time window to obtain the data to be detected, and the target distributed node is deployed with a preset container set pod.
[0204] In a possible implementation, the device further includes:
[0205] The acquisition module is further configured to acquire resource utilization of each distributed node, affinity between the preset pod and each distributed node, and preset trusted hardware supported by each distributed node;
[0206] A calculation module, configured to calculate a security score of each distributed node based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node;
[0207] An allocation module, configured to allocate the preset pod to a trusted distributed node with the highest security score;
[0208] The acquisition module is further used to acquire historical resource data and historical central processing unit CPU load data of the trusted distributed node;
[0209] The prediction module is configured to input the historical resource data of the trusted distributed node and the historical CPU load data into a load prediction model to obtain a first prediction result, wherein the second prediction result includes CPU load data within a preset time period in the future;
[0210] The allocation module is further configured to allocate computing resources and storage resources to the trusted distributed nodes according to the first prediction result.
[0211] In a possible implementation, the computing module is specifically configured to:
[0212] For each distributed node, calculating a resource utilization score of the distributed node according to the resource utilization of the distributed node;
[0213] For each distributed node, calculate the affinity score of the distributed node based on the affinity between the preset pod and the distributed node;
[0214] For each distributed node, calculating a trusted hardware support score of the computing node based on whether the computing node supports preset trusted hardware;
[0215] For each distributed node, a weighted sum is performed on the resource utilization score, the affinity score, and the trusted hardware support score to obtain a security score of the distributed node.
[0216] It should be noted that the data detection device is a device corresponding to the above-mentioned method for data detection applied to the central node. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of the device and can achieve the same technical effects.
[0217] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0218] The electronic device may include a processor 701 and a memory 702 storing computer program instructions.
[0219] Specifically, the processor 701 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0220] The memory 702 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 702 may include removable or non-removable (or fixed) media. Where appropriate, the memory 702 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 702 is a non-volatile solid-state memory.
[0221] In certain embodiments, the memory 702 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0222] The processor 701 implements any one of the data detection methods in the above embodiments by reading and executing computer program instructions stored in the memory 702 .
[0223] In one example, the electronic device may further include a communication interface 703 and a bus 710. Figure 7 As shown, the processor 701 , the memory 702 , and the communication interface 703 are connected via a bus 704 and communicate with each other.
[0224] The communication interface 703 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0225] The bus 704 includes hardware, software, or both that couples components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Super Transmission (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 704 may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.
[0226] In addition, in conjunction with the data detection method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the data detection methods in the above embodiments is implemented.
[0227] An embodiment of the present application also provides a computer program product, including a computer program, which implements any one of the data detection methods in the above embodiments when the computer program is processed and executed.
[0228] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0229] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (erasable read-only memory, EROM), floppy disks, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical discs, hard disks, optical fiber media, radio frequency (Radio Frequency, RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0230] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0231] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0232] The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application.
Claims
1. A data detection method, characterized in that: Applied to distributed nodes, the method includes: Receiving data to be detected from a central node in real time, wherein the data to be detected includes non-text data and text data; Performing a taint analysis on the non-text data to obtain a first detection result, where the first detection result indicates whether tainted data in the non-text data has been leaked; Inputting the text data into a privacy data detection model, performing semantic analysis on the text data using the privacy data detection model, and obtaining a second detection result output by the privacy data detection model, wherein the second detection result indicates whether private data in the text data has been leaked; Feedback the first detection result and the second detection result to the central node.
2. The method according to claim 1, characterized in that Inputting the text data into a privacy data detection model, performing semantic analysis on the text data using the privacy data detection model, and obtaining a second detection result output by the privacy data detection model includes: Inputting the text data into the privacy data detection model, performing feature extraction on the text data through a feature extraction network in the privacy data detection model to obtain a text feature vector; The text feature vector is semantically analyzed by the context analysis network in the privacy data detection model to determine whether the privacy data in the text data has been leaked, and the second detection result is output.
3. The method according to claim 1 or 2, characterized in that The second detection result includes a probability score of leakage of the private data in the text data; before feeding back the first detection result and the second detection result to the central node, the method further includes: When the likelihood score is greater than a first preset threshold, extracting preset indicator feature data from the text data; Using a support vector machine to calculate the distance between the preset indicator feature data and a decision boundary of a normal behavior model to obtain an anomaly score, wherein the normal behavior model is trained using the preset indicator feature data without privacy data leakage behavior; When the abnormality score is greater than a second preset threshold, it is determined that privacy data leakage exists in the text data.
4. A data detection method, characterized in that: Applied to a central node, the method includes: Sending the data to be detected to each distributed node, so that the distributed node performs a taint analysis on non-text data in the data to be detected to obtain a first detection result, and performing a semantic analysis on the text data in the data to be detected using a privacy data detection model to obtain a second detection result output by the privacy data detection model; Receive the first detection result and the second detection result sent by each distributed node.
5. The method according to claim 4, characterized in that Before sending the data to be detected to each distributed node, the method further includes: Using a data change trend prediction model to predict the indicators to be detected that have data changes within a preset time window, the data change trend prediction model is trained using historical data of the preset indicators; Within the preset time window, data corresponding to the indicator to be detected is acquired from a target distributed node to obtain the data to be detected, where a preset container set pod is deployed on the target distributed node.
6. The method according to claim 5, characterized in that Before using the data change trend prediction model to predict the indicator to be detected that has undergone data changes within a preset time window, the method further includes: Obtaining resource utilization of each distributed node, affinity between the preset pod and each distributed node, and preset trusted hardware supported by each distributed node; Calculate a security score for each distributed node based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node; Allocate the preset pod to the trusted distributed node with the highest security score; Obtain historical resource data and historical central processing unit (CPU) load data of the trusted distributed node; Inputting the historical resource data of the trusted distributed node and the historical CPU load data into a load prediction model to obtain a first prediction result, wherein the first prediction result includes CPU load data within a preset time period in the future; Allocate computing resources and storage resources to the trusted distributed node based on the first prediction result.
7. The method according to claim 6, characterized in that The security score of each distributed node is calculated based on the resource utilization of each distributed node, the affinity between the preset pod and each distributed node, and the preset trusted hardware supported by each distributed node, including: For each distributed node, calculating a resource utilization score of the distributed node according to the resource utilization of the distributed node; For each distributed node, calculate the affinity score of the distributed node based on the affinity between the preset pod and the distributed node; For each distributed node, calculating a trusted hardware support score of the distributed node according to whether the distributed node supports preset trusted hardware; For each distributed node, a weighted sum is performed on the resource utilization score, the affinity score, and the trusted hardware support score to obtain a security score of the distributed node.
8. A data detection device, characterized in that: Applied to a distributed node, the device includes: A receiving module, configured to receive data to be detected sent by a central node in real time, wherein the data to be detected includes non-text data and text data; an analysis module, configured to perform a taint analysis on the non-text data to obtain a first detection result, wherein the first detection result indicates whether taint data in the non-text data has been leaked; a detection module, configured to input the text data into a privacy data detection model, perform semantic analysis on the text data using the privacy data detection model, and obtain a second detection result output by the privacy data detection model, wherein the second detection result indicates whether private data in the text data has been leaked; A feedback module is used to feed back the first detection result and the second detection result to the central node.
9. A data detection device, characterized in that: Applied to a central node, the device includes: a sending module, configured to send the data to be detected to each distributed node, so that the distributed node performs a taint analysis on the non-text data in the data to be detected to obtain a first detection result, and performs a semantic analysis on the text data in the data to be detected using a privacy data detection model to obtain a second detection result output by the privacy data detection model; A receiving module is used to receive the first detection result and the second detection result sent by each distributed node.
10. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the data detection method according to any one of claims 1 to 3 or claims 4 to 7 is implemented.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the data detection method according to any one of claims 1 to 3 or claims 4 to 7.
12. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the data detection method according to any one of claims 1 to 3 or claims 4 to 7.