Data security protection method and related equipment thereof
By analyzing and assessing data risks from multiple dimensions, and combining regular expression rules and dynamic feature models, this approach addresses the inadequacy of existing data security protection methods, achieving more efficient data identification and security protection.
Patent Information
- Application Number
- CN202511361942.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-12-12
AI Technical Summary
Existing data security protection methods are insecure and cannot cope with data obfuscation, encrypted transmission, or new leakage methods, resulting in a high risk of data leakage.
By collecting traffic data and related information from the target data link, we use regular expression rules and dynamic feature models to judge data risks, combine multi-dimensional analysis to determine whether the data is risky, and implement security protection when risks are identified.
It improves the accuracy and security of data identification, reduces the false negative rate, enhances the ability to identify sensitive data, and effectively addresses complex leakage scenarios such as data distortion and tampering.
Smart Images

Figure CN121125268A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of information security, and particularly relates to a data security protection method and a related device thereof. BACKGROUND
[0002] Network data security protection refers to a series of security protection measures and management strategies taken to protect data transmitted, processed and stored in a network from being leaked, tampered with, lost or illegally used.
[0003] In the prior art, a traditional data loss prevention (DLP) system relies on static rules or a single data feature (such as a mobile phone number format) to determine whether data is safe, which is difficult to deal with data obfuscation, encrypted transmission or new leakage methods (such as document similarity leakage), resulting in poor security of data security protection. SUMMARY
[0004] The purpose of the embodiments of the application is to provide a data security protection method and a related device thereof, which can solve the problem of poor security of the existing data security protection method.
[0005] In a first aspect, the embodiments of the application provide a data security protection method, which comprises: collecting traffic data of a target data link and first information and second information corresponding to the traffic data, the first information being used to indicate a feature of a device associated with the traffic data, and the second information being information associated with an action of accessing the traffic data; determining third information of the traffic data based on the traffic data and a preset regular expression rule, the third information being used to indicate whether the traffic data is sensitive data; determining a risk score corresponding to the traffic data according to the second information; judging whether the traffic data is risk data based on at least one of the first information, the third information and the risk score; performing security protection on the traffic data in a case where the traffic data is risk data.
[0006] Optionally, after collecting the traffic data of the target data link, the method further comprises: obtaining comparison data, the traffic data being data collected at a first node, the comparison data being data collected at a second node and being of the same type as the traffic data, and the second node being a previous node of the first node in a data transmission direction of the target data link; judging whether the traffic data is risk data based on a similarity between the comparison data and the traffic data.
[0007] Optionally, the determining whether the traffic data is risk data based on the similarity between the comparison data and the traffic data comprises: calculating a first fingerprint string corresponding to the traffic data and a second fingerprint string corresponding to the comparison data based on a document fingerprint algorithm; in a case where the similarity between the first fingerprint string and the second fingerprint string is greater than a preset similarity threshold, determining that the traffic data is risk data.
[0008] Optionally, the determining whether the traffic data is risk data based on at least one of the first information, the third information or the risk score comprises at least one of the following: in a case where the third information indicates that the traffic data is sensitive data, determining that the traffic data is risk data; in a case where the risk score is greater than a preset risk score threshold, determining that the traffic data is risk data; in a case where the first information indicates that an interface, an application, an Internet Protocol (IP) address or an account corresponding to the traffic data has a risk behavior, determining that the traffic data is risk data.
[0009] Optionally, the collecting traffic data of a target data link comprises: collecting, by a network probe arranged at a target node of the target data link, the traffic data of the target node, the target node comprising an application programming interface gateway, an inter-microservice communication interface, a database access agent, a cloud storage access point, a cross-cloud channel, a message middleware and a terminal device software development kit (SDK).
[0010] Optionally, after the security protection on the traffic data, the method further comprises: storing a security log, the security log comprising the traffic data and information indicating the security protection; determining an attack chain path of the traffic data based on the security log and node information of the traffic data; generating a control policy of a node corresponding to the traffic data based on the attack chain path and the node information, the control policy being used for data access on the node corresponding to the traffic data.
[0011] In a second aspect, an embodiment of the present application provides a data security protection device, the device comprising: The acquisition module is used to acquire traffic data of the target data link and first information and second information corresponding to the traffic data. The first information is used to indicate the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data. The first determining module is used to determine third information of the traffic data based on the traffic data and preset regular expression rules, wherein the third information is used to indicate whether the traffic data is sensitive data; The second determining module is used to determine the risk score corresponding to the traffic data based on the second information; The first judgment module is used to determine whether the traffic data is risky data based on at least one of the first information, the third information, and the risk score; A security module is used to provide security protection for the traffic data when the traffic data is risky data.
[0012] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, they implement the steps of the data security protection method as described in the first aspect.
[0013] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the data security protection method as described in the first aspect are implemented.
[0014] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the data security protection method as described in the first aspect.
[0015] In this embodiment, traffic data of the target data link and first and second information corresponding to the traffic data are collected. The first information indicates the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data. Based on the traffic data and preset regular expression rules, third information of the traffic data is determined, which indicates whether the traffic data is sensitive data. A risk score corresponding to the traffic data is determined according to the second information. Based on at least one of the first information, the third information, and the risk score, it is determined whether the traffic data is risky data. If the traffic data is risky data, security protection is provided for the traffic data. This method analyzes the traffic data from multiple dimensions, including the first information, the third information, and the risk score, to determine whether the traffic data is risky. Compared with existing technologies, this method has more multi-dimensional detection dimensions and a stronger ability to identify sensitive data content, thus reducing the false negative rate for identifying risky data and improving security. Attached Figure Description
[0016] Figure 1 A flowchart illustrating the data security protection method provided in this application embodiment; Figure 2 This is a schematic diagram of the data security protection device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0019] The data security protection method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0020] like Figure 1 As shown in the embodiments of this application, the data security protection method includes the following steps: Step 101: Collect traffic data of the target data link and first information and second information corresponding to the traffic data. The first information is used to indicate the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data.
[0021] The first information may include interface profiles, application profiles, Internet Protocol (IP) profiles, and account profiles. Interface profiles may include: Queries Per Second (QPS), response code distribution, and frequency of sensitive fields. Application profiles may include: data lineage graphs (indicating the generation, flow, processing, storage, and consumption paths of data throughout the system), and data dependencies between microservices (in a microservice architecture, applications are broken down into many independent but collaborative modules, and data is often called, transferred, shared, or transformed between microservices) (clearly indicating which microservices provide or depend on which data, who gives data to whom, and who needs whose data). IP profiles may include: the geographical location of the IP address and the risk level of historical attack records (obtained through a threat intelligence database of associated data). Account profiles may include: account operation baselines, abnormal logins, and privilege escalation.
[0022] The second data may include the data flow direction, access frequency, and device correlation of traffic data.
[0023] Step 102: Based on the traffic data and the preset regular expression rules, determine the third information of the traffic data, which is used to indicate whether the traffic data is sensitive data.
[0024] In this step, for sensitive data with specific formats, such as mobile phone numbers, SIM card numbers, and email addresses, which are private data for users, the presence of sensitive data can be determined by checking whether the content of the data traffic conforms to the corresponding regular expression rules. For example, if the data traffic successfully matches a field that satisfies the regular expression rule for a mobile phone number, it means that the user's mobile phone number can be obtained through the data traffic, which indicates that the data traffic poses a privacy leakage risk, and the third information indicates that the data traffic is sensitive data.
[0025] Step 103: Determine the risk score corresponding to the traffic data based on the second information.
[0026] In this step, the second information can be input into a pre-trained dynamic feature model. The dynamic feature model is used to construct a dynamic risk scoring matrix for data flow direction, access frequency, and device correlation. Understandably, the higher the output score of the corresponding dimensions such as data flow direction, access frequency, and device correlation in the dynamic risk scoring matrix, the greater the probability that the traffic data is risky.
[0027] The training process for a dynamic feature model can be as follows: Historical secondary information of historical traffic data and manually labeled dynamic risk scoring matrix are collected. The historical secondary information is used as the input of the model, and the manually labeled dynamic risk scoring matrix is used as the label to supervise the training of the model. After training, the dynamic feature model is obtained.
[0028] Step 104: Based on at least one of the first information, the third information, and the risk score, determine whether the traffic data is risky data.
[0029] In this step, traffic data is analyzed from multiple dimensions, including primary information, secondary information, and risk scoring, to determine whether it is risky data. If the traffic data is risky, it indicates a significant risk of privacy breach. Furthermore, the identified risky data can be displayed in a table format, allowing users to more intuitively view the terminal, IP address, application, and sensitive tags associated with the risk response, thus improving the user experience.
[0030] Step 105: If the traffic data is risky, perform security protection on the traffic data.
[0031] When traffic data is identified as risky, security measures should be implemented to reduce the risk of privacy breaches. These security measures can include a five-level dynamic blocking strategy, comprising: 1) Traffic limiting: Implement bandwidth restrictions on suspicious connections; 2) Connection Reset: Forcefully terminates the current TCP session; 3) Account Lockout: Prohibit further operations on the account in question; 4) Service Circuit Breaker: Suspend API calls to the relevant microservices; 5) Network isolation: Blacklist the attack source IP and isolate it in a sandbox environment.
[0032] Security protection operations can also be implemented through a multi-channel alarm platform, including: 1) Notify security management personnel via SMS / email 2) Automatically generate maintenance work orders and push them to management systems such as Jira. 3) Integrates with Security Information and Event Management (SIEM) systems to support aggregated analysis of alarm events.
[0033] In the method of this application embodiment, traffic data is analyzed by integrating multiple dimensions such as first information, third information, and risk score to determine whether the traffic data is risky data. Compared with the prior art, the detection dimensions are more multi-dimensional and the identification capability of sensitive data content is stronger. Therefore, the false negative rate of identifying risky data can be reduced, thereby improving security.
[0034] Optionally, after collecting traffic data from the target data link, the method further includes: Obtain comparison data, wherein the traffic data is data collected at the first node, and the comparison data is data of the same type as the traffic data collected at the second node, wherein the second node is the node preceding the first node in the data transmission direction of the target data link; Based on the similarity between the comparison data and the traffic data, it is determined whether the traffic data is risky data.
[0035] In this embodiment, the traffic data is compared with the same type as the traffic data in the previous node (second node). By judging the similarity between the traffic data and the comparison data, it can be determined whether the traffic data has been tampered with between the previous node and the current node (first node). Judging whether the traffic data is risky data from the dimension of whether the data has been tampered with can further reduce the false negative rate of risky data identification.
[0036] Optionally, determining whether the traffic data is risky data based on the similarity between the comparison data and the traffic data includes: Based on the document fingerprinting algorithm, calculate the first fingerprint string corresponding to the traffic data and the second fingerprint string corresponding to the comparison data; If the similarity between the first fingerprint string and the second fingerprint string is greater than a preset similarity threshold, the traffic data is determined to be risky data.
[0037] In this embodiment, the document fingerprint algorithm can be a Locality-Sensitive Hash (LSH) algorithm for text or set features. This algorithm can generate a fixed-length binary fingerprint for an object (such as a piece of text or a dataset) through feature processing and hash mapping. For similar objects, the smaller the Hamming distance between the generated fingerprints, the better. Using the document fingerprint algorithm to determine the similarity between traffic data and comparison data can improve the accuracy of the similarity results.
[0038] Optionally, determining whether the traffic data is risky data based on at least one of the first information, the third information, or the risk score includes at least one of the following: If the third piece of information indicates that the traffic data is sensitive data, then the traffic data is determined to be risky data. If the risk score is greater than a preset risk score threshold, the traffic data is determined to be risky data. If the first information indicates that the interface, application, Internet Protocol IP address, or account corresponding to the traffic data has risky behavior, the traffic data is determined to be risky data.
[0039] In this embodiment, traffic data can be judged as risky data based on the first information, the third information, or the risk score. This multi-dimensional judgment method can reduce the false negative rate of risky data compared to the single-dimensional judgment method of the prior art.
[0040] Optionally, the traffic data of the target data link being collected includes: Traffic data of the target node is collected by a network probe set on the target node of the target data link. The target node includes an application programming interface (API) gateway, a microservice communication interface, a database access proxy, a cloud storage access point, a cross-cloud channel, a message middleware, and a terminal device software development kit (SDK).
[0041] In this embodiment, a lightweight network probe cluster is deployed at seven key nodes in the data link (API gateway, microservice inter-communication interface, database access proxy, cloud storage access point, cross-cloud channel, message middleware, and terminal device SDK) to achieve the following functions: Traffic mirroring: Capture full network traffic through port mirroring or bypass listening technology.
[0042] Metadata extraction: Records contextual information such as source / destination IP, protocol type, and timestamp of data packets.
[0043] Preprocessing filtering: Filter non-sensitive business traffic based on a whitelist mechanism to reduce the backend analysis load.
[0044] By deploying a full-domain data link probe matrix, we can avoid the situation where risk data is missed due to incomplete monitoring coverage of the data link, and further reduce the false negative rate.
[0045] Optionally, after providing security protection for the traffic data, the method further includes: Store a security log, the security log including the traffic data and information for indicating the security protection; Based on the node information of the security logs and the traffic data, the attack chain path of the traffic data is determined; Based on the attack chain path and the node information, a control policy for the node corresponding to the traffic data is generated. The control policy is used to access data on the node corresponding to the traffic data.
[0046] In this embodiment, the data link log analysis platform can store raw traffic, risk events, and interception logs, with a retention period of ≥180 days; based on the node information of the security logs and the traffic data, it can perform multi-dimensional correlation analysis to reconstruct the attack chain path; and automatically generate an audit report that conforms to the ISO 27001 standard.
[0047] Specifically, control policies for the nodes corresponding to the traffic data can be generated based on the attack chain path and the node information. The control policies include: 1) Hybrid access control: Combining role-based permissions and attribute-based permissions to dynamically adjust data access strategies; 2) PKI certificate chain verification: Binding devices, accounts and data operation behaviors through digital certificates.
[0048] The above control policies are built into compliance templates and the policy library is updated automatically. Before traffic data is identified as risky, access to traffic data can be controlled through these policies, reducing the risk of privacy leaks.
[0049] In addition, a unified policy configuration library and log repository can be established to achieve the following functions: 1) Policy configuration library: Stores metadata such as detection rules, blocking policies, and compliance templates, and supports version management and canary releases; 2) Log repository: Columnar storage (such as Apache Parquet) is used to optimize the retrieval efficiency of massive logs; 3) Feedback loop: The strategy recommendations output by the audit optimization layer are synchronized to each level of module through the configuration library, forming a self-evolving security ecosystem.
[0050] The security protection method of this application embodiment is implemented by a security protection network architecture, including a data acquisition layer, an analysis and monitoring layer, a security response layer, and an audit optimization layer. The data acquisition layer is used to collect traffic data of the target data link and the first information and second information corresponding to the traffic data. The analysis and monitoring layer is used to determine the third information of the traffic data based on the traffic data and preset regular expression rules. The third information is used to indicate whether the traffic data is sensitive data. The risk score corresponding to the traffic data is determined according to the second information. Based on at least one of the first information, the third information, and the risk score, it is determined whether the traffic data is risky data. The security response layer is used to perform security protection on the traffic data when the traffic data is risky data. The audit optimization layer generates the control strategy in the above embodiment. By implementing security protection methods through a secure network architecture, the accuracy and adaptability of sensitive data identification can be significantly improved, effectively addressing complex leakage scenarios such as data distortion and document tampering. Relying on a full-domain business link probe matrix deployment, seamless monitoring and risk visualization management of the entire data flow lifecycle can be achieved. Combined with real-time defense mechanisms and dynamic trust domain control models, attack response time is significantly shortened, while simultaneously ensuring flexible adaptation and automated implementation of compliance policies. Ultimately, a comprehensive, rapid-response, and precisely controlled data security protection system is formed, enhancing the confidentiality and integrity of privacy data while meeting multi-dimensional compliance requirements and providing reliable assurance for business continuity.
[0051] like Figure 2 As shown in the illustration, this application also provides a data security protection device 300, which includes: The acquisition module 301 is used to acquire traffic data of the target data link and first information and second information corresponding to the traffic data. The first information is used to indicate the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data. The first determining module 302 is used to determine the third information of the traffic data based on the traffic data and the preset regular expression rules, wherein the third information is used to indicate whether the traffic data is sensitive data; The second determining module 303 is used to determine the risk score corresponding to the traffic data based on the second information; The first judgment module 304 is used to determine whether the traffic data is risky data based on at least one of the first information, the third information, and the risk score; The security module 305 is used to provide security protection for the traffic data when the traffic data is risky data.
[0052] Optionally, the data security protection device 300 also includes: The acquisition module is used to acquire comparison data, wherein the traffic data is data collected at the first node, and the comparison data is data of the same type as the traffic data collected at the second node, wherein the second node is the node preceding the first node in the data transmission direction of the target data link; The first judgment module is used to determine whether the traffic data is risky data based on the similarity between the comparison data and the traffic data.
[0053] Optionally, the first judgment module includes: The first calculation submodule is used to calculate the first fingerprint string corresponding to the traffic data and the second fingerprint string corresponding to the comparison data based on the document fingerprint algorithm. The first determination submodule is used to determine the traffic data as risk data when the similarity between the first fingerprint string and the second fingerprint string is greater than a preset similarity threshold.
[0054] Optionally, the first judgment module 304 includes: The second determination submodule is used to determine that the traffic data is risky data when the third information indicates that the traffic data is sensitive data; The third determination submodule is used to determine that the traffic data is risky data when the risk score is greater than a preset risk score threshold. The fourth determination submodule is used to determine that the traffic data is risky data when the first information indicates that the interface, application, Internet Protocol IP address or account corresponding to the traffic data has risky behavior.
[0055] Optionally, the acquisition module 301 is also used for: Traffic data of the target node is collected by a network probe set on the target node of the target data link. The target node includes an application programming interface gateway, a microservice communication interface, a database access proxy, a cloud storage access point, a cross-cloud channel, a message middleware, and a terminal device software development kit (SDK).
[0056] Optionally, the data security protection device 300 also includes: A storage module is used to store security logs, the security logs including the traffic data and information for indicating the security protection; The third determining module is used to determine the attack chain path of the traffic data based on the node information of the security log and the traffic data; The generation module is used to generate a control policy for the node corresponding to the traffic data based on the attack chain path and the node information. The control policy is used to access the data of the node corresponding to the traffic data.
[0057] It should be noted that the data security protection device 300 provided in this application embodiment can achieve the following: Figure 2 The entire technical process of the data security protection method shown in the embodiment, and the same technical effect, will not be described again here to avoid repetition.
[0058] The data security protection device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. Non-mobile electronic devices can also be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0059] Optionally, such as Figure 3 As shown, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described data security protection method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0060] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0061] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described data security protection method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0062] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0063] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the data security protection method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0064] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0065] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0066] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data security protection method, characterized in that, The method includes: Collect traffic data of the target data link and first information and second information corresponding to the traffic data. The first information is used to indicate the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data. Based on the traffic data and preset regular expression rules, a third piece of information about the traffic data is determined, which is used to indicate whether the traffic data is sensitive data. The risk score corresponding to the traffic data is determined based on the second information; Based on at least one of the first information, the third information, and the risk score, determine whether the traffic data is risky data; If the traffic data is considered risky, security measures will be implemented for the traffic data.
2. The method as described in claim 1, characterized in that, After collecting traffic data from the target data link, the method further includes: Obtain comparison data, wherein the traffic data is data collected at the first node, and the comparison data is data of the same type as the traffic data collected at the second node, wherein the second node is the node preceding the first node in the data transmission direction of the target data link; Based on the similarity between the comparison data and the traffic data, it is determined whether the traffic data is risky data.
3. The method as described in claim 2, characterized in that, The step of determining whether the traffic data is risky data based on the similarity between the comparison data and the traffic data includes: Based on the document fingerprinting algorithm, calculate the first fingerprint string corresponding to the traffic data and the second fingerprint string corresponding to the comparison data; If the similarity between the first fingerprint string and the second fingerprint string is greater than a preset similarity threshold, the traffic data is determined to be risky data.
4. The method as described in claim 1, characterized in that, The determination of whether the traffic data is risky data based on at least one of the first information, the third information, or the risk score includes at least one of the following: If the third piece of information indicates that the traffic data is sensitive data, then the traffic data is determined to be risky data. If the risk score is greater than a preset risk score threshold, the traffic data is determined to be risky data. If the first information indicates that the interface, application, Internet Protocol IP address, or account corresponding to the traffic data has risky behavior, the traffic data is determined to be risky data.
5. The method according to any one of claims 1 to 4, characterized in that, The traffic data of the target data link being collected includes: Traffic data of the target node is collected by a network probe set on the target node of the target data link. The target node includes an application programming interface gateway, a microservice communication interface, a database access proxy, a cloud storage access point, a cross-cloud channel, a message middleware, and a terminal device software development kit (SDK).
6. The method according to any one of claims 1 to 4, characterized in that, After providing security protection for the traffic data, the method further includes: Store a security log, the security log including the traffic data and information for indicating the security protection; Based on the node information of the security logs and the traffic data, the attack chain path of the traffic data is determined; Based on the attack chain path and the node information, a control policy for the node corresponding to the traffic data is generated. The control policy is used to access data on the node corresponding to the traffic data.
7. A data security protection device, characterized in that, The device includes: The acquisition module is used to acquire traffic data of the target data link and first information and second information corresponding to the traffic data. The first information is used to indicate the characteristics of the device associated with the traffic data, and the second information is information associated with the behavior of accessing the traffic data. The first determining module is used to determine third information of the traffic data based on the traffic data and preset regular expression rules, wherein the third information is used to indicate whether the traffic data is sensitive data; The second determining module is used to determine the risk score corresponding to the traffic data based on the second information; The first judgment module is used to determine whether the traffic data is risky data based on at least one of the first information, the third information, and the risk score; A security module is used to provide security protection for the traffic data when the traffic data is risky data.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the data security protection method as described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the steps of the data security protection method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the data security protection method as described in any one of claims 1 to 6.