A data security detection method, device, equipment, medium and product

CN120915546BActive Publication Date: 2026-09-29CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511142190.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2026-09-29
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

[0005]本申请实施例提供一种数据安全检测方法、装置、设备、介质及产品,以解决现有的主流数据安全评估方案的安全评估结果滞后,且评估效率和评估准确性较低的技术问题

Benefits of technology

[0047]在本申请实施例中,首先获取待检测业务系统的与安全检测相关的数据,并基于数据构建知识图谱,基于知识图谱对数据进行分析,得到安全评估结果,知识图谱用于表示业务系统中的数据之间的关联关系,且能够直观地展示各个实体(数据)之间的关系,帮助安全分析人员快速理解待检测业务系统的结构和潜在的安全隐患,且通过知识图谱,可以将割裂的各个数据进行综合分析。由此,能够更有效地识别和评估安全风险,不仅提高了安全检测的效率和准确性,还为安全管理提供了科学依据,增强了整体的安全防护能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915546B_ABST
    Figure CN120915546B_ABST
Patent Text Reader

Abstract

The application provides a data security detection method, device, equipment, medium and product, the method comprises: obtaining the data related to security detection of the to-be-detected business system; constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph to obtain a security evaluation result, wherein the knowledge graph is used to represent the association relationship between the data in the business system. Wherein, the knowledge graph can intuitively show the relationship between each entity, help the security analyst quickly understand the structure and potential security risks of the to-be-detected business system, and through the knowledge graph, the fragmented data can be comprehensively analyzed. Thus, the security risks can be more effectively identified and evaluated, not only improving the efficiency and accuracy of security detection, but also providing a scientific basis for security management and enhancing the overall security protection capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular to a data security detection method, apparatus, equipment, medium and product. Background Technology

[0002] In today's rapidly developing information technology landscape, data security has become one of the major challenges facing enterprises and organizations. As business systems become increasingly complex, effectively assessing and monitoring the security status of systems to ensure the confidentiality, integrity, and availability of data has become a research hotspot and an urgent need for practical applications.

[0003] Currently, the main data security assessment schemes are as follows. First, many enterprises use manual assessment methods, creating a list of assessment items and having professional security assessors use assessment tools to evaluate the security of their business systems. While this method can identify potential risks to some extent, it is inefficient and susceptible to subjective influences due to its reliance on manual processes. Second, some enterprises have established security management platforms to monitor data classification and grading, interface usage status, etc., of their business systems in real time and audit operation logs. These platforms provide some real-time monitoring capabilities but often lack dynamic assessment of security risks and cannot respond promptly to changes in the current security status of the business systems. In addition, some machine learning-based security assessment methods have emerged in recent years, such as using Support Vector Machines (SVM) for multi-source data security assessment. These methods acquire classification features of multi-source data in real time, construct data flow supervision models, and thus achieve data security assessment. This method has certain advantages in improving assessment accuracy and training efficiency, but it still has some shortcomings.

[0004] In summary, existing mainstream data security assessment solutions generally suffer from the following technical problems: First, they cannot dynamically assess the current security risk status of business systems, resulting in delayed security assessment results; second, they lack the ability to automatically correlate and analyze security detection data from multiple dimensions of business systems, still requiring manual operation, which greatly increases the workload of security assessment personnel. These problems limit the efficiency and accuracy of data security assessment. Summary of the Invention

[0005] This application provides a data security detection method, apparatus, device, medium, and product to address the technical problems of existing mainstream data security assessment schemes having lagging security assessment results and low assessment efficiency and accuracy.

[0006] To solve the above-mentioned technical problems, this application is implemented as follows:

[0007] In a first aspect, embodiments of this application provide a data security detection method, the method comprising:

[0008] Obtain security testing-related data from the business system to be tested;

[0009] A knowledge graph is constructed based on the data, and the data is analyzed based on the knowledge graph to obtain a security assessment result. The knowledge graph is used to represent the relationships between the data in the business system.

[0010] Optionally, before constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result, the method further includes:

[0011] The data is converted to a serialized data structure and then stored. The amount of data in the data structure is dynamically updated based on the data.

[0012] Determine whether the data structure meets preset conditions, the preset conditions including at least one of the following: the continuous storage duration corresponding to the data structure is greater than or equal to the currently set first threshold, and the change rate of the data volume of the data structure is greater than or equal to the currently set second threshold.

[0013] If so, then the step of constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result is performed, wherein the knowledge graph is used to represent the relationship between the data in the business system.

[0014] Optionally, the first threshold and / or the second threshold are both variables, and the method further includes:

[0015] Based on historical security assessment results, adjust the first threshold and / or the second threshold, and determine the adjusted first threshold as the currently set first threshold, and / or determine the adjusted second threshold as the currently set second threshold.

[0016] Optionally, the historical security assessment results include: multiple historical security issues, which are ordered according to their chronological order of occurrence. Adjusting the first threshold based on the historical security assessment results includes:

[0017] In the sorted historical security issues, the time interval between any two consecutive historical security issues is determined, and the time interval of the next security issue is predicted based on a preset time interval prediction model.

[0018] If the time interval between the next impending security issue is greater than a first threshold, then the first threshold is updated to the time interval between the impending security issue.

[0019] Optionally, the historical security assessment results include: multiple historical security issues; adjusting the second threshold based on the historical security assessment results includes:

[0020] Determine the data volume change rate corresponding to each historical security issue, and based on the preset change rate prediction model, predict the expected data volume change rate when the next security issue is about to occur.

[0021] If the expected data volume change rate is greater than the second threshold, then the second threshold is updated to the expected data volume change rate.

[0022] Optionally, the data includes at least one of the following: account list, operation logs, system permissions, data permissions, and data classification and grading; constructing a knowledge graph based on the data includes:

[0023] Security element entities are extracted from the data, wherein the security element entities include at least one of the following: account list entity, operation log entity, system permissions, data permissions entity, and data classification and grading entity;

[0024] Tag the security element entities to mark their characteristic attributes;

[0025] Establish the association relationships between the aforementioned security element entities;

[0026] The knowledge graph is constructed based on the security element entities, the feature attributes, and the relationships.

[0027] Optionally, when the data includes: an account list and operation logs, and the account list includes accounts, and the security element entity includes an account list entity and an operation log entity, and the account list corresponding to the account list entity includes accounts, the security assessment results obtained by analyzing the data based on the knowledge graph include:

[0028] Determine the association rules between the security element entities from the knowledge graph;

[0029] The account and the operation log are evaluated to obtain account evaluation results and operation log evaluation results, respectively.

[0030] Based on the data, the association rules, the account evaluation results, and the operation log evaluation results, a comprehensive evaluation vector is generated;

[0031] The comprehensive evaluation vector is compared with the preset security evaluation rule set to obtain the security evaluation result.

[0032] Optionally, determining the association rules between the security element entities from the knowledge graph includes:

[0033] From the knowledge graph, at least one candidate rule F(x,y)→P(x,y) is generated, where F(x,y) is the association between the account and the operation log, and P(x,y) is the candidate rule between the account and the operation log that can be derived from F(x,y).

[0034] Calculate the support and confidence of each candidate rule;

[0035] Based on preset support and confidence thresholds and the support and confidence of each candidate rule, abnormal candidate rules are filtered out, and the candidate rules other than the abnormal candidate rules are determined as the association rules.

[0036] Optionally, the comprehensive evaluation vector is compared with a preset security evaluation rule set to obtain the security evaluation result, including:

[0037] The comprehensive evaluation vector is compared with each rule in the security evaluation rule set, wherein the security evaluation rule set includes at least one of the following: account overdue but not disabled status verification rule, system permission validity period verification rule, unauthorized access behavior identification rule, sensitive data access behavior identification rule, sensitive data access permission verification rule, and sensitive data access approval compliance verification rule;

[0038] If the comprehensive evaluation vector is consistent with the rule in the security evaluation rule set, then the score of the comprehensive evaluation vector on the rule is determined to be full.

[0039] If there is a discrepancy, the score of the comprehensive evaluation vector on the rule is determined to be zero.

[0040] All scores are aggregated, and the aggregated scores are determined as the security assessment result.

[0041] Secondly, embodiments of this application provide a data security detection device, the device comprising:

[0042] The acquisition module is used to acquire data related to security testing from the business system to be tested.

[0043] An execution module is used to construct a knowledge graph based on the data, analyze the data based on the knowledge graph, and obtain a security assessment result, wherein the knowledge graph is used to represent the relationship between the data in the business system.

[0044] Thirdly, embodiments of this application provide a network device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of a data security detection method as described in the first aspect.

[0045] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a data security detection method as described in the first aspect.

[0046] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of a data security detection method as described in the first aspect.

[0047] In this embodiment, data related to security testing from the business system under test is first acquired. A knowledge graph is then constructed based on this data. The data is analyzed using the knowledge graph to obtain security assessment results. The knowledge graph represents the relationships between data within the business system and visually displays the relationships between various entities (data), helping security analysts quickly understand the structure and potential security vulnerabilities of the business system under test. Furthermore, the knowledge graph allows for the comprehensive analysis of fragmented data. This enables more effective identification and assessment of security risks, improving not only the efficiency and accuracy of security testing but also providing a scientific basis for security management and enhancing overall security protection capabilities. Attached Figure Description

[0048] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0049] Figure 1 A flowchart illustrating a data security detection method provided in this application embodiment;

[0050] Figure 2 A flowchart illustrating a data security detection method provided in this application embodiment;

[0051] Figure 3 A structural block diagram of a data security detection system provided in this application embodiment;

[0052] Figure 4 A structural block diagram of a data security detection device provided in an embodiment of this application;

[0053] Figure 5This is a structural block diagram of a network device provided in an embodiment of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] Figure 1 This application illustrates a data security detection method according to an embodiment of the present application, such as... Figure 1 As shown, the method includes:

[0056] Step S101: Obtain security testing-related data from the business system to be tested;

[0057] Step S102: Construct a knowledge graph based on the data, analyze the data based on the knowledge graph, and obtain the security assessment results;

[0058] Knowledge graphs are used to represent the relationships between data in a business system.

[0059] It should be noted that, Figure 1 The method described aims to assess the security of a business system under test through systematic processing. The overall process first involves acquiring all security-related data from the business system, including at least one of the following: account list, operation logs, system permissions, data permissions, and data classification and grading. Then, a comprehensive security test is performed on this data. This stage specifically involves constructing a knowledge graph based on the acquired data. This knowledge graph is specifically designed to represent the relationships between data points in the business system, connecting scattered data entities into a structured network graph. Next, this knowledge graph is used for deep data analysis to identify potential threats, vulnerabilities, or anomaly patterns. Ultimately, this analysis process generates a clear security assessment result, providing a quantitative or qualitative judgment of the overall security status of the system. In summary, Figure 1 The method shown can efficiently integrate security-related data from business systems and improve the accuracy and comprehensiveness of detection through the association analysis function of knowledge graphs. It can help users quickly discover hidden security risks and strengthen the system's defense capabilities.

[0060] In one possible implementation, before constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result, the method further includes: converting the data format to obtain a serialized data structure and storing it, wherein the amount of data in the data structure is dynamically updated according to the data; determining whether the data structure meets preset conditions, the preset conditions including at least one of the following: the continuous storage duration corresponding to the data structure is greater than or equal to a currently set first threshold, and the change rate of the amount of data in the data structure is greater than or equal to a currently set second threshold; if so, then the steps of constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result are performed, wherein the knowledge graph is used to represent the relationship between data in the business system.

[0061] It should be noted that, firstly, the data to be processed must be converted into a serialized data structure through format conversion (e.g., converted into a binary stream or standardized text format). This serialization process ensures the data has compatibility for cross-system transmission and storage. Subsequently, the serialized result is persistently stored, and the amount of data contained in this data structure is not fixed but dynamically updated according to the real-time changes in the input data. That is, the system continuously tracks data operations such as addition, modification, or deletion and adjusts the storage scale in real time. Based on this, the system needs to pre-set two key operating parameters: firstly, determine the persistent storage duration for this data structure (i.e., the time span from initial storage to the current moment); secondly, calculate the rate of change of data volume in the data structure (usually referring to the percentage increase or decrease of data per unit time). The security detection mechanism is triggered when the continuous monitoring process meets any of the following conditions: when the measured persistent storage duration exceeds or equals the preset first threshold (time threshold), or when the calculated rate of change of data volume exceeds or equals the preset second threshold (change rate threshold).

[0062] Therefore, an intelligent data storage and detection scheduling mechanism can be established, which can not only force detection to discover potential risks after long-term data storage, but also respond to anomalies in a timely manner when the data volume changes drastically. This can significantly improve detection efficiency and resource utilization, while strengthening adaptability to dynamic business environments.

[0063] In one possible implementation, the first threshold and / or the second threshold are both variables, and the method further includes: adjusting the first threshold and / or the second threshold according to historical security assessment results, and determining the adjusted first threshold as the currently set first threshold, and / or determining the adjusted second threshold as the currently set second threshold.

[0064] It should be noted that both the first threshold (continuous storage duration threshold) and the second threshold (data volume change rate threshold) are designed as variables rather than fixed values. Before initiating the data security detection process, the system performs the following adaptive operations: based on historically accumulated security assessment results (such as the number of vulnerabilities, threat level, anomaly frequency, etc.), it analyzes the changing trends and risk characteristics through an algorithmic model, and then automatically adjusts the values ​​of the first threshold and / or the second threshold. After adjustment, the system immediately applies the adjusted first threshold as the currently set first threshold, and simultaneously applies the adjusted second threshold as the currently set second threshold. This process forms a closed-loop feedback. For example, when historical assessment results show that data stored for more than 5 days frequently presents high-risk threats, the system may shorten the first threshold from 7 days to 4 days; if periods of sudden data volume increases are often accompanied by covert attacks, the second threshold may be lowered to trigger detection more quickly.

[0065] In one possible implementation, historical security assessment results include: multiple historical security issues; these issues are ordered chronologically according to their occurrence; and the first threshold is adjusted based on the historical security assessment results, including:

[0066] In the sorted historical security issues, the time interval between any two consecutive historical security issues is determined, and the time interval of the next security issue is predicted based on a preset time interval prediction model.

[0067] If the time interval between the next impending security issue exceeds the first threshold, then the first threshold is updated to the time interval between the impending security issue.

[0068] It should be noted that when analyzing historical security assessment results, the system extracts multiple historical security issues (such as vulnerability outbreaks, malicious intrusions, data breaches, etc.) and calculates the time interval between adjacent historical security issues in chronological order (e.g., 72 hours between event A and event B, 36 hours between event B and event C). These time intervals are then input into a preset time interval prediction model (such as a long short-term memory network based on time series analysis or a probabilistic statistical model). The model analyzes the fluctuation patterns (such as periodicity and trends) of historical time intervals to predict the time interval of the next impending security issue (e.g., outputting "the security issue is most likely to occur within the next 24 hours"). Based on this prediction, the system performs a threshold adjustment: if the predicted time interval of the next security issue is greater than the currently set first threshold (e.g., the predicted interval is 48 hours, while the current threshold is 24 hours), it indicates that the system can withstand a longer storage period without increasing risk, and the first threshold is updated to this larger predicted time interval value.

[0069] Therefore, this method automatically synchronizes the timing of security detection with the dynamic risk trends of the business system. By predicting the occurrence cycle of security issues in real time and adjusting the threshold accordingly, the threshold can be proactively shortened during periods of high attack incidence to improve response speed (e.g., immediately adjusting the threshold to 24 hours when an attack is predicted to occur within 24 hours), and the threshold can be extended during periods of stability to avoid resource waste (e.g., relaxing the threshold when the predicted risk interval is 30 days), ultimately achieving precise matching between security detection resource allocation and threat trends.

[0070] In one possible implementation, the historical security assessment results include: multiple historical security issues; adjusting the second threshold based on the historical security assessment results includes: determining the data volume change rate corresponding to the occurrence of each historical security issue, predicting the expected data volume change rate when the next security issue is about to occur based on a preset change rate prediction model; if the expected data volume change rate is greater than the second threshold, then the second threshold is updated to the expected data volume change rate.

[0071] It should be noted that in the implementation of the data security detection method, multiple historical security issues recorded in the historical security assessment results (such as data breaches, unauthorized operation alerts, or malicious attack events) can be used to dynamically and adaptively adjust the second threshold. The specific process is as follows: First, for each historical security issue, trace the data volume change ratio of the business system corresponding to the time the event occurred; then, input the extracted sequence of data volume change ratios associated with historical security issues into a preset change ratio prediction model. By analyzing the distribution pattern, fluctuation trend, and event correlation of historical change ratios, predict the expected data volume change ratio corresponding to the next security issue; after obtaining this prediction value, if the predicted expected data volume change ratio is greater than the second threshold set in the current system (i.e., the triggering benchmark value for the data volume change ratio), then update the second threshold to this expected data volume change ratio value.

[0072] Therefore, when the prediction model determines that the next risk event will be accompanied by high data fluctuations (e.g., a sudden increase of 70% in the predicted data volume), the system will automatically raise the second threshold (e.g., from the original 30% to 70%), so that security detection is only triggered when the data fluctuates drastically, avoiding frequent false triggers caused by business peaks or normal data fluctuations, significantly reducing the resource overhead of invalid detection, and avoiding resource waste.

[0073] In one possible implementation, such as Figure 2 As shown, building a knowledge graph based on data includes:

[0074] Step S201: Extract security element entities from the data;

[0075] Among them, security element entities include at least one of the following: account list entity, operation log entity, system permissions, data permissions entity, and data classification and grading entity;

[0076] Step S202: Tag the safety element entities to mark their characteristic attributes;

[0077] Step S203: Establish the association relationships between security element entities;

[0078] Step S204: Construct a knowledge graph based on security element entities, feature attributes, and relationships.

[0079] It should be noted that, Figure 2 The method described starts with the raw data from the business system. First, it extracts security element entities, specifically covering key entity types such as account list entities (recording system account information), operation log entities (storing user behavior records), system permissions (defining access control rules), data permission entities (specifying the scope of data operations), and data classification and grading entities (identifying data sensitivity levels). Then, it annotates the extracted entities with feature attributes—that is, through tagging operations (e.g., labeling account entities with privileged user attributes, and data grading entities with confidentiality level attributes)—transforming the security characteristics of the entities into structured metadata. Based on this, the system proactively mines and establishes relationships between security element entities (e.g., identifying the causal relationship of user A modifying confidential data through the operation log entity). Finally, it integrates security element entities, annotated feature attributes, and relationships between entities to construct a multi-layered networked knowledge graph. This graph visually maps the interactions of security elements within the business system through a topological structure of nodes (entities) and edges (relationships).

[0080] Therefore, fragmented information can be transformed into a computable relational network by structuring and parsing dispersed security data. Entity coverage of the knowledge graph (accounts / permissions / logs / data hierarchies, etc.) ensures that no core security elements are omitted, feature attribute annotation (such as marking high-privilege accounts) enables entity risk profiling, and relationship mining (such as the binding of log operations and data access) exposes potential attack paths. This improves the efficiency, comprehensiveness, and accuracy of subsequent data security assessments.

[0081] In one possible implementation, when the data includes an account list and operation logs, and the account list includes accounts, and the security element entities include an account list entity and an operation log entity, and the account list corresponding to the account list entity includes accounts, the data is analyzed based on a knowledge graph to obtain security assessment results, including: determining the association rules between security element entities from the knowledge graph; evaluating accounts and operation logs to obtain account assessment results and operation log assessment results respectively; generating a comprehensive assessment vector based on the data, association rules, account assessment results, and operation log assessment results; and comparing the comprehensive assessment vector with a preset security assessment rule set to obtain the security assessment result.

[0082] It should be noted that the system first extracts association rules between security element entities from the constructed knowledge graph (e.g., logical patterns such as "privileged accounts accessing sensitive data outside of working hours" or "the same account logging in from multiple locations in a short period of time"). These rules reflect the causal or statistical correlation characteristics between entities. Then, two independent assessments are executed in parallel: risk analysis of account information (e.g., detecting weak passwords and abnormal permission allocation) to generate account assessment results, and behavioral analysis of operation logs (e.g., identifying brute-force attacks and abnormal command sequences) to generate operation log assessment results. Next, the system integrates four types of inputs: raw data, mined association rules, account assessment results, and operation log assessment results, constructing a comprehensive assessment vector through a feature vectorization engine. Finally, this vector is matched and compared with a preset security assessment rule set to output a quantitative security assessment result (e.g., "High risk: a privileged account was detected performing data export operations during non-compliant periods").

[0083] Therefore, by integrating the association rule analysis of knowledge graphs, the dual-path independent evaluation of accounts and operation logs, and the generation of a comprehensive evaluation vector through multi-dimensional data fusion, and finally using a preset rule set to achieve accurate matching decisions, the accuracy of data security assessment can be effectively improved.

[0084] In one possible implementation, determining the association rules between security element entities from the knowledge graph includes: generating at least one candidate rule F(x,y)→P(x,y) from the knowledge graph, where F(x,y) is the association relationship between accounts and operation logs, and P(x,y) is a candidate rule for accounts and operation logs that can be derived from F(x,y); calculating the support and confidence of each candidate rule; based on preset support and confidence thresholds and the support and confidence of each candidate rule; filtering out abnormal candidate rules, and determining the candidate rules other than the abnormal candidate rules as association rules.

[0085] It should be noted that the system first extracts the original association between accounts and operation logs from the knowledge graph as the basic condition F(x,y), and derives the implicit candidate rules P(x,y) that it may trigger, forming at least one candidate rule F(x,y)→P(x,y). Then, it quantitatively calculates the support and confidence of each candidate rule. Finally, it filters all candidate rules based on preset support and confidence thresholds—only retaining candidate rules with both support and confidence above the thresholds, while eliminating those that do not meet the criteria (i.e., abnormal candidate rules). Ultimately, the qualified candidate rules are formally established as valid association rules. This allows for the selection of strongly correlated rules from a massive number of candidate rules, significantly improving the accuracy of security detection while eliminating random noise interference, and providing quantifiable and interpretable reasoning for subsequent comprehensive evaluation.

[0086] In one possible implementation, comparing the comprehensive evaluation vector with a preset security evaluation rule set to obtain the security evaluation result includes: comparing the comprehensive evaluation vector with each rule in the security evaluation rule set, wherein the security evaluation rule set includes at least one of the following: account expired but not disabled status verification rule, system permission validity period verification rule, unauthorized access behavior identification rule, sensitive data access behavior identification rule, sensitive data access permission verification rule, and sensitive data access approval compliance verification rule; if the comprehensive evaluation vector matches the rule in the security evaluation rule set, the score of the comprehensive evaluation vector on that rule is determined to be full marks; if they do not match, the score of the comprehensive evaluation vector on that rule is determined to be zero marks; summarizing all scores, and determining the summed scores as the security evaluation result.

[0087] It should be noted that the system will compare the comprehensive evaluation vector (an array of quantitative indicators including account risk, log anomalies, and rule weights) with each rule in the preset security evaluation rule set. This rule set covers multi-dimensional verification standards, specifically including at least one core rule such as the account overdue but not disabled status verification rule (detecting invalid accounts that have not been deactivated for a long time), system permission validity period verification rule (verifying whether the permission allocation is within the validity period), unauthorized access behavior identification rule (identifying operations beyond the scope of permissions), sensitive data access behavior identification rule (monitoring operations involving classified data), sensitive data access permission verification rule (checking whether the visitor has the necessary permissions), and sensitive data access approval compliance verification rule (verifying the compliance of the approval process). During the comparison process, if the comprehensive evaluation vector fully meets the requirements of the rule currently being compared, the score for that rule is determined to be full marks; if the vector does not match the rule conditions, the score for that rule is determined to be zero marks. Finally, the system summarizes the scores of all rules (e.g., 4 out of 6 rules get full marks and 2 get zero marks), and outputs this summary score as the final security assessment result (e.g., a total score of 80 / 100).

[0088] The data security detection method shown in the embodiments of this application is now described from a system perspective, such as... Figure 3 As shown, the system includes: a business system integration module, a rule base, and a dynamic self-inspection center.

[0089] 1. Business System Integration Module: This module is deployed in the backend server cluster and is used to receive various security detection data from the business system.

[0090] The business system integration module includes:

[0091] (1) Data Acquisition Unit: The security testing data recorded by the business system is formed into batch data files and periodically transmitted to the business system interface module. Typical data includes, but is not limited to, account lists, operation logs, system permissions, and data permission data, which are periodically transmitted to the dynamic self-inspection center through the interface.

[0092] (2) Data Fetching Unit: Automatically fetches security testing data using business system URLs and database tables. Typical application scenarios include, but are not limited to: For password management, specific password management requirements can be configured, such as whether passwords are encrypted, whether passwords are updated regularly, and password complexity requirements. The business system provides the password management URL and password storage table. For data display pages, fetch the fields displayed on the page and display the data.

[0093] (3) Self-testing and adaptive unit: The security detection data is saved as a serialized data structure A', and dynamic detection is controlled by a timer and a threshold timer tmax and a threshold Smax. When the storage time ta>tmax or the data change ratio Sa>Smax, the transmission mechanism is started and the data structure A' is transmitted to the data receiving unit through an encrypted channel.

[0094] (4) Data receiving unit: Obtain the security detection data structure A', establish internal identifiers, perform structured aggregation, establish a hash table with the account as the primary key, establish the relationship between data such as account list, operation log, system permissions, data permissions, data classification and grading, and transmit the data to the "Dynamic Self-Inspection Center".

[0095] 2. Rule base module

[0096] The rule base module includes:

[0097] (1) Rule configuration unit: used to configure adaptive dynamic detection rules, including the initial values ​​of timer tmax and threshold Smax, and the preset security assessment rule vector set I (including whether system permissions have expired, whether access is unauthorized, whether access is sensitive data, and whether access to sensitive data requires treasury approval).

[0098] (2) Rule adaptive adjustment unit: After the rules are configured, the timer tmax and threshold Smax are automatically adjusted based on the data security detection results using intelligent adaptive technology.

[0099] (3) Establish an adaptive dynamic detection time interval model:

[0100] Tn'=β0+β1Tn-1+β2Tn-2+...+βpTn-p+εt

[0101] Where Tn' represents the time interval for predicting the next detection item problem, Tn-1 represents the time interval for the (n-1)th actual detection item problem, p is the number of the most recent time intervals to be counted, β0, β1, β2...βp are parameters, and εt is the error term, which is adjusted according to the actual business system.

[0102] (4) If Tn'>tmax, then set tmax=Tn'. When the storage time ta of the data structure A of the "self-test adaptive unit" is greater than tmax, the transmission mechanism is started, and the data structure A is transmitted to the data receiving unit through the encrypted channel for security assessment.

[0103] 3. Dynamic Self-Inspection Center: Deployed in the backend server cluster. The Dynamic Self-Inspection Center categorizes and summarizes the initial data structure A extracted from the data receiving unit of the business system integration module, and transmits it to the knowledge graph module to construct a knowledge graph. Each time a dynamic data structure A' is collected, it is transmitted to the account evaluation unit and the operation log evaluation unit. The resulting security evaluation results are sent to the data calculation module, and the calculation results are sent to the evaluation module for final evaluation to obtain the final evaluation result.

[0104] The dynamic self-test center includes the following modules:

[0105] (1) Knowledge graph module, the knowledge graph module is used for

[0106] 1) Entity extraction: This involves converting account lists, operation logs, system permissions, data permissions, and data classification / grading into data entity vector sets. Account list entities include account, validity period, role, etc.; operation log data entities include IP, port, account, time, etc.; system permission entities include account, role, menu, valid start time, valid end time, etc.; data permissions include account, database name, table name, data field name, valid start time, valid end time, etc.; data classification / grading entities include database name, table name, data field name, data security level, data security category, and whether the data is important, etc.

[0107] 2) For the tagged data entity vectors, perform association matching on the extracted data entities: establish the association relationship between system permissions, data permissions, account list, and operation logs through accounts; establish the association relationship between operation logs, data permissions, and data classification and grading through data table names and field names.

[0108] 3) Based on accounts, logs, system permissions, data permissions, and data classification and grading entities, a knowledge representation K(Node, Prop, Edge) is used to construct a system operation knowledge graph. Node represents data entities such as accounts, logs, system permissions, data permissions, and data classification and grading; Prop represents the attribute of the data entity; and Edge represents the relationship between two data entities. Candidate rules are generated from the knowledge graph using the Association Rule Mining (AMIE) method. Let x be an account data entity, y be a log data entity, F(x,y) be the relationship between the two data entities, and P(x,y) be the relationship between the two data entities that can be deduced from F(x,y), i.e., (F(x,y)→P(x,y), which is the knowledge graph vector).

[0109] 4) For each candidate rule, calculate its support and confidence. Define the rule's support as `supp = count(F(x,y) → P(x,y))`, which is the number of distinct topic-object pairs in all instantiated headers. Define the rule's confidence as `conf = supp / count(F(x,y))`, which is the ratio of support to header relationship pairs. Set support and confidence thresholds `ZCmax` and `ZXmax`, and adjust these thresholds to retain rules with support and confidence values ​​greater than the thresholds, thus filtering out meaningless rules.

[0110] (2) Data computation module: The knowledge graph is transformed into a low-dimensional vector using graph embedding technology. The computation logic in the data computation module is designed to capture the relationships and attributes between nodes in the graph. The computation logic is as follows:

[0111] 1) Account Evaluation Unit: This module evaluates accounts and obtains an account evaluation result vector set ZH. p (Is the account disabled? Has the account expired? Has the account not been logged in for 3 months?)

[0112] 2) Operation Log Evaluation Unit: This module evaluates the operation logs and obtains an operation log evaluation result vector set CZ. d (Whether to access data, whether to access sensitive data)

[0113] 3) Based on the association rules of the knowledge graph module, the account evaluation result vector set ZH is comprehensively considered in the first calculation unit. pThe business system integration module uses data structure A', configures calculation rules Qzh, performs summary calculations, and generates account violation vector ZHWG. x1;

[0114] ZHWG x1 =Qzh(knowledge graph vector, ZH) p ,A')

[0115] 4) In the first calculation unit, taking into account the operation log evaluation result vector set CZd and the data structure A of the business system interface module, the calculation rule Qcz is configured, and the summary calculation is performed to generate the operation violation vector CZWG. xl ;

[0116] CZWG xl =Qcz(Knowledge Graph Vector, CZ) d ,A')

[0117] 5) The comprehensive vector Zh obtained through the second calculation unit xs The account violation vector is ZHWG. xl With operation violation vector CZWG xl Perform correlation calculations and G analysis to obtain the comprehensive vector Zh xs

[0118] Zh xs =G(ZH p CZ d )

[0119] 6) In the evaluation module, the comprehensive vector Zh calculated by the data calculation module is used. xs The rule set (including the following rules) is compared with the rule configuration unit preset in the rule base module. The comparison is made to determine whether each rule in the rule set (including the following rules) is consistent with each rule in the rule set (including the following rules): whether the account has expired and not been disabled, whether the system permissions have expired, whether the access is unauthorized, whether the sensitive data has been accessed, whether the access to sensitive data has been authorized, and whether the access to sensitive data has been approved by the treasury. If they are consistent, the item scores 10 points, which means that the security control requirements are met and there is no data security risk. If they are inconsistent, the item scores 0 points, which means that the security control requirements are not met and there is a data security risk.

[0120] Therefore, the system provided in this application embodiment controls the frequency of data security detection through a dual mechanism of timer Tmax and business system data change threshold Smax. Furthermore, the system incorporates an adaptive algorithm to achieve automated and dynamic management of data security detection. Specifically, for business systems with minimal overall data changes and few security risks, the system automatically reduces the security detection frequency; conversely, for business systems with minimal overall data changes or frequent security risks, the system automatically increases the security detection frequency. In addition, information such as account details, system permissions, data permissions, operation logs, and data classification and grading is comprehensively processed and stored in the knowledge graph module. By extracting data entities and constructing a knowledge graph, the system can identify potential association rules, thereby assisting in identifying potential security risks. This enables a comprehensive assessment of system data risks and improves the effectiveness of data security management.

[0121] Figure 4 A data security detection device is shown, the device 40 comprising:

[0122] The acquisition module 401 is used to acquire data related to security testing from the business system to be tested.

[0123] Execution module 402 is used to construct a knowledge graph based on data, analyze the data based on the knowledge graph, and obtain security assessment results. The knowledge graph is used to represent the relationship between data in the business system.

[0124] In one possible implementation, execution module 402 is further configured to: convert the data into a serialized data structure and store it before constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result; determine whether the data structure meets preset conditions, which include at least one of the following: the continuous storage duration corresponding to the data structure is greater than or equal to a currently set first threshold, and the change rate of the data volume of the data structure is greater than or equal to a currently set second threshold; if so, execute the steps of constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result, wherein the knowledge graph is used to represent the relationship between data in the business system.

[0125] In one possible implementation, the first threshold and / or the second threshold are both variables. The execution module 402 is further configured to adjust the first threshold and / or the second threshold based on historical security assessment results, and determine the adjusted first threshold as the currently set first threshold, and / or determine the adjusted second threshold as the currently set second threshold.

[0126] In one possible implementation, the historical security assessment results include: multiple historical security issues, which are ordered according to their chronological order of occurrence. Execution module 402 is also used to determine the time interval between every two consecutive historical security issues in the ordered set of issues, and to predict the time interval of the next security issue based on a preset time interval prediction model.

[0127] If the time interval between the next impending security issue exceeds the first threshold, then the first threshold is updated to the time interval between the impending security issue.

[0128] In one possible implementation, the historical security assessment results include: multiple historical security issues; the execution module 402 is also used to determine the data volume change rate corresponding to the occurrence of each historical security issue, and predict the expected data volume change rate when the next security issue is about to occur based on a preset change rate prediction model.

[0129] If the expected data volume change rate is greater than the second threshold, then the second threshold will be updated to the expected data volume change rate.

[0130] In one possible implementation, the data includes at least one of the following: account list, operation log, system permissions, data permissions, and data classification and grading; the execution module 402 is further used to extract security element entities from the data, wherein the security element entities include at least one of the following: account list entity, operation log entity, system permissions, data permissions entity, and data classification and grading entity;

[0131] Tag the safety element entities to mark their characteristic attributes;

[0132] Establish the relationships between security element entities;

[0133] A knowledge graph is constructed based on security element entities, characteristic attributes, and relationships.

[0134] In one possible implementation, when the data includes: an account list and an operation log, and the account list includes accounts, and the security element entity includes an account list entity and an operation log entity, and the account list corresponding to the account list entity includes accounts, the execution module 402 is further used to determine the association rules between security element entities from the knowledge graph.

[0135] The account and operation logs are evaluated to obtain the account evaluation results and operation log evaluation results, respectively.

[0136] A comprehensive evaluation vector is generated based on data, association rules, account evaluation results, and operation log evaluation results.

[0137] The comprehensive evaluation vector is compared with the preset security evaluation rule set to obtain the security evaluation result.

[0138] In one possible implementation, the execution module 402 is further configured to generate at least one candidate rule F(x,y)→P(x,y) from the knowledge graph, where F(x,y) is the association between the account and the operation log, and P(x,y) is the candidate rule between the account and the operation log that can be derived from F(x,y).

[0139] Calculate the support and confidence of each candidate rule;

[0140] Based on preset support and confidence thresholds and the support and confidence of each candidate rule, abnormal candidate rules are filtered out, and the candidate rules other than the abnormal candidate rules are identified as association rules.

[0141] In one possible implementation, the execution module 402 is further configured to compare the comprehensive evaluation vector with each rule in the security evaluation rule set, wherein the security evaluation rule set includes at least one of the following: account overdue but not disabled status verification rule, system permission validity period verification rule, unauthorized access behavior identification rule, sensitive data access behavior identification rule, sensitive data access permission verification rule, and sensitive data access approval compliance verification rule.

[0142] If the comprehensive evaluation vector is consistent with the rules in the security evaluation rule set, then the score of the comprehensive evaluation vector on the rule is determined to be full.

[0143] If there is a discrepancy, the score of the comprehensive evaluation vector on the rule is determined to be zero.

[0144] All scores are aggregated, and the aggregated scores are used as the security assessment result.

[0145] In summary, the embodiments of this application can adaptively adjust the detection frequency of detection and evaluation items according to the current security risk status of the business system. Furthermore, through knowledge graphs, it performs deep learning analysis on the fragmented security evaluation values ​​to identify behavioral characteristics. This enables the computer to realize the multi-dimensional comprehensive security risk monitoring function of humans, which can more effectively identify and evaluate security risks. This not only improves the efficiency and accuracy of security detection but also provides a scientific basis for security management and enhances the overall security protection capability.

[0146] This application provides a network device 50, such as... Figure 5 As shown, the network device 50 includes a processor 501, a memory 502, and a program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the steps of a data security detection method as shown in the above embodiment.

[0147] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of a data security detection method as shown in the above embodiments, achieving the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0148] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the steps of the data security detection method shown in the above embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0151] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A data security detection method, characterized in that, The method includes: Obtain security testing-related data from the business system to be tested; A knowledge graph is constructed based on the data, and the data is analyzed based on the knowledge graph to obtain a security assessment result. The knowledge graph is used to represent the relationship between the data in the business system. The data includes at least one of the following: account list, operation logs, system permissions, data permissions, and data classification and grading; the knowledge graph constructed based on the data includes: Security element entities are extracted from the data, wherein the security element entities include at least one of the following: account list entity, operation log entity, system permissions, data permissions entity, and data classification and grading entity; Tag the security element entities to mark their characteristic attributes; Establish the association relationships between the aforementioned security element entities; The knowledge graph is constructed based on the security element entities, the feature attributes, and the relationships. When the data includes an account list and operation logs, and the account list includes accounts, and the security element entity includes an account list entity and an operation log entity, and the account list corresponding to the account list entity includes accounts, the security assessment results obtained by analyzing the data based on the knowledge graph include: Determine the association rules between the security element entities from the knowledge graph; The account and the operation log are evaluated to obtain account evaluation results and operation log evaluation results, respectively. Based on the data, the association rules, the account evaluation results, and the operation log evaluation results, a comprehensive evaluation vector is generated; The comprehensive evaluation vector is compared with a preset set of security evaluation rules to obtain the security evaluation result; Determining the association rules between the security element entities from the knowledge graph includes: From the knowledge graph, at least one candidate rule F(x,y)→P(x,y) is generated, where F(x,y) is the association between the account and the operation log, and P(x,y) is the candidate rule between the account and the operation log that can be derived from F(x,y). Calculate the support and confidence of each candidate rule; Based on preset support and confidence thresholds and the support and confidence of each candidate rule, abnormal candidate rules are filtered out, and the candidate rules other than the abnormal candidate rules are determined as the association rules.

2. The method according to claim 1, characterized in that, Before constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result, the method further includes: The data is converted to a serialized data structure and then stored. The amount of data in the data structure is dynamically updated based on the data. Determine whether the data structure meets preset conditions, the preset conditions including at least one of the following: the continuous storage duration corresponding to the data structure is greater than or equal to the currently set first threshold, and the change rate of the data volume of the data structure is greater than or equal to the currently set second threshold. If so, then the step of constructing a knowledge graph based on the data, analyzing the data based on the knowledge graph, and obtaining a security assessment result is performed, wherein the knowledge graph is used to represent the relationship between the data in the business system.

3. The method according to claim 2, characterized in that, The first threshold and / or the second threshold are both variables, and the method further includes: Based on historical security assessment results, adjust the first threshold and / or the second threshold, and determine the adjusted first threshold as the currently set first threshold, and / or determine the adjusted second threshold as the currently set second threshold.

4. The method according to claim 3, characterized in that, The historical security assessment results include: multiple historical security issues, which are ordered chronologically according to their occurrence. Adjusting the first threshold based on these historical security assessment results includes: In the sorted historical security issues, the time interval between any two consecutive historical security issues is determined, and the time interval of the next security issue is predicted based on a preset time interval prediction model. If the time interval between the next impending security issue is greater than a first threshold, then the first threshold is updated to the time interval between the impending security issue.

5. The method according to claim 3, characterized in that, The historical security assessment results include: multiple historical security issues; adjusting the second threshold based on the historical security assessment results includes: Determine the data volume change rate corresponding to each historical security issue, and based on the preset change rate prediction model, predict the expected data volume change rate when the next security issue is about to occur. If the expected data volume change rate is greater than the second threshold, then the second threshold is updated to the expected data volume change rate.

6. The method according to claim 1, characterized in that, The comprehensive evaluation vector is compared with a preset security evaluation rule set to obtain the security evaluation result, including: The comprehensive evaluation vector is compared with each rule in the security evaluation rule set, wherein the security evaluation rule set includes at least one of the following: account overdue but not disabled status verification rule, system permission validity period verification rule, unauthorized access behavior identification rule, sensitive data access behavior identification rule, sensitive data access permission verification rule, and sensitive data access approval compliance verification rule; If the comprehensive evaluation vector is consistent with the rule in the security evaluation rule set, then the score of the comprehensive evaluation vector on the rule is determined to be full. If there is a discrepancy, the score of the comprehensive evaluation vector on the rule is determined to be zero. All scores are aggregated, and the aggregated scores are determined as the security assessment result.

7. A data security detection device, characterized in that, The device includes: The acquisition module is used to acquire data related to security testing from the business system to be tested. An execution module is used to construct a knowledge graph based on the data, analyze the data based on the knowledge graph, and obtain a security assessment result, wherein the knowledge graph is used to represent the relationship between the data in the business system; The data includes at least one of the following: account list, operation logs, system permissions, data permissions, and data classification and grading; the knowledge graph constructed based on the data includes: Security element entities are extracted from the data, wherein the security element entities include at least one of the following: account list entity, operation log entity, system permissions, data permissions entity, and data classification and grading entity; Tag the security element entities to mark their characteristic attributes; Establish the association relationships between the aforementioned security element entities; The knowledge graph is constructed based on the security element entities, the feature attributes, and the relationships. When the data includes an account list and operation logs, and the account list includes accounts, and the security element entity includes an account list entity and an operation log entity, and the account list corresponding to the account list entity includes accounts, the security assessment results obtained by analyzing the data based on the knowledge graph include: Determine the association rules between the security element entities from the knowledge graph; The account and the operation log are evaluated to obtain account evaluation results and operation log evaluation results, respectively. Based on the data, the association rules, the account evaluation results, and the operation log evaluation results, a comprehensive evaluation vector is generated; The comprehensive evaluation vector is compared with a preset set of security evaluation rules to obtain the security evaluation result; Determining the association rules between the security element entities from the knowledge graph includes: From the knowledge graph, at least one candidate rule F(x,y)→P(x,y) is generated, where F(x,y) is the association between the account and the operation log, and P(x,y) is the candidate rule between the account and the operation log that can be derived from F(x,y). Calculate the support and confidence of each candidate rule; Based on preset support and confidence thresholds and the support and confidence of each candidate rule, abnormal candidate rules are filtered out, and the candidate rules other than the abnormal candidate rules are determined as the association rules.

8. A network device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a data security detection method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of a data security detection method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of a data security detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Network security situation early warning method and system based on knowledge graph

    CN119603058A