Data management optimization method and system

By collecting and processing data from each sub-node, and utilizing a self-built risk identification model and alarm mechanism, the system has solved the coordination difficulties and resource waste issues in enterprise-level data security assessment and analysis, achieved real-time risk identification and optimized data processing, and improved system performance and response speed.

CN120915485APending Publication Date: 2025-11-07CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510848033.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies cannot conduct data security risk assessment and analysis at the enterprise-wide level, making it difficult to detect threats across sub-nodes, resulting in coordination difficulties, long response times, serious resource waste, increased data processing burden, impacting system performance and response speed, and failing to detect new risks in a timely manner.

Method used

By collecting logs and business data from each sub-node, calculating the data collection success rate and latency assessment value, performing multiple types of data processing, and using a self-built risk event identification model to identify risk events and calculate risk scores, an alarm mechanism is triggered when the score exceeds a threshold.

Benefits of technology

It enables synchronous risk event identification and real-time alarms in a multi-child node environment, optimizes the data collection and processing process, improves data collection rate and processing efficiency, accurately assesses risk events, and reduces resource waste and response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915485A_ABST
    Figure CN120915485A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a data management optimization method and system. The method comprises the following steps: based on an acquisition strategy, acquiring log and business data of each child node, and calculating a data acquisition success rate and a data acquisition delay evaluation value to measure an acquisition condition; transmitting the collected logs and service data of each child node to a central node, carrying out multi-class data processing, and calculating and determining a quantitative index of a data processing delay condition so as to optimize a data processing process; monitoring the data after processing the multiple types of data in real time, identifying a risk event by using a self-constructed risk event identification model, and calculating a risk score of the risk event, including analyzing and calculating a risk monitoring result so as to update a monitoring rule and a feature library of the risk event identification model; and when the calculated risk score exceeds a preset threshold value, an alarm mechanism is triggered, and alarm information is pushed to safety management personnel of the related child nodes. According to the invention, the multi-node data acquisition and data processing process is effectively optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a data management optimization method and system. BACKGROUND

[0002] The existing decentralized security management lacks a global perspective, each sub-node operates independently, cannot carry out data security risk assessment and analysis from the overall level of the enterprise, and cannot detect threats across sub-nodes; the response efficiency is low, when a risk event involves multiple sub-nodes, it needs to be handled at each sub-node, coordination is difficult and the response time is long, which can easily expand the risk event; meanwhile, there is a problem of resource waste, each sub-node independently deploys a security monitoring system, resulting in repeated investment in hardware resources, software licenses, etc., increasing the operating cost of the enterprise.

[0003] The simple centralized security management has a large data processing pressure, directly transmitting raw logs and data to the central node can occupy a high network bandwidth, increase the data processing burden of the central node, and can affect the system performance and response speed; the data quality is poor, the format of the raw logs and data is not unified, contains a large amount of noise and redundant information, direct storage and analysis can affect the accuracy of risk monitoring; the rule flexibility is lacking, the preset rules are difficult to adapt to complex and variable risk scenarios, and new risks cannot be detected in time.

[0004] Therefore, there is an urgent need for a new data management optimization method and system to solve the above problems. SUMMARY

[0005] The present application aims to provide a data management optimization method and system to solve the technical problems in the prior art that the data security risk cannot be assessed and analyzed from the overall level of the enterprise, it is difficult to detect threats across sub-nodes, when a risk event involves multiple sub-nodes, multi-node coordination is difficult and the response time is long, which can lead to low response efficiency and easily expand the risk event; each sub-node independently deploys a security monitoring system, resulting in repeated investment in hardware resources, software licenses, etc., which causes resource waste and increases the operating cost of the enterprise; directly transmitting raw logs and data to the central node increases the data processing burden of the central node, which can affect the system performance and response speed, the data quality is poor, direct storage and analysis can affect the accuracy of risk monitoring; the existing method is difficult to adapt to complex and variable risk scenarios, and new risks cannot be detected in time, etc. The technical problems to be solved by the present application are solved by the following technical solutions.

[0006] The first aspect of the present application provides a data management optimization method, comprising: collecting log and business data of each sub-node based on a collection strategy, calculating a data collection success rate and a data collection delay evaluation value to measure the collection situation; transmitting the collected log and business data of each sub-node to a central node and performing multi-type data processing to calculate a quantitative indicator of data processing delay condition to optimize the data processing process; performing real-time monitoring on the data after multi-type data processing, identifying a risk event by using a self-constructed risk event identification model, and calculating a risk score of the risk event, wherein the risk monitoring result is analyzed and calculated to update the monitoring rules and feature library of the risk event identification model; when the calculated risk score exceeds a preset threshold, triggering an alarm mechanism and pushing alarm information to a security management personnel of a related sub-node.

[0007] The second aspect of the present application provides a data management optimization system which executes the data management optimization method of the first aspect of the present application, comprising: in the data collection layer, collecting log and business data of each sub-node based on a collection strategy, calculating a data collection success rate and a data collection delay evaluation value to measure the collection situation; in the data aggregation layer, transmitting the collected log and business data of each sub-node to a central node, performing multi-type data processing in the data processing layer, and calculating a quantitative indicator of data processing delay condition to optimize the data processing process; in the risk monitoring layer, performing real-time monitoring on the data after multi-type data processing, identifying a risk event by using a self-constructed risk event identification model, and calculating a risk score of the risk event, wherein the risk monitoring result is analyzed and calculated to update the monitoring rules and feature library of the risk event identification model; in the alarm issuing layer, when the calculated risk score exceeds a preset threshold, triggering an alarm mechanism and pushing alarm information to a security management personnel of a related sub-node.

[0008] The third aspect of the present application provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the data management optimization method of the first aspect of the present application.

[0009] The fourth aspect of the present application provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the data management optimization method of the first aspect of the present application.

[0010] The embodiments of the present application have the following advantages:

[0011] Compared with the prior art, the application is based on the collection strategy, collects the logs and service data of each sub-node, calculates the data collection success rate and data collection delay evaluation value to measure the collection situation, and can effectively improve the data collection rate of the sub-node; the collected logs and service data of each sub-node are transmitted to the central node, and multi-type data processing is performed, a quantitative index of data processing delay condition is calculated and determined, and the data processing process can be effectively optimized; the data after multi-type data processing is monitored in real time, a risk event recognition model is recognized by using the self-constructed risk event recognition model, the risk score of the risk event is calculated, the monitoring rules and feature library of the risk event recognition model are updated by analyzing and calculating the risk monitoring result; when the calculated risk score exceeds the preset threshold, an alarm mechanism is triggered, and the alarm information is pushed to the security management personnel of the related sub-node, the data collection process and data processing process of the multiple sub-nodes are optimized, and synchronous risk event recognition and real-time alarm are effectively realized.

[0012] In addition, the data processing delay index is calculated to evaluate the processing efficiency of the data analysis layer on the data, so as to optimize the data processing process; by continuously feeding back the related parameters of data processing to the background server in real time for statistics and calculation, the quantitative index of data processing delay condition can be accurately output, and the system running efficiency can be accurately evaluated.

[0013] In addition, by calculating the risk score of the to-be-processed risk event according to the severity, influence range and occurrence frequency of the risk event and other factors, the to-be-processed risk event can be accurately quantified. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a flowchart of an example of the data management optimization method of the application;

[0015] Figure 2 is a framework diagram of an application example of the data management optimization method of the application;

[0016] Figure 3 is a structure diagram of an example of the data management optimization system of the application;

[0017] Figure 4 is a structure diagram of an electronic equipment embodiment according to the application;

[0018] Figure 5 is a structure diagram of a computer readable medium embodiment according to the application. DETAILED DESCRIPTION

[0019] Exemplary embodiments of the present application will now be described more fully hereinafter with reference to the accompanying drawings. The exemplary embodiments of the present application may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these exemplary embodiments of the present application are provided so that this disclosure will be thorough and complete, and will fully convey the inventive concept to those skilled in the art. Like reference numerals refer to like elements throughout the specification.

[0020] In the case where the technical concept of the present application is met, the features, structures, characteristics or other details described in a certain embodiment can not exclude being combined in a suitable manner in one or more other embodiments.

[0021] In the description of the specific embodiments, the features, structures, characteristics or other details described in the present application are to make those skilled in the art fully understand the embodiments. However, it does not exclude that one or more of the specific features, structures, characteristics or other details can not be practiced by those skilled in the art without the technical solution of the present application.

[0022] The flowcharts shown in the drawings are only exemplary illustrations, and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.

[0023] The block diagrams shown in the drawings are only functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0024] It should be understood that although the first, second, third, etc. denotative adjectives can be used herein to describe various devices, elements, components or parts, this should not be limited by these adjectives. These adjectives are used to distinguish one from another. For example, the first device can also be referred to as the second device without departing from the essential technical solution of the present application.

[0025] The term "and / or" or "and / or" includes all combinations of any one and one or more of the associated listed items.

[0026] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0027] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0028] In view of the above problems, the present application proposes a data management optimization method and system, which gathers the security logs and data of each node, performs efficient cleaning and analysis, establishes a unified risk monitoring, alarm issuing and event handling mechanism, and realizes data aggregation processing under enterprise-level multi-layer architecture. Specifically, based on the collection strategy, the logs and business data of each sub-node are collected, the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation, which can effectively improve the data collection rate of the sub-node; the collected logs and business data of each sub-node are transmitted to the central node and subjected to multi-type data processing, the quantitative indicators of data processing delay status are calculated, which can effectively optimize the data processing process; the data after multi-type data processing is monitored in real time, the risk event recognition model is recognized by using the self-constructed risk event recognition model, the risk score of the risk event is calculated, the monitoring rules and feature library of the risk event recognition model are updated by analyzing and calculating the risk monitoring results; when the calculated risk score exceeds the preset threshold, the alarm mechanism is triggered, and the alarm information is pushed to the security management personnel of the related sub-node, thereby effectively realizing synchronous risk event recognition and real-time alarm while optimizing the data collection process and data processing process of multiple sub-nodes.

[0029] It should be noted that in the data management optimization method proposed by the present application, an organized or multi-node business data management mode is adopted, specifically taking the central node as the core and cooperating with multiple sub-nodes to operate together, and the security logs and data generated by each sub-node are comprehensively collected and accurately aggregated to the central node. Based on the improved risk event recognition model, the central node accurately identifies potential or emerging risk events and timely issues alarm information so that relevant personnel can quickly carry out disposal work, thereby effectively ensuring data security.

[0030] Embodiment 1

[0031] The content of the present application will be described in detail below with reference to Figure 1 , Figure 2 and Figure 3 .

[0032] Figure 1 is a step flow chart of an example of the data management optimization method of the present application.

[0033] As shown in Figure 1 , the data management optimization method of the present application includes the following steps.

[0034] In step S101, based on the collection strategy, the logs and business data of each sub-node are collected, and the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation.

[0035] Figure 2 is a framework schematic diagram of an application example of the data management optimization method of the present application.

[0036] From Figure 2 application examples, the application includes a central node, a plurality of sub-nodes communicable (transmissible data) with the central node, wherein the central node serves as the core control center of the entire system (such as a data management system or a data management platform), has strong computing and storage capabilities, adopts a high-performance server cluster architecture, and can be flexibly expanded according to the data volume and business needs of various enterprises. After deploying multiple high-performance servers in the system to form a central node cluster at a suitable location, installing and configuring big data processing frameworks such as Hadoop, Spark, and related software for data management, analysis, and monitoring, and performing network configuration and security settings on the backend servers, centralized management of each sub-node is achieved, responsible for receiving, storing, and processing security logs and business data from the sub-nodes, ensuring centralized management and efficient processing of data, and ensuring stable operation and data security of the central node. Multiple sub-nodes are distributed in different business systems or network areas of the enterprise, and are deployed with lightweight and adaptive collection agents that can automatically adjust the collection frequency based on the running state of the business system, automatically reduce the collection frequency during peak business periods to reduce the impact on the performance of the business system, and appropriately increase the collection frequency during the trough to ensure timely data collection. Each sub-node is responsible for collecting security logs and data from its own node and uploading them to the central node through a secure communication protocol such as SSL / TLS. In actual deployment, the sub-node collection program is deployed in each business system or network area through software installation packages or containerization technologies such as Docker, and parameters such as collection frequency, data source address, and collection content are configured, and network connection testing is performed to ensure that the sub-nodes can communicate normally with the central node to ensure the safe transmission of collected security logs and data.

[0037] As Figure 3 shown, a multi-layer distributed architecture design is adopted to achieve efficient monitoring and operation processing of multi-node or organized data security. Specifically, from bottom to top, it is divided into data collection layer, data aggregation layer, data analysis layer, risk monitoring layer, alarm issuing layer and event handling layer (for details, see Figure 3 ), each layer is independent and works together to jointly ensure enterprise data security, and thus forms a data management system with good scalability and maintainability, facilitating subsequent function upgrade and optimization.

[0038] In a specific embodiment, in the data collection layer, the security logs and business data of each sub-node are collected.

[0039] Specifically, based on the collection strategy, the logs and business data of each sub-node are collected, and the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation.

[0040] The collection strategy includes configuring collection parameters according to actual needs, calculating evaluation parameters, and judging whether the evaluation parameters meet preset values, wherein the evaluation parameters include a data collection success rate and a data collection delay evaluation value. For the collection parameters, for example: setting the type of collected data (such as system logs, application logs, security audit logs, etc.), the collection frequency (such as every minute, every hour, etc.), the limit of the amount of collected data, etc.

[0041] It should be noted that a light, high-compatibility customized data collection agent is deployed in each type of system in each sub-node, and the types of collected data include network traffic data, system logs, database access logs, application logs, etc.

[0042] Specifically, the customized data collection agent (hereinafter also referred to as a collection agent) is used to collect business data, and the collected business data includes log information generated by various systems such as hosts, VPNs, bastion hosts, databases, interfaces, and asset information, alarm records, security event details, and other key security data.

[0043] Specifically, the host log includes information such as the running state and operation behavior of the host, such as user login, file access, process startup, etc., which helps to discover risk activities and potential security threats at the host level. The VPN log includes the use of the virtual private network, including connection time, user identity, accessed internal resources, etc., which can monitor the security status of remote access. The bastion host log focuses on the audit of operation and maintenance operations, including operation instructions, operation time, and operation objects of operation and maintenance personnel, to ensure the traceability and compliance of operation and maintenance operations. The database log contains operation records such as query, modification, and deletion of the database.

[0044] The interface log includes the calling situation of the interface between systems, including calling time, calling parameters, and return results, which helps to analyze the security and stability of the interface.

[0045] The security log includes asset information such as the model, configuration, and location of the device.

[0046] The collection strategy includes configuring collection parameters according to actual needs, calculating evaluation parameters, and judging whether the evaluation parameters meet preset values, wherein the evaluation parameters include a data collection success rate and a data collection delay evaluation value.

[0047] The following expression is used to calculate the data collection success rate, which is an indicator for measuring the integrity and reliability of the collection agent in obtaining security logs and business data from each sub-node:

[0048]

[0049] C 成功represents the success rate of data collection, which is obtained by summing up the success rates of data collection from all the sub-nodes and then calculating the average; L i represents the actual amount of data collected at the i-th sub-node, Z i represents the data integrity coefficient of the data collected at the i-th sub-node; E i represents the data reliability coefficient of the data collected at the i-th sub-node, n is the total number of sub-nodes, i represents the i-th sub-node, i and n are positive integers, specifically 1, 2,..., n; a represents the first weight corresponding to the actual amount of data collected at the i-th sub-node; b represents the second weight corresponding to the data integrity coefficient of the data collected at the i-th sub-node; g represents the third weight corresponding to the data reliability coefficient of the data collected at the i-th sub-node; G i represents the actual amount of data collected at the i-th sub-node, Z

[0050] Specifically, the data integrity coefficient is used to measure whether the collected data is complete, which can be determined by calculating the proportion of missing fields or records in the collected data. For example, if a piece of data should contain 10 fields, but only 8 fields are actually collected, the integrity coefficient is 0.8.

[0051] The data reliability coefficient is used to evaluate the reliability of the collected data, which can be determined by data verification, error detection, etc. For example, if the collected data is verified and found to have a 5% error rate, the reliability coefficient is 0.95.

[0052] The first weight a, the second weight b, and the third weight g satisfy a+b+g=1. These weights can be adjusted according to actual needs to reflect the importance of different factors in the success rate of data collection.

[0053] Alternatively, the value of a is 0.2-0.6, the value of b is 0.2-0.8, and the value of g is 0.2-0.7.

[0054] For the data integrity coefficient, the field and record level verification are combined, and the dynamic adjustment factor reflects the historical trend. The data integrity coefficient is calculated using the following expression:

[0055]

[0056] wherein, Z i represents the data integrity coefficient of the data collected at the i-th sub-node;

[0057] z jdenotes the missing field weight of the jth missing field according to the field importance distribution, m denotes the number of missing fields, j and m are positive integers, j is specifically 1, 2, …, m, for example, the missing field weight corresponding to the user ID is 0.8, and the missing field weight corresponding to the timestamp is 0.2; p j denotes the field missing ratio, i.e., the missing rate, for example, if a certain field is missing 10%, the missing rate j = 0.1; z k denotes the penalty coefficient set according to the kth log record type, p is the number of log types, k and p are positive integers, k is specifically 1, 2, …, p, z k the value of z is 0.5-1.5, for example, the penalty coefficient of the key log record type is 1.5, and the penalty coefficient corresponding to the ordinary log record type is 0.5; p k denotes the missing number of the kth record; z 动态调整因子 denotes the parameter dynamically adjusted according to the historical data integrity trend, z 动态调整因子 the initial value of z is 1, for example, if the integrity is lower than 80% for three consecutive times, z 动态调整因子 is reduced to 0.9.

[0058] The data reliability coefficient is introduced into time decay and redundancy check to evaluate the data accuracy, and the following expression is used to calculate the data reliability coefficient:

[0059]

[0060] wherein, E i denotes the data reliability coefficient of the data successfully collected at the ith sub-node, i is a positive integer; p a denotes the pass rate, which is specifically calculated by performing hash check, format check, and logic check on the ath segment of the data successfully collected at the ith sub-node, for example, the hash check pass rate is 90%, and the format check pass rate is 100%; E a denotes the field missing ratio, i.e., the missing rate, of the ath field, q denotes the number of missing fields, a and q are positive integers, and a is specifically 1, 2, …, q; denotes the time decay representation value using an exponential decay function, wherein λ is the decay coefficient (for example, λ = 0.1), t 时间延迟 denotes the data collection delay time of the business data collected at the ith sub-node, the unit is second (in other embodiments, it can also be minute and hour); E 冗余校验因子 denotes the check result adjustment factor according to the redundancy data source. For example, there are three redundancy data sources, and two redundancy data sources are consistent, and the check result adjustment factor is 0.95.

[0061] The first weight, the second weight and the third weight are calculated using the following expression based on the sensitivity of business requirements and dynamic allocation of historical data, which adapts to different scenarios.

[0062]

[0063] wherein a represents the first weight corresponding to the actual amount of data successfully collected at the i-th sub-node; β represents the second weight corresponding to the data integrity coefficient of the data successfully collected at the i-th sub-node; γ represents the third weight corresponding to the data reliability coefficient of the data successfully collected at the i-th sub-node; y 整 represents the business requirement integrity sensitivity determined according to the business scenario, L 整 represents the average value of the integrity of the specified historical time period (including the historical time period of seven days, ten days, fifteen days, thirty days, sixty days from the current time to the past), i.e. the historical integrity average, y 靠 represents the business requirement reliability sensitivity determined according to the business scenario, L 靠 represents the average value of the reliability of the specified historical time period (including the historical time period of seven days, ten days, fifteen days, thirty days, sixty days from the current time to the past), i.e. the historical reliability average, y 数 represents the business requirement data volume sensitivity determined according to the business scenario, L 数 represents the average value of the data volume of the specified historical time period (including the historical time period of seven days, ten days, fifteen days, thirty days, sixty days from the current time to the past), i.e. the historical data volume average.

[0064] For example, in the database log audit scenario, the integrity sensitivity is 0.7, the reliability is 0.2, and the data volume is 0.1. For example, the weights are recalculated every week according to the changes in business requirements to dynamically adjust the weights.

[0065] For judging whether the evaluation parameters meet the preset values, specifically including whether the calculated data collection success rate is greater than the preset value (e.g. 99.5%). In order to make the calculated data collection success rate greater than the preset value (e.g. 99.5%), the collection agent sets an automatic retransmission mechanism, which automatically resends the data within a set time interval (such as 30 seconds) when the data transmission fails, with a maximum of 3 retries. At the same time, a data collection state monitoring mechanism is established to monitor the collection state of each sub-node in real time, and when it is detected that the collection success rate is lower than the preset value, an alarm is triggered immediately and corresponding measures are taken, such as checking the network connection, optimizing the collection configuration, etc.

[0066] The following expression is used to calculate the data collection delay evaluation value, which is used to evaluate the response speed of the collection agent to the data and reflects the real-time performance of data collection:

[0067]

[0068] wherein, C 延迟 represents the response speed generated by the collection agent when collecting data at the individual node; C 基础 represents the basic delay determined by the data generation time, collection completion time, and transmission time consumption of each sub-node; B 惩罚 represents the fluctuation penalty parameter characterized by the historical delay mean or historical delay standard deviation in a specific historical time period; F 事件 represents the event delay parameter involved in the risk event, which is specifically determined according to the number of risk events, the number of sub-nodes affected by each risk event, the occurrence event of the risk event, and the business tolerance duration; min 延迟 represents the theoretical minimum delay parameter in an ideal state (such as close to 0 when collecting locally, and the network transmission theoretical minimum value when collecting across the network); D 补偿 represents dynamic benchmark compensation.

[0069] The fluctuation penalty parameter is calculated using the following expression:

[0070]

[0071] wherein, B 惩罚 represents the fluctuation penalty parameter characterized by the historical delay mean or historical delay standard deviation in a specific historical time period (for example, a historical time period of seven days, ten days, fourteen days, thirty days, etc. from the current time to the past); d 延迟 represents the current delay; c 差 represents the historical delay mean / standard deviation, which is based on the delay statistics of the past N times of collection; b 波 represents the fluctuation sensitivity coefficient, which is set according to the business requirement for delay stability. For example, in a high real-time scenario, the fluctuation sensitivity coefficient b 波 is 1.5, and in a general scenario, the fluctuation sensitivity coefficient b 波 is 0.8.

[0072] The abnormal event delay is calculated using the following expression:

[0073]

[0074] wherein, F 事件 represents the event delay parameter involved in the risk event, which is specifically determined according to the number of risk events, the number of sub-nodes affected by each risk event, the occurrence event of the risk event, and the business tolerance duration; q e represents the weight corresponding to the e-th risk type, which is calculated and allocated according to the risk severity, r' represents the number of risk types, e and r' are both positive integers, e is specifically 1, 2,..., r', and he represents the duration of the risk event of the e-th risk type, i.e. the length of time that the risk event affects the collection. For q e The calculation process is as follows:

[0075] q e = (f 影响范围 × t 持续时间 × L 频率惩罚 ) × (1 + Y G ) × Y 动态

[0076] wherein q e represents the weight corresponding to the e-th risk type, which is calculated and distributed according to the risk severity;

[0077] f 影响范围 represents the number of affected range systems, wherein Y 受影响 represents the number of affected nodes, i.e. the number of collection nodes or subsystems affected by the risk event; Y 总 represents the total number of nodes, i.e. the total number of collection nodes; Y 重要性 represents the node importance factor of each node, which is distributed according to the node business value, for example, the node importance factor of the node corresponding to the core database is 1.5, and the node importance factor of the node corresponding to the ordinary log sending is 0.8; t 持续时间 represents the duration coefficient, wherein h 持续 represents the duration of the risk event, i.e. the length of time from the occurrence to the recovery of the risk event; h 容忍 represents the business tolerance duration, i.e. the maximum duration of the risk event that the business can accept (for example, the network interruption tolerance duration is 30 seconds); h 敏感 represents the time sensitivity index, which is adjusted according to the sensitivity of the business to delay (for example, the time sensitivity index corresponding to the real-time transaction system is 1.8, and the time sensitivity index corresponding to the offline analysis system is 1.2); L 频率惩罚 represents the historical frequency penalty, wherein L 发生 represents the historical occurrence frequency, i.e. the number of occurrences of the risk type in a specific historical time period (for example, the historical time period of seven days, ten days, fourteen days, thirty days, etc. from the current time to the past); L 基准 represents the historical reference frequency, i.e. the expected risk occurrence frequency (for example, once a week); L 敏感 represents the frequency sensitivity coefficient, i.e. the punishment degree of the risk (for example, the frequency sensitivity coefficient corresponding to the high-risk risk event is 1.5); Y G represents the business sensitivity addition, represents the bth business rule weight, which is assigned according to the importance of the business rule (for example, the business rule weight corresponding to the compliance rule is 1.2, and the business rule weight corresponding to the performance rule is 0.9); represents the bth rule matching degree, which is the matching degree of the risk event and the business rule (for example, 1.0 for complete matching and 0.5 for partial matching), n represents the number of business rules, b and n are positive integers, and b is specifically 1, 2,..., n; 动态 represents a dynamic adjustment factor, Y 动态 = 1 + (Y 实环 × W 外部 ), wherein Y 实环 represents a real-time environment factor, which is dynamically adjusted according to the current system load, network state, and the like (for example, the value of Y 实环 is 0.2 when the load is high, and the value of Y 实环 is 0 when the load is low); W 外部 represents an external influence coefficient, which considers the influence of external factors (network attack warning) on the risk weight (for example, the value of W 外部 is 0.3 when there is a network attack warning).

[0078] The following expression is used to calculate the dynamic benchmark compensation:

[0079] D 补偿 = y 延迟 × r 容忍

[0080] wherein D 补偿 represents the dynamic benchmark compensation; y 延迟 represents a historical delay quantile, which is a delay quantile in a specific historical time period (for example, a historical time period of seven days, ten days, fourteen days, thirty days, or the like calculated from the current time to the past), and is used to reflect an acceptable upper limit of delay, for example, the 90th quantile of delay in the past seven days; r 容忍 represents a business tolerance coefficient, which is adjusted according to business needs (for example, set to 1.2 when delay fluctuations are tolerated, and set to 1.0 when strict real-time performance is required).

[0081] The following expression is used to calculate the real-time weight factor:

[0082]

[0083] wherein S 实时 represents the real-time weight factor; D 当前 represents the current delay; Y 最大允许 represents the business SLA delay, which is the maximum allowed delay specified in the service level agreement; s 优先represents a data priority coefficient, which is assigned according to the importance of the data (for example, 1.2 for critical logs and 0.8 for ordinary logs); a1represents a first weight factor; b1represents a second weight factor, a1+b1=1, the value range of a1is 0.2-0.8, and the value range of b1is 0.3-0.8. According to the business focus adjustment, when the SLA compliance is focused, a1=0.7, and when the data priority is focused, b1=0.7.

[0084] For judging whether the evaluation parameter meets the preset value, specifically, the average data collection delay evaluation value of each subnode is less than a predetermined value (for example, 5 seconds, which is dynamically adjusted according to different business data types). For example, for critical business data, the data collection delay evaluation value is less than 2 seconds. In order to make the calculated data collection delay evaluation value less than the preset value, the collection agent adopts a customized collection strategy, which includes an event-triggered collection mode, which collects immediately when critical business data is generated, reducing unnecessary polling. The event trigger specifically refers to the generation of critical business data, which includes client-side business data related to the purchase operation of domestic traffic, international traffic, and targeted traffic. The critical business data includes operation management-side business data related to the purchase operation of domestic traffic, international traffic, and targeted traffic.

[0085] Then, the related parameters of data collection are real-timely fed back to the background server for statistics and calculation, and two quantitative indexes, data collection success rate and data collection delay evaluation value, are output to ensure accurate judgment of the quality of the collected data.

[0086] By using a dynamic collection strategy, the security logs and business data of each subnode are collected, and the data collection success rate and the data collection delay evaluation value are calculated, which can effectively ensure the accuracy of data collection and the quality of data.

[0087] In a specific embodiment, the data collection task is to collect database log data from 20 subnodes. The risk event is that network interruption causes five core nodes (node importance factor = 1.2) to be unable to transmit data for 60 seconds (business tolerance time = 30 seconds). The historical data is that the risk event has occurred three times in the past seven days (benchmark frequency = 1 time / week). The business rules include compliance rules (corresponding weight = 1.2, complete match) and performance rules (corresponding weight 0.9, partial match). The system state is that the current load is moderate (real-time environment factor is 0.1), and there is no external warning (external influence coefficient is 0).

[0088] Specifically, the step of calculating the weight corresponding to the risk type is performed, and the weight corresponding to the risk type = (0.3 x 3.48 x 6.73) x (1 + 1.65) x 1.0 ≈ 19.3.

[0089] Then, the data collection delay evaluation value is calculated. For example, the data generation time is 10:00:00.000, the collection completion time is 10:00:00.500, the transmission time consumption is 0 ms, thus the basic delay is 500 ms, the historical delay average is 300 ms, the standard deviation is 100 ms, the fluctuation sensitivity coefficient is 1.2, thus, The risk duration is 60 ms (network interruption causes 60 ms delay), the weight corresponding to the risk type is 19.3, thus the risk event delay = 19.3 x 60 = 1158 ms, the theoretical minimum delay is 100 ms, the 90th percentile of the delay in the past seven days is 600 ms, the business tolerance coefficient is 1.1, thus the dynamic benchmark compensation is 660 ms, the business SLA delay is 800 ms, the data priority coefficient is 1.0, a = 0.6, b = 0.4, thus, That is, 0.625 is obtained. In summary, the data collection delay evaluation value is 136.6%.

[0090] As can be seen from the above example, the weight corresponding to the risk type is 19.3, indicating that the risk event has a greater impact on the system and needs to be processed in priority. The data collection delay evaluation value is 136.6%, which significantly exceeds the business tolerance upper limit (800 ms) and is mainly affected by the risk event (risk delay contribution 1158 ms). For example, by immediately investigating the reason for the network interruption and optimizing the fault tolerance capability of the collection agent, such as a built-in cache (Redis) in the agent, temporarily storing data during network interruption and automatically supplementing transmission after recovery.

[0091] It should be noted that the above is only described as an optional example and cannot be understood as a limitation of the present application.

[0092] Next, in step S102, the collected log and business data of each sub-node are transmitted to the central node, and multi-type data processing is performed to calculate a quantitative index for determining the data processing delay condition to optimize the data processing process.

[0093] Specifically, each sub-node transmits the collected security log and business data to the data aggregation server of the central node through, for example, a secure channel (application of incremental synchronization, breakpoint resume, data encryption, etc., to ensure data real-time, confidentiality and integrity). In the data transmission process, the security log and business data are compressed to reduce the transmission amount, reduce the network bandwidth occupation, and thus improve the data transmission efficiency.

[0094] It should be noted that the secure channel refers to a network path that ensures data is not stolen, tampered with or forged during transmission through encryption, authentication and integrity protection mechanisms. Its core goal is to achieve data confidentiality, integrity and availability, prevent man-in-the-middle attacks, data leakage or malicious tampering. During transmission, the data to be transmitted is compressed to reduce transmission volume, reduce network bandwidth occupancy and improve data transmission efficiency.

[0095] Optionally, the security log and business data are encrypted using, for example, AES-256 and RSA / ECC encryption algorithms, and both the child node and the central node need to verify the identity of the other party (such as TLS certificate verification) to prevent the security log and business data from being stolen or tampered with during data transmission.

[0096] For example, when the data aggregation server receives the security log and business data, it performs preliminary preprocessing operations on the security log and business data. Preprocessing operations include data format conversion, which converts different formats of data sent by different child nodes to a standard format, and integrity checking to ensure that the received security log and business data are not lost or damaged.

[0097] Multi-class data processing includes data cleaning, data format conversion, etc.

[0098] Advanced data cleaning algorithms are used to deeply clean the aggregated security log and business data. First, based on data characteristics and business rules, noise data in the data is identified and removed, such as records containing invalid information such as garbled code and incorrect timestamps. Second, duplicate data is detected and removed, ensuring data uniqueness through hash algorithms or similarity comparison methods. In addition, data is standardized to unify data formats and units, such as converting different date and time formats to standard timestamp format.

[0099] In addition, the parsed unstructured or semi-structured log data is converted into structured data by, for example, a parsing engine. The parsing engine supports parsing of multiple log formats, such as Apache log format, Windows event log format, database log format, etc. For each log format, according to the predetermined parsing rules, key fields such as time, source IP, operation type, target object, etc. are extracted and stored in structured database tables for subsequent query and analysis.

[0100] Since data processing speed plays a crucial role in system efficiency, relevant parameters of data processing are continuously transmitted back to the background server for statistics and calculation, which can accurately output quantitative indicators of data processing delay and accurately assess system efficiency.

[0101] The data processing delay index is calculated by using the following expression to evaluate the processing efficiency of the data analysis layer on the data, so as to optimize the data processing process:

[0102]

[0103] Y represents the data processing delay index of the current data of the current child node in the data processing layer, which is used to measure the real-time performance of the current data processing;h 基础 represents the basic time consumption for different data processing, specifically including query processing; 动态 represents the dynamic delay penalty generated by dynamic delay processing, specifically including additional time consumption caused by resource competition, data skew, and system load; 理论 represents the theoretical optimal processing time determined according to the optimal time consumption in the specified historical time period and hardware factors, and the specified historical time period includes historical time periods of seven days, ten days, fifteen days, thirty days, and sixty days calculated from the current time to the past; 实时 represents the real-time performance impact factor determined according to the data collection priority and data invalidity factor of the business data, Q b represents the weight of the bth business rule, b the bth rule matching degree, b and n are positive integers, n represents the number of business rules, b is 1, 2,..., n, and the business rules include real-time transaction rules, offline report rules, and data operation rules of child nodes.

[0104] Specifically, the theoretical optimal time consumption h 理论 , i.e., the theoretical optimal processing event, is determined according to the optimal time consumption in the specified historical time period and hardware factors (such as hardware upgrade compensation coefficient), and the specified historical time period includes historical time periods of seven days, ten days, fifteen days, thirty days, and sixty days calculated from the current time to the past.

[0105] For example, the theoretical optimal time consumption is represented by using the following expression:

[0106] h 理论 = h 最优耗时 × (1-h 补偿系数 ).

[0107] wherein h 理论 represents the theoretical optimal processing time determined according to the optimal time consumption in the specified historical time period and hardware factors;h 最优耗时 represents the optimal time consumption in the specified historical time period, and the specified historical time period includes historical time periods of seven days, ten days, fifteen days, thirty days, and sixty days calculated from the current time to the past; 补偿系数 represents the hardware upgrade compensation coefficient.

[0108] Specifically, when the calculated processing delay index is greater than a specified value (for example, 120%), resource isolation is performed on the resources required for processing the current data, or dynamic expansion is performed to improve the processing speed of the current data.

[0109] The dynamic delay penalty is calculated using the following expression:

[0110] d 动态 = Z 资源 x S 倾斜 x f 负载

[0111] wherein d 动态 represents a dynamic delay, specifically including additional time consumption caused by resource competition, data skew, and system load;

[0112] Z 资源 represents a resource competition penalty, CPU 当前使用率 represents the current usage rate of the CPU, I 当前使用率 represents the current usage rate of the memory, z 系数 represents a competition sensitivity coefficient, which can be determined according to the corresponding sub-node priority and the corresponding data priority; S 倾斜 represents a data skew penalty, q 倾斜 represents a data skew sensitivity index, which is determined according to the relationship between the data amount of the current partition and the average data amount of the partition, q 平均 represents the average data amount of all partitions; f 负载 represents a system load factor of the data management optimization system, r 当前 represents the current concurrent task amount, r 基准 represents the reference concurrent amount, f 指数 represents a load sensitivity index.

[0113] In an embodiment, the weight corresponding to the real-time transaction rule is 1.5, and the weight corresponding to the offline report rule is 0.8.

[0114] For the case of hardware upgrade (for example, CPU performance improvement of 20%), the hardware upgrade compensation coefficient is 0.2, which is used to determine the theoretical optimal processing time.

[0115] In a scenario example, when processing the collected database log data, for example, the base processing time is 200 ms, the CPU occupancy is 90%, the memory occupancy is 80%, the competition sensitivity coefficient is 1.2, the maximum partition data volume is 4 times the average partition data volume, the data skew sensitivity index is 1.5, the current concurrent task volume is 3 times the benchmark concurrent volume, the load sensitivity coefficient is 1.0, the weight corresponding to the matched real-time transaction rule is 1.5, the matching degree is 1.0, the historical optimal time consumption (i.e., the optimal time consumption in the historical time period) is 150 ms, and the system hardware is not upgraded, then the calculated data processing delay index is 217.46%.

[0116] As can be seen from the above scenario example, the data processing delay index is 217.46%, which significantly exceeds the specified value (for example, 120%), indicating that the processing efficiency is far lower than the theoretical optimum, and the resource competition and data skew problems need to be optimized first. For example, by resource isolation and dynamic expansion, the delay is quickly relieved, and by data re-partitioning and skew key processing, the problem is radically solved, and the delay is reduced to the theoretical optimum value (for example, 100%).

[0117] It should be noted that in some embodiments, the average data processing delay does not exceed 10 seconds. For critical business data (such as business data related to purchase operations of targeted traffic), the processing delay does not exceed 5 seconds. A distributed computing framework such as Apache Spark is used to perform parallel processing of data to improve data processing speed. At the same time, the data processing algorithm and process are optimized to reduce unnecessary calculations and storage operations. The above is only an optional example for illustration and cannot be understood as a limitation of the present application.

[0118] Next, in step S103, the data after multi-type data processing is monitored in real time, and a risk event recognition model is used to recognize risk events and calculate the risk score of the risk events, wherein the risk monitoring results are analyzed and calculated to update the monitoring rules and feature library of the risk event recognition model.

[0119] Specifically, the data collected from each sub-node is subjected to multi-type data processing, which specifically includes data cleaning, compression, encryption, decryption, and the like, to form log and event records after multi-type data processing.

[0120] A risk event recognition model is used to recognize risk events.

[0121] Optionally, external threat intelligence (such as CVE vulnerability library, IP blacklist) is used to identify risk events.

[0122] Specifically, structured data includes database logs, security event records (fields such as timestamp, IP, event type). Unstructured data includes text logs (such as web server logs), free text descriptions (such as event handling notes).

[0123] It should be noted that data cleaning includes data deduplication, specifically removing duplicate records (such as duplicate login failure logs). Fill in missing values (such as filling in numerical fields with mean values, filling in categorical fields with "unknown"). Standardize formats (such as uniform timestamp format ISO 8601).

[0124] For the construction of the risk event identification model, the following algorithms are used to construct one or more risk event identification models: random forest algorithm, XGBoost algorithm, neural network algorithm. Or use ensemble learning, use the first risk event identification model and the second risk event identification model established by random forest and XGBoost algorithm respectively, and weighted voting of the prediction results of the first risk event identification model and the second risk event identification model (such as the weight of the first risk event identification model is 0.6 and the weight of the second risk event identification is 0.4).

[0125] For the risk event identification model including the training process, the following data of each sub-node collected is labeled with risk type to obtain a training data set for training the risk event identification model: sub-node log data after various data processing, event record process; risk types include a first risk class related to vulnerabilities, a second risk class related to login failures, a third risk class related to IP traffic, and a fourth risk class related to cracking and data leakage. Non-risk type is normal user login.

[0126] For the definition of risk type label, for example, "10 failed logins within 5 minutes as a risk event".

[0127] In addition, it also includes statistical features, behavioral features, and time series features. Among them, the statistical features include login frequency (the number of user logins per hour), request distribution (standard deviation of API interface request volume to reflect volatility). Behavioral features include access path entropy (user access URL diversity, where high entropy indicates scanning behavior), permission usage (proportion of users accessing sensitive data, such as financial personnel accessing database proportion). Time series features include sliding window statistics (such as the number of risk events in the past 1 hour) and trend analysis (including the change rate of login failure rate, such as a sudden increase in the change rate).

[0128] Further feature transformation such as normalization, encoding, dimensionality reduction, etc. For example, scale numerical features to [0, 1] (e.g., number of logins divided by the maximum value). For example, convert categorical features to numerical (e.g., One-Hot encode the country of origin of an IP address). For example, reduce high-dimensional features using PCA (e.g., TF-IDF vectors of log text).

[0129] In one embodiment, the collected logs of the current child node and the real-time parsed data of the central node are input into the risk monitoring layer, relevant features (specifically including at least two of statistical features, behavioral features, and time series features) are calculated, and the calculated relevant features are input into the risk event identification model to output a risk score and alarm information.

[0130] For example, the child node uploads logs, and the central node performs real-time parsing. A sliding window calculates statistical features (e.g., a login failure rate in the past 5 minutes).

[0131] The collected logs of the current child node and the real-time parsed data of the central node are input into the risk monitoring layer to output a risk score (0-1). In combination with a rule engine, a final alarm is generated (e.g., "risk score > 0.8 and matching SQL injection rule = high-risk alarm").

[0132] In another embodiment, the security logs and business data processed by the multi-class data are specifically monitored in real time.

[0133] In the central node, a risk event identification model based on a combination of machine learning algorithms and rule engines is constructed to monitor and analyze the security logs and business data processed by the multi-class data in real time.

[0134] A combination of supervised learning and unsupervised learning is adopted. Using the above training data set, a risk event identification model is established to identify potential risks or risk events.

[0135] An unsupervised learning algorithm is used to discover risk patterns and potential risk events in the data. For example, a clustering algorithm can detect a risk cluster of data, and an outlier detection algorithm can find data points that differ greatly from the normal data distribution. A rule engine is based on expert experience and industry standards to predefine a series of security rules, such as abnormal access to specific ports, illegal operations on sensitive data, and frequent login failures. When the data meets the rule conditions or the machine learning algorithm detects abnormal patterns, it is determined to be a risk event.

[0136] For example, web server logs are analyzed to determine user access records (timestamp, IP, URL, HTTP status code), K-Means clustering algorithm is used to group user access behaviors, and risk events are identified. Here, the risk event refers to a risk cluster.

[0137] Specifically comprising the following steps.

[0138] Feature extraction is performed on the web server log to obtain an access behavior representation vector (such as access frequency, access time distribution, and access URL diversity) of each user.

[0139] A K-Means clustering algorithm is used to perform clustering calculation on the access behavior representation vector of the user to cluster normal user behaviors into different clusters such as "office time access" and "night low-frequency access".

[0140] For example, cluster A is that a certain IP accesses sensitive paths (such as / admin and / backup) at 1:00-3:00 in the morning at a high frequency (automated scanning). Cluster B is that multiple IPs collectively access an undisclosed API interface in a short period of time (the first risk category related to vulnerabilities).

[0141] Preferably, a risk score mechanism is set in the risk monitoring layer to calculate a risk score for each to-be-processed risk event to quantify each risk event. A risk false alarm rate of the to-be-processed risk event is calculated to optimize the parameters of the risk event identification model.

[0142] The risk score mechanism includes three input indicators: severity, impact range, and occurrence frequency, which can be quantified according to actual business conditions.

[0143] According to the severity, impact range, and occurrence frequency of the risk event, the following expression is used to calculate the risk score of the to-be-processed risk event:

[0144] F 评分 =(W1×B0+W2×F0+W3×K0)×Y0

[0145] Wherein, F 评分 represents the risk score quantified according to the severity, impact range, and occurrence frequency of the current risk event; B0 represents the severity of the current risk event; F0 represents the impact range of the current risk event; K0 represents the occurrence frequency of the current risk event; W1, W2, and W3 are the first weight coefficient, the second weight coefficient, and the third weight coefficient corresponding to the severity, the impact range, and the occurrence frequency, respectively; Y0 represents a business priority coefficient, which is adjusted according to business strategy. For example, the business priority coefficient corresponding to the core business is 1.2, and the business priority coefficient corresponding to the edge business is 0.8.

[0146] The following expression is used to calculate the severity of the current risk event:

[0147]

[0148] Wherein, B0 represents the severity of the current risk event; B1 represents the direct economic loss caused by the current risk event; B2 represents the long-term income loss caused by the current risk event; B3 represents the regulatory penalty caused by the current risk event; B4 represents the total annual income; B5 represents the reputation loss coefficient caused by the current risk event. For details, see Table 1 below.

[0149] Table 1

[0150]

[0151] Table 1 is an example table of parameters for calculating the severity of the current risk event.

[0152] The following expression is used to calculate the impact range of the current risk event:

[0153]

[0154] Wherein, F0 represents the impact range of the current risk event; F1 represents the number of affected businesses affected by the current risk event; F2 represents the total number of businesses; F3 represents the number of affected users affected by the current risk event; F4 represents the total number of users; F5 represents the user sensitivity index of users to the current risk event. For details, see Table 2 below.

[0155] Table 2

[0156]

[0157] Table 2 is an example table of parameters for calculating the impact range of the current risk event.

[0158] The following expression is used to calculate the occurrence frequency of the current risk event:

[0159]

[0160] Wherein, K0 represents the occurrence frequency of the current risk event; K1 represents the number of historical risk events; K2 represents the observation period; K3 represents the trend growth rate of the number of current risk events; K4 represents the time sensitivity coefficient. For details, see Table 3 below. Table 3 is an example table of parameters for calculating the occurrence frequency of the current risk event.

[0161] Table 3

[0162]

[0163]

[0164] For the first weight coefficient, the second weight coefficient, and the third weight coefficient, the weight is recalculated according to the latest data every quarter. When the ratio of the historical impact event and the historical high-frequency event exceeds a threshold value (such as 30%), the weight update is triggered immediately.

[0165] First, the severity weight, the impact range weight, and the occurrence frequency weight are calculated respectively, and then the first weight coefficient, the second weight coefficient, and the third weight coefficient are obtained after normalization.

[0166] The severity weight is calculated as follows:

[0167]

[0168] Wherein, W'1 represents the severity weight; b1 represents the industry risk preference coefficient; b2 represents the business loss sensitivity index; b3 represents the industry average severity; and b4 represents the historical severity event proportion. For details, refer to Table 4 below. Table 4 is an example table of parameters for calculating the severity weight.

[0169] Table 4

[0170]

[0171] The impact range weight is calculated as follows:

[0172]

[0173] Wherein, W'2 represents the impact range weight; f1 represents the user scale sensitivity index; f2 represents the business tolerance to interruption; f3 represents the industry average impact range; and f4 represents the historical impact event proportion. For details, refer to Table 5 below. Table 5 is an example table of parameters for calculating the impact range weight.

[0174] Table 5

[0175]

[0176] The occurrence frequency weight is calculated as follows:

[0177]

[0178] Wherein, W'3 represents the occurrence frequency weight; g1 represents the technology maturity coefficient; g2 represents the risk controllable value; g3 represents the industry average occurrence frequency; and g4 represents the proportion of the number of historical high-frequency events in the total number of risk events. For details, refer to Table 6 below. Table 6 is an example table of parameters for calculating the occurrence frequency weight.

[0179] Table 6

[0180]

[0181]

[0182] Then, the severity weight, the influence range weight and the occurrence frequency weight are normalized to obtain a first weight coefficient, a second weight coefficient and a third weight coefficient:

[0183]

[0184] By calculating the risk score of the to-be-processed risk event according to the severity, the influence range and the occurrence frequency of the risk event, the to-be-processed risk event can be accurately quantified.

[0185] Further, according to the calculated risk score, the risk event is divided into different levels, and corresponding disposal measures are taken.

[0186] The following expression is used to calculate the risk false positive rate for reflecting the misjudgment of the risk event identification model on normal events:

[0187]

[0188] Wherein, F 误报率 represents the risk false positive rate of mistaking the normal event as a risk event; A i represents the number of false positive events mistaking the normal event as the i-th risk event; B j represents the number of normal events correctly identifying the j-th event as a normal event; C i represents the resource consumption coefficient of processing the i-th false positive event; C j represents the resource consumption coefficient of processing the j-th false positive event; D i represents the business sensitivity index of the i-th false positive event; D j represents the business sensitivity index of the j-th false positive event; M represents the total number of false positive events, j and M are positive integers, j is specifically 1, 2,..., M; N represents the total number of normal events, i and N are positive integers, i is specifically 1, 2,..., N. For details, see Table 7 below. Table 7 is an example table of parameters for calculating the risk false positive rate.

[0189] Table 7

[0190]

[0191]

[0192] For example, the threshold of false positive rate is 8%, i.e. the calculated false positive rate should not exceed 8%. By optimizing the threshold settings and rule definitions of the risk monitoring model, the occurrence of false positives is reduced. The false positive events are analyzed, the reasons for false positives are summarized, and the model parameters or rules are adjusted accordingly, including adjusting the threshold, optimizing feature extraction, adjusting the model complexity, refining the rule conditions, adjusting the rule priority, etc. At the same time, a false positive feedback mechanism is established, which allows security management personnel to mark and feedback false positive events so that the model can learn and improve in a timely manner.

[0193] The risk miss rate is calculated using the following expression to measure the omission of the risk event identification model to security threats:

[0194]

[0195] Wherein, F 漏报率 represents the risk miss rate of correctly identifying the missed risk event; L h represents the number of missed events that actually occurred; P h represents the harm weight of the hth missed event to the business or system; R h represents the business loss coefficient of the hth missed event, H represents the total number of missed events, h and H are positive integers, h is specifically 1, 2, …, H; O q represents the actual number of risk events including the qth missed risk event that is not identified; P q represents the harm weight of the qth missed event to the business or system; R q represents the business loss coefficient of the qth missed event; Q represents the total number of actual risk events, q and Q are positive integers, q is specifically 1, 2, …, Q. See Table 8 below for details. Table 8 is an example table of parameters for calculating the risk miss rate.

[0196] Table 8

[0197]

[0198]

[0199] For example, the threshold of risk false negative rate is 5%, that is, the calculated risk false negative rate should not exceed 5%. In order to reduce the false negative rate, the research and analysis of new risk events need to be strengthened, and the rules and feature library of the risk monitoring model need to be updated constantly, including adding new attack mode rules (such as new risk software, APT attack, supply chain attack, etc.), refining rule conditions, introducing dynamic rules (such as threat intelligence, behavior baseline rules, etc.), updating risk feature library, updating model training data, etc. The method of multi-model fusion is adopted, combined with different machine learning algorithms and rule engines, to improve the comprehensiveness and accuracy of risk monitoring. The risk monitoring results are reviewed and analyzed regularly to find possible false negative situations and take timely measures to improve them

[0200] It should be noted that the above is only described as an optional example and cannot be understood as a limitation of the present application.

[0201] Next, in step S104, when the calculated risk score exceeds the preset threshold, the alarm mechanism is triggered, and the alarm information is pushed to the security management personnel of the relevant sub-node.

[0202] According to the risk score calculated (or output) by the risk monitoring layer, when the risk score exceeds the alarm threshold, the alarm mechanism is triggered. The alarm threshold can be flexibly set according to the security needs and risk bearing capacity of the enterprise. For example, the specific risk levels can be determined by referring to Table 9 below. For high-risk events, a lower alarm threshold is set; for low-risk events, a higher alarm threshold is set. The alarm information includes detailed description of the risk event, occurrence time, impact range, risk score, etc. The detailed description includes the type of risk event, possible causes, affected systems or data, etc., so that the security management personnel can quickly understand the overview of the risk event. Table 9 is an example table of specific classification of risk levels.

[0203] Table 9

[0204] Risk level Score range Alarm trigger condition Example of applicable scenario High risk 5-15 points Alarm triggered when risk score ≥ 6 points Key system attack, data leakage, etc. Medium risk 3-4 points Alarm triggered when risk score ≥ 3.5 points Abnormal login, abuse of authority, etc. Low risk 1-2 points Alarm triggered when risk score ≥ 2 points Normal user misoperation, access to non-sensitive resources

[0205] Specifically, it includes various alarm methods such as email, short message, instant messaging tool, etc., which can timely push alarm information to the security management personnel of the relevant sub-node. At the same time, the alarm issuing layer has alarm filtering and aggregation functions to avoid a large number of repeated alarms interfering with the management personnel. For example, for the repeated occurrence of the same risk event in different sub-nodes, aggregation processing can be performed, and only one comprehensive alarm information is sent. For example, the event handling layer provides a unified event handling platform to display all risk events to be handled, including detailed information of the event, current state, etc. By recording the handling process and results of each event, an event handling knowledge base is formed. The knowledge base contains detailed information of the event, handling methods, lessons learned, etc., providing a reference for the handling of similar events in the future. When a new risk event occurs, the knowledge base can be queried to draw on past handling experience to improve the efficiency and quality of event handling. At the same time, the knowledge base is regularly updated and maintained to ensure the accuracy and timeliness of the knowledge. At the same time, a standardized event handling process is developed, including event confirmation, cause analysis, disposition measure development and execution, disposition result verification, etc.

[0206] In the event confirmation link, the management personnel verifies the alarm information to confirm whether it is a real risk event. False positives are recorded and the monitoring rules are adjusted, and real risk events enter the cause analysis link.

[0207] In the cause analysis link, the root cause of the risk event is determined by collecting and analyzing relevant logs and data. In the disposition measure development and execution link, appropriate disposition measures are developed according to the type and severity of the risk event, and after approval, relevant personnel are arranged to execute and keep communication and coordination. In the disposition result verification link, the effectiveness of the disposition measures is evaluated to ensure that the risk event is effectively resolved. If not, the disposition measures are redeveloped, and if resolved, the disposition results are recorded. During the event handling process, according to the severity of the risk event, different types of events are subject to time limits from occurrence to completion of disposition to measure the efficiency of event handling. For example, high-risk events are completed within 18 hours. For example, medium-risk events are completed within 36 hours. For example, low-risk events are completed within 48 hours.

[0208] An event handling tracking and supervision mechanism is established to monitor the progress of the handling of each risk event in real time. When the event handling time approaches the specified limit, the relevant responsible person is reminded to speed up the handling progress.

[0209] It should be noted that the above is only described as a preferred example and cannot be understood as a limitation of the present application.

[0210] In addition, the drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended for limiting purposes. It is readily understood that the processes shown in the drawings do not indicate or limit the chronological order of these processes. In addition, it is readily understood that these processes can be executed synchronously or asynchronously, for example, in a plurality of modules.

[0211] Compared with the prior art, the present application is based on a collection strategy, collects the logs and service data of each sub-node, calculates the data collection success rate and data collection delay evaluation value to measure the collection situation, and can effectively improve the data collection rate of the sub-node; transmits the collected logs and service data of each sub-node to the central node, performs multi-type data processing, calculates and determines the quantitative indicators of the data processing delay situation, and can effectively optimize the data processing process; performs real-time monitoring on the data after multi-type data processing, identifies risk events using a self-constructed risk event identification model, calculates the risk score of the risk events, analyzes and calculates the risk monitoring results to update the monitoring rules and feature library of the risk event identification model; when the calculated risk score exceeds a preset threshold, triggers an alarm mechanism, and pushes alarm information to the security management personnel of the relevant sub-node, thereby effectively realizing synchronous risk event identification and real-time alarm while optimizing the data collection process and data processing process of multiple sub-nodes.

[0212] In addition, the data processing delay indicators are calculated to evaluate the processing efficiency of the data analysis layer on the data, thereby optimizing the data processing process; the related parameters of data processing are continuously fed back to the background server in real time for statistics and calculation, which can accurately output the quantitative indicators of the data processing delay situation and accurately evaluate the system running efficiency

[0213] In addition, by calculating the risk score of the to-be-processed risk event according to the severity, influence range and occurrence frequency of the risk event, the to-be-processed risk event can be accurately quantified.

[0214] Embodiment 2

[0215] The following is an embodiment of the system of the present application, which executes the method embodiment of the data management optimization method of embodiment 1. For details not disclosed in the system embodiment of the present application, please refer to the method embodiment of the present application.

[0216] Figure 3 is a structural schematic diagram of an example of the data management optimization system of the present application.

[0217] With reference to Figure 3 , the second aspect of the present disclosure provides a data management optimization system, which adopts the data management optimization method of the first aspect of the present application.

[0218] Specifically, the data management optimization system 100 comprises a data collection layer 10, a data collection layer 20, a data processing layer 30, a risk monitoring layer 40, an alarm issuing layer 50, and an event processing layer 60.

[0219] In the data collection layer 10, based on the collection strategy, the log and business data of each sub-node are collected, and the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation. The data collection layer 20 is used to transmit the collected log and business data of each sub-node to the central node, and the multi-type data processing is carried out in the data processing layer 30, and the quantitative index of data processing delay condition is calculated to optimize the data processing process. In the risk monitoring layer 40, the data after multi-type data processing is monitored in real time, and the risk event recognition model is used to identify the risk event, and the risk score of the risk event is calculated, wherein the risk monitoring result is analyzed and calculated to update the monitoring rules and features of the risk event recognition model. In the alarm issuing layer 40, when the calculated risk score exceeds the preset threshold, the alarm mechanism is triggered, and the alarm information is pushed to the security management personnel of the related sub-node.

[0220] According to an optional embodiment, the following expression is used to calculate the data processing delay index for evaluating the processing efficiency of the data analysis layer on the data to optimize the data processing process:

[0221]

[0222] Wherein Y represents the processing delay index of the current data of the current sub-node in the data processing layer, which is used to measure the real-time performance of the current data processing; h 基础 represents the basic time consumption for different data processing, specifically including query processing; 动态 represents the dynamic delay penalty generated by dynamic delay processing, specifically including additional time consumption caused by resource competition, data skew, and system load; 理论 represents the theoretical optimal processing time determined according to the optimal time consumption in the specified historical period and hardware factors, and the specified historical period includes the historical period of seven days, ten days, fifteen days, thirty days, and sixty days from the current time to the past; 实时 represents the real-time performance influence factor, which is determined according to the collection priority of the business data and the data invalidation factor, Q b represents the weight of the bth business rule, P b represents the bth rule matching degree, b and n are positive integers, n represents the number of business rules, and the business rules include real-time transaction rules, offline report rules, and data operation rules of sub-nodes.

[0223] Specifically, when the calculated processing delay index is greater than a specified value, resource isolation is performed on resources required for processing the current data, or dynamic expansion is performed to improve the processing speed of the current data.

[0224] According to the scientific research implementation, the following expression is used to calculate the dynamic delay penalty:

[0225] d 动态 = Z 资源 * S 倾斜 * f 负载

[0226] wherein d 动态 represents a dynamic delay, specifically including additional time consumption caused by resource competition, data skew, and system load;

[0227] Z 资源 represents a resource competition penalty, CPU 当前使用率 represents the current usage rate of CPU, I 当前使用率 represents the current usage rate of memory, z 系数 represents a competition sensitivity coefficient, which can be determined according to the corresponding sub-node priority and the corresponding data priority; S 倾斜 represents a data skew penalty, q 倾斜 represents a data skew sensitivity index, which is determined according to the relationship between the data amount of the current partition and the average data amount of the partition, q 平均 represents the average data amount of all partitions; f 负载 represents a system load factor of a data management optimization system, r 当前 represents the current concurrent task amount, r 基准 represents a reference concurrent amount, f 指数 represents a load sensitivity index.

[0228] According to an optional implementation, based on a collection strategy, the logs and business data of each sub-node are collected, and a data collection success rate and a data collection delay evaluation value are calculated to measure the collection situation, including:

[0229] The collection strategy includes configuring collection parameters according to actual needs, calculating evaluation parameters, and determining whether the evaluation parameters meet preset values, wherein the evaluation parameters include a data collection success rate and a data collection delay evaluation value, and the following expression is used to calculate the data collection success rate:

[0230]

[0231] wherein C 成功 represents a data collection success rate from the i th sub-node to the n th sub-node; L irepresents the actual amount of data successfully collected at the i-th sub-node, Z i represents the data integrity coefficient of the data successfully collected at the i-th sub-node; E i represents the data reliability coefficient of the data successfully collected at the i-th sub-node, n is the total number of sub-nodes, i represents the i-th sub-node, i and n are positive integers, specifically 1, 2,..., n; a represents a first weight corresponding to the actual amount of data successfully collected at the i-th sub-node; b represents a second weight corresponding to the data integrity coefficient of the data successfully collected at the i-th sub-node; g represents a third weight corresponding to the data reliability coefficient of the data successfully collected at the i-th sub-node; G i represents the expected amount of data that should be collected at the i-th sub-node.

[0232] According to an optional embodiment, the following expression is used to calculate the data collection delay evaluation value for evaluating the response speed of the collection agent to the data:

[0233]

[0234] wherein C 延迟 represents the response speed of the collection agent when collecting data at each sub-node; C 基础 represents the base delay determined by the data generation time, collection completion time, and transmission time at each sub-node; B 惩罚 represents a fluctuation penalty parameter represented by the historical delay mean or historical delay standard deviation in a specific historical time period; F 事件 represents an event delay parameter involved in the risk event, which is specifically determined according to the number of risk events, the number of sub-nodes affected by each risk event, the occurrence time of the risk event, and the business tolerance duration.

[0235] According to an optional embodiment, one or more risk event identification models are constructed using the following algorithms: random forest algorithm, XGBoost algorithm, neural network algorithm.

[0236] For a risk event identification model including a training process, the following data of each sub-node collected is labeled with a risk type to obtain a training data set for training the risk event identification model: sub-node log data after multiple data processing, event recording process; the risk type includes vulnerability threat type, login failure type, IP flow risk type, and attack threat type.

[0237] According to an optional embodiment, the collection log of the current sub-node and the real-time analysis input of the central node are input into the risk monitoring layer, the relevant features are calculated, and the calculated relevant features are input into the risk event identification model to output a risk score and alarm information.

[0238] According to an optional embodiment, a risk score of each to-be-processed risk event is calculated according to a risk scoring mechanism of the risk monitoring layer to quantify each risk event, and a risk false alarm rate of the to-be-processed risk event is calculated to optimize parameters of a risk event identification model.

[0239] It should be noted that the data management optimization method performed by the data management optimization system in Embodiment 2 is the same as the content of the data management optimization method in Embodiment 1, and therefore the same part is omitted.

[0240] Compared with the prior art, the present application collects the log and business data of each sub-node based on a collection strategy, calculates a data collection success rate and a data collection delay evaluation value to measure the collection situation, and can effectively improve the data collection rate of the sub-node; transmits the collected log and business data of each sub-node to the central node and performs multi-type data processing, calculates a quantitative index of the data processing delay condition, and can effectively optimize the data processing process; monitors the data after multi-type data processing in real time, identifies a risk event by using a self-constructed risk event identification model, calculates a risk score of the risk event, analyzes and calculates the risk monitoring result to update the monitoring rules and feature library of the risk event identification model; when the calculated risk score exceeds a preset threshold, triggers an alarm mechanism, and pushes alarm information to the security management personnel of the related sub-node, thereby effectively realizing synchronous risk event identification and real-time alarm while optimizing the data collection process and the data processing process of the multiple sub-nodes.

[0241] In addition, the data processing delay index is calculated to evaluate the processing efficiency of the data by the data analysis layer, thereby optimizing the data processing process; the related parameters of the data processing are continuously fed back to the background server in real time for statistics and calculation, thereby accurately outputting the quantitative index of the data processing delay condition and accurately evaluating the system running efficiency

[0242] In addition, the risk score of the to-be-processed risk event is calculated according to the severity, influence range and occurrence frequency of the risk event, thereby accurately quantifying the to-be-processed risk event.

[0243] Figure 4 is a structural schematic diagram of an electronic device embodiment according to the present application.

[0244] As shown in Figure 4 , the electronic device is in the form of a general computing device. The processor can be one or multiple and work cooperatively. The present application does not exclude distributed processing, i.e., the processor can be dispersed in different entity devices. The electronic device of the present application is not limited to a single entity, but can also be the sum of multiple entity devices.

[0245] The memory stores computer executable programs, usually machine readable codes. The computer readable programs can be executed by the processor to enable the electronic device to perform the method of the present application, or at least part of the steps in the method.

[0246] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and can also include non-volatile memory, such as read only memory (ROM).

[0247] Optionally, in this embodiment, the electronic device further comprises an I / O interface for data exchange between the electronic device and external devices. The I / O interface can be one or more of several types of bus structures, including memory bus or memory controller, peripheral bus, graphics acceleration port, processing unit, or local bus using any of the bus structures.

[0248] It should be understood that, Figure 4 The electronic device shown is only an example of the present application, and the electronic device of the present application can also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as display screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute the computer readable programs in the memory to implement the method of the present application or at least part of the steps in the method, it can be considered as an electronic device covered by the present application.

[0249] From the above description of the embodiments, those skilled in the art will readily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, as Figure 5 As shown, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or on a network, and includes a number of commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to perform the above-mentioned method according to the embodiments of the present application.

[0250] The software product can employ any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0251] The computer readable storage medium can include a data signal transported over a carrier wave and can be baseband or propagated along with carriers. The propagated carrier can take any suitable form, including but not limited to electro-magnetic, optical, or any suitable combination thereof. A computer readable storage medium can be any medium (tangible or non-tangible) that can store data for use by or in connection with the computer system or device. The program code embodied on the computer readable storage medium can be transmitted using any carrier wave appropriate to the communication context, including but not limited to wireless, wire line, optical, radio frequency (RF), or any suitable combination thereof.

[0252] The program code can be executed by one or more programmable processors, which can be implemented as one or more microprocessors, microcontrollers, digital signal processors, embedded processors, programmable logic devices, FPGAs, PLDs, controller, state machines, gated logic, discrete hardware circuits, and / or other devices that can be configured to operate based on a program. It will be appreciated that the described embodiments can be implemented in computer languages including, but not limited to, Java, C++, C# or similar languages. It will be further appreciated that the program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider (ISP).

[0253] The computer readable medium described above can bear one or more programs, when the one or more programs are executed by the device, the computer readable medium implements the data interaction method of the present disclosure.

[0254] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, and can also be changed in one or more devices different from the embodiments. The modules of the above-mentioned embodiments can be combined into one module, or can be further split into a plurality of sub-modules.

[0255] Through the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of commands to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the method according to the embodiments of the present application.

[0256] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs.

[0257] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments can be used, and other changes can be made, without departing from the spirit or scope of the subject matter presented herein.

[0258] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data management optimization method, characterized by, The data management optimization method comprises: Based on the collection strategy, the log and business data of each sub-node are collected, the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation; The collected log and business data of each sub-node are transmitted to the central node, and multi-type data processing is performed to calculate the quantitative indicators of data processing delay condition to optimize the data processing process; Real-time monitoring is performed on the data after multi-type data processing, a risk event identification model is constructed to identify risk events, and the risk score of the risk event is calculated, wherein the risk monitoring result is analyzed and calculated to update the monitoring rules and feature library of the risk event identification model; When the calculated risk score exceeds a preset threshold, an alarm mechanism is triggered, and alarm information is pushed to the security management personnel of the relevant sub-node.

2. The data management optimization method of claim 1, wherein the data processing delay indicator is calculated using the following expression to evaluate the processing efficiency of the data analysis layer on the data to optimize the data processing process: Specifically, when the calculated processing delay indicator is greater than a specified value, resource isolation is performed on the resources required for processing the current data, or dynamic expansion is performed to improve the processing speed of the current data. Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing; 基础 h represents a basic time consumption for different data processing, and specifically includes query processing; 动态 h represents a dynamic delay penalty generated by dynamic delay processing, and specifically includes additional time consumption caused by resource competition, data skew, and system load; 理论 S represents a theoretical optimal processing time determined according to an optimal time consumption in a specified historical time period and a hardware factor, and the specified historical time period includes a historical time period of seven days, ten days, fifteen days, thirty days, or sixty days calculated from the current time to the past; 实时 Q represents a real-time performance impact factor determined according to a data collection priority and a data invalidation factor of a business, Q represents a real-time performance impact factor determined according to a data collection priority and a data invalidation factor of a business, b P represents a weight of the bth business rule, b P represents a weight of the bth business rule, b and n are positive integers, n represents a number of business rules, and the business rules include real-time transaction rules, offline report rules, and data operation rules of child nodes. Further comprising:

3. The data management optimization method of claim 2, wherein, The dynamic delay penalty is calculated using the following expression: The collection strategy comprises configuring collection parameters according to actual needs, calculating evaluation parameters, and determining whether the evaluation parameters meet the preset value, wherein the evaluation parameters include the data collection success rate and the data collection delay evaluation value, and the data collection success rate is calculated using the following expression: d 动态 = Z 资源 x S 倾斜 x f 负载 wherein d 动态 represents dynamic delay, specifically including additional time consumption caused by resource contention, data skew, system load; Z 资源 represents resource competition penalty, CPU 当前使用率 represents the current usage rate of CPU, I 当前使用率 represents the current usage rate of memory, z 系数 represents the competition sensitivity coefficient, which can be determined according to the corresponding child node priority and the corresponding data priority; S 倾斜 denotes a data skew penalty, q 倾斜 denotes a data skew sensitive index, determined by the relationship between the data volume of the current partition and the partition average data volume, q 平均 denotes the partition average data volume of all partitions; f 负载 a system load factor indicative of data management optimization system load, 当前 r represents a current concurrency level, 基准 f represents a baseline concurrency level, 指数 f represents a load sensitivity index.​ 4. The data management optimization method of claim 1, wherein, Further comprising: The data collection delay evaluation value is calculated using the following expression to evaluate the response speed of the collection agent to the data: wherein C 成功 represents the data collection success rate from the i-th sub-node to the n-th sub-node; L i represents the actual collection amount of the data successfully collected at the i-th sub-node, Z i represents the data integrity coefficient of the data successfully collected at the i-th sub-node; E i represents the data reliability coefficient of the data successfully collected at the i-th sub-node, n is the total number of sub-nodes, i represents the i-th sub-node, i and n are positive integers, specifically 1, 2, …, n; α represents a first weight corresponding to the actual collection amount of the data successfully collected at the i-th sub-node; β represents a second weight corresponding to the data integrity coefficient of the data successfully collected at the i-th sub-node; γ represents a third weight corresponding to the data reliability coefficient of the data successfully collected at the i-th sub-node; G i represents the data collection amount of the data that should be collected at the i-th sub-node.

5. The data management optimization method of claim 4, wherein, Further comprising: One or more risk event identification models are constructed using the following algorithms: random forest algorithm, XGBoost algorithm, and neural network algorithm; C 延迟 represents the response speed generated by the collection agent when collecting data at each sub-node; C 基础 represents the basic delay determined by the data generation time, collection completion time, and transmission time consumption of each sub-node; B 惩罚 represents a fluctuation penalty parameter characterized by using the historical delay mean or historical delay standard deviation in a specific historical time period; F 事件 represents an event delay parameter involved in the risk event, which is specifically determined according to the number of risk events, the number of sub-nodes affected by each risk event, the occurrence time of the risk event, and the business tolerance duration.

6. The data management optimization method of claim 1, wherein, For the risk event identification model including a training process, the following data of each sub-node are labeled by risk type to obtain a training data set for training the risk event identification model: sub-node log data after multi-type data processing and event record process; The risk types include vulnerability threat type, login failure type, IP flow risk type, and attack threat type. Further comprising: The collection log of the current sub-node and the real-time analysis input of the central node are input into the risk monitoring layer, the relevant features are calculated, and the calculated relevant features are input into the risk event identification model to output the risk score and alarm information.

7. The data management optimization method of claim 6, wherein, 8. The data management optimization method of claim 6, wherein The risk score of each risk event to be processed is calculated according to the risk score mechanism of the risk monitoring layer to quantify each risk event; The risk false positive rate of the risk event to be processed is calculated to optimize the parameters of the risk event identification model. ​ ​ 9. A data management optimization system, characterized by, The data management optimization system comprises: In the data collection layer, based on the collection strategy, the log and business data of each sub-node are collected, and the data collection success rate and data collection delay evaluation value are calculated to measure the collection situation; The data aggregation layer transmits the collected log and business data of each sub-node to the central node, performs multi-type data processing in the data processing layer, calculates a quantitative index for determining the data processing delay condition, and optimizes the data processing process; In the risk monitoring layer, the data after multi-type data processing is monitored in real time, a risk event recognition model is constructed, a risk event is recognized, and a risk score of the risk event is calculated, wherein the risk monitoring result is analyzed and calculated to update the monitoring rules and feature library of the risk event recognition model; In the alarm issuing layer, when the calculated risk score exceeds a preset threshold, an alarm mechanism is triggered, and alarm information is pushed to the security management personnel of the related sub-node.

10. The data management optimization system of claim 9, wherein, Comprise: The following expression is used to calculate the data processing delay index for evaluating the processing efficiency of the data analysis layer on the data to optimize the data processing process: Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing;h 基础 Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing;d 动态 Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing;h 理论 Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing;S 实时 Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing; Q b Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing;P b Y represents a processing delay indicator of current data of a current child node in a data processing layer, and is used to measure the real-time performance of the current data processing; Specifically, when the calculated processing delay index is greater than a specified value, resource isolation or dynamic expansion is performed on the resources required for processing the current data to improve the processing speed of the current data.

Citation Information

Patent Citations

  • Method for handling security event

    CN107453909A

  • Data acquisition intelligent management system based on data analysis

    CN116800517A

  • Data security risk assessment system and method

    CN117632633A

  • Online shopping contract generation method and system based on block chain

    CN119809770A

  • Data security log analysis system and method

    CN120162776A

Cited By

  • Operation and maintenance method and device of analog quantity measurement front end, storage medium and electronic equipment

    CN121481520A

  • Data acquisition and analysis method and system for bulk commodity supply and demand indexes

    CN121599711A