Edge data processing method and related device of information security monitoring and early warning AI platform
By standardizing the security alarm data collected by edge computing devices and decoupling the objective function, the optimal task decomposition solution is generated, which solves the complexity of edge computing resource coordination and scheduling, and improves the system's processing efficiency and security situation evaluation capabilities.
Patent Information
- Application Number
- CN202510148164.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In a large-scale distributed network environment, the real-time collection, processing and analysis of massive security alarm data puts higher requirements on the system's computing power and resource utilization efficiency, and the computing resources and storage capacity of edge computing devices are limited. How to efficiently coordinate and schedule these scattered computing resources has become a key issue.
By standardizing the original security alarm data collected by multiple edge computing devices, an objective function containing processing efficiency terms and resource consumption terms is constructed, and the objective function is decoupled into multiple sub-optimization problems of edge computing nodes using Lagrangian multipliers, an initial task allocation scheme is obtained, and an optimal task decomposition scheme is generated through local linearization processing and convex optimization decomposition.
It realizes the optimal configuration of edge computing resources, improves the overall processing efficiency of the system, ensures the convergence and stability of the task allocation plan, solves the load balancing problem of edge nodes, and improves the correlation analysis and situation evaluation capabilities of security alarm data.
Smart Images

Figure CN119621287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge data processing technology, and in particular to an edge data processing method and related devices for an information security monitoring and early warning AI platform. Background Art
[0002] As network security threats become increasingly complex and diverse, traditional centralized security monitoring and early warning systems can no longer meet the needs of real-time processing and rapid response. Especially in large-scale distributed network environments, the real-time collection, processing and analysis of massive security alarm data places higher demands on the system's computing power and resource utilization efficiency.
[0003] Under the edge computing architecture, although delegating data processing tasks to edge nodes can reduce the computing pressure of central nodes, due to the limited computing resources and storage capacity of edge computing devices, how to efficiently coordinate and schedule these decentralized computing resources has become a key issue that needs to be solved. Especially when processing multi-source heterogeneous security data, the spatiotemporal correlation and data quality differences between different data sources further increase the complexity of task scheduling. Summary of the invention
[0004] The main purpose of the present invention is to provide an edge data processing method and related devices for an information security monitoring and early warning AI platform. The present invention improves the utilization efficiency of edge computing resources.
[0005] To achieve the above objectives, the present invention provides an edge data processing method for an information security monitoring and early warning AI platform, comprising the following steps:
[0006] Standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data;
[0007] Based on the standardized feature data, an objective function including a processing efficiency term and a resource consumption term is constructed, and the objective function is decoupled into multiple sub-optimization problems of edge computing nodes using Lagrange multipliers to obtain an initial task allocation solution;
[0008] Performing local linearization and convex optimization decomposition on the initial task allocation scheme to obtain an optimal task decomposition scheme;
[0009] Constructing a state transition probability matrix according to the optimal task decomposition scheme and generating a dynamic task scheduling strategy;
[0010] The dynamic task scheduling strategy is assigned to each edge computing node, and time series analysis and spatial clustering are performed on multi-source data to obtain security situation assessment results.
[0011] The present invention also provides an edge data processing device of an information security monitoring and early warning AI platform, comprising:
[0012] A standardization module is used to standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data;
[0013] A construction module, configured to construct an objective function including a processing efficiency term and a resource consumption term based on the standardized feature data, and decouple the objective function into a plurality of sub-optimization problems of edge computing nodes using Lagrange multipliers to obtain an initial task allocation scheme;
[0014] A decomposition module, used for performing local linearization processing and convex optimization decomposition on the initial task allocation scheme to obtain an optimal task decomposition scheme;
[0015] A generation module, used to construct a state transition probability matrix according to the optimal task decomposition scheme and generate a dynamic task scheduling strategy;
[0016] The allocation module is used to allocate the dynamic task scheduling strategy to each edge computing node, perform time series analysis and spatial clustering on multi-source data, and obtain security situation assessment results.
[0017] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0018] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0019] In summary, the technical solution provided by the present invention realizes the optimal configuration of edge computing resources and improves the overall processing efficiency of the system by constructing an objective function including processing efficiency terms and resource consumption terms, and using the Lagrange multiplier method for task decomposition. The continuous convex approximation method is used to process non-convex optimization problems, and the convergence and stability of the task allocation scheme are guaranteed through iterative optimization and dynamic adjustment. The Markov decision process is introduced to construct a dynamic task scheduling strategy, and the task allocation is adaptively adjusted according to the real-time load status, which effectively solves the load balancing problem of edge nodes. A multi-source data processing mechanism based on time series analysis and spatial clustering is designed to realize the correlation analysis and situation assessment of security alarm data, and improve the threat detection accuracy of the system. Through distributed data processing and parallel computing framework, the response delay of the system is reduced, and the utilization efficiency of edge computing resources is improved through resource sharing and task coordination mechanism. A complete performance evaluation and optimization mechanism is established, and the stable operation of the system under different load conditions is ensured by dynamically adjusting the processing strategy and resource allocation scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a schematic diagram of the steps of an edge data processing method of an information security monitoring and early warning AI platform in one embodiment of the present invention;
[0021] Figure 2 This is a structural block diagram of an edge data processing device of an information security monitoring and early warning AI platform in one embodiment of the present invention;
[0022] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0023] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0025] Reference Figure 1 , this embodiment provides an edge data processing method of an information security monitoring and early warning AI platform, including the following steps:
[0026] S1, standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data;
[0027] Among them, the original security alarm data is collected from multiple edge computing devices, including multi-dimensional attribute information such as alarm type, source IP address, target IP address, port number, and alarm level. Outlier detection and redundant data elimination are performed on the original data. Outlier detection identifies data that significantly deviates from the normal range by setting reasonable thresholds or using statistical methods, such as rules based on mean and standard deviation; while redundant data elimination filters duplicate alarm records to reduce the amount of data and improve the reliability of data to obtain pre-processed data. The timestamp information of the pre-processed data is processed to achieve time alignment. Since edge computing devices are distributed in different network environments, the data they collect are not synchronized in time, resulting in errors in the analysis results between multi-source data. The time consistency between each data record is ensured by uniformly aligning the data timestamps. The data is corrected according to a unified time reference using a method based on the Network Time Protocol (NTP) or other synchronization algorithms to obtain time synchronization data. The alarm feature attributes in the time synchronization data are structured. For example, unstructured text or log information is converted into structured field representation to generate structured standard data that meets the analysis requirements. The structured standard data is feature classified and encoded, and a feature vector template is established for the data based on the alarm type, alarm level, and network address information (such as IP address and port number). Each template consists of several specific dimensions. For example, the alarm type and alarm level are mapped to discrete values, while the IP address and port number are converted into numerical features through hash coding or word embedding technology. Through this step, the original non-numerical data is converted into a coded feature sequence suitable for machine learning and statistical analysis. In order to explore the potential relationship in the data, the coded feature sequence is subjected to correlation analysis. By calculating the correlation index between different features (such as Pearson correlation coefficient or mutual information), the intrinsic connection between the alarm type and the network behavior pattern is revealed, and the associated feature data is generated accordingly. The associated feature data is subjected to data dimensionality reduction and standardization. In the dimensionality reduction process, the main feature dimensions in the data are extracted by methods such as principal component analysis or linear discriminant analysis, thereby reducing the dimension of the feature space. At the same time, through standardization processing, the feature values of different dimensions are normalized to the same dimensional range, such as compressing them to [0,1] or normalizing them in the form of zero mean and unit variance, and finally obtaining standardized feature data.
[0028] S2, constructs an objective function including processing efficiency and resource consumption terms based on the standardized feature data, and uses Lagrange multipliers to decouple the objective function into multiple sub-optimization problems of edge computing nodes to obtain the initial task allocation plan;
[0029] Specifically, the task complexity is calculated for the standardized feature data. For the collected standardized feature data, the complexity of each task is calculated in combination with the total amount of alarm data, the dimension of the feature, and the time requirement for task processing, and a task processing efficiency evaluation matrix is generated. Each element of the matrix represents the processing efficiency of a specific task on a certain edge computing node, and the efficiency is determined by the task completion time and computing performance. In the process of constructing the processing efficiency evaluation matrix, the possible processing performance of each task on different computing nodes is quantified by weighing the task characteristics and the node computing power. At the same time, in order to evaluate the resource usage of the edge computing node, the resource status of each node is monitored in real time. The detection of resource status includes multiple key dimensions such as computing load, memory occupancy, and energy consumption indicators, and these monitoring results are used to construct a resource consumption evaluation matrix. Each element of the matrix describes the resource consumption required to process a task on a specific node, including key information such as the load of hardware resources and energy consumption. By combining the task processing efficiency evaluation matrix with the resource consumption evaluation matrix, the trade-off relationship between the system's performance and resources during task allocation is described. Based on the processing efficiency term and the resource consumption term, an objective function is constructed to simultaneously optimize the task processing efficiency and resource consumption. The design of the objective function needs to consider multiple variables, including task processing time, computing load, and node energy consumption. At the same time, certain constraints need to be set for the objective function, such as task completion time constraints and resource utilization constraints. The task completion time constraint ensures that all tasks can be completed within the specified time, while the resource utilization constraint is used to limit the resource usage ratio of a single node, avoiding overload of some nodes while ensuring the overall balanced operation of the system. Through these constraints, a constrained optimization model is formed. In order to effectively solve the constrained optimization model, Lagrange multipliers are introduced to convert the constraints of the objective function into additional Lagrangian terms to form a Lagrangian function. The construction of the Lagrangian function is based on the dual decomposition principle. By converting the original optimization problem into a Lagrangian relaxation problem, the complex global optimization problem is gradually decomposed into smaller sub-problems. When constructing the Lagrangian function, Lagrangian multipliers are introduced to represent the weights of the constraints so that the constraints are reflected in the relaxation problem. The variables of the Lagrangian relaxation problem are separated. The global optimization problem is decomposed into multiple sub-optimization problems at the edge computing node level, where each sub-problem is only associated with the computing power and resource status of a specific node. This decomposition method effectively reduces the complexity of problem solving, allowing each edge node to independently solve its corresponding sub-problem without relying on the calculation results of other nodes. The solution of node-level sub-problems requires comprehensive consideration of the resource status and task characteristics of the current node, and determines the optimal task allocation plan by optimizing the balance between computing performance and resource consumption.By solving all sub-optimization problems and based on the real-time resource status and task characteristics of each edge computing node, the global task allocation matching calculation is completed and the initial task allocation plan is generated.
[0030] A variable correlation matrix is constructed for the Lagrangian relaxation problem, which is used to quantify the coupling relationship between task processing delay and resource consumption. Each element in the variable correlation matrix represents the degree of correlation between a pair of variables. By analyzing these correlations, we can identify which variables have strong dependencies, thereby guiding the subsequent grouping process. According to the analysis results, the variables are subjected to correlation analysis, and the variables with strong correlation are grouped together to obtain the variable grouping results. Based on the variable grouping results, the main variables and the subordinate variables are extracted, and the variable separation criterion matrix is constructed according to the dependency relationship between the variables. The criterion matrix defines the rules for separating variables to ensure that the changes in the main variables can effectively drive the responses of the subordinate variables and form a variable separation scheme. The task processing capability indicators of the edge computing nodes are evaluated according to the variable separation scheme. The evaluation of task processing capability depends not only on the total amount of computing resources of the node, but also on its performance in processing delay, resource utilization and load balancing. By combining these three indicators, a node capability scoring vector is constructed to quantify the relative capability of each node in different task scenarios. The calculation results of the node capability scoring vector divide each node into different capability levels, thereby generating a node grading result to reflect the performance differences between nodes. Based on the node classification results, the global optimization problem is processed in blocks. The global problem is divided into several independent sub-problems according to the mapping relationship between tasks and resources. A sub-problem mapping table is constructed, which describes the allocation relationship between tasks and nodes. At the same time, the specific task scope that each node should undertake is clarified by combining the computing resource capacity and task processing requirements. The constraints are reasonably allocated. By analyzing the resource sharing relationship and task dependency relationship between nodes, a set of constraint allocation rules are established. These rules are used to guide the distribution and adjustment of constraints to ensure that the decomposed sub-problems still meet the global constraints. The allocation of constraints needs to comprehensively consider the resource boundaries and task characteristics of the nodes. For example, task constraints can be appropriately reduced on high-load nodes, while more computing tasks can be allocated to low-load nodes. Through this process, a sub-problem constraint set is generated to describe the conditions that each sub-problem must meet when solving it. Based on the sub-problem constraint set, the boundaries of the decomposed optimization problem are regularized. The solution domain of each sub-problem is clarified to ensure that the node completes the task allocation within its own computing resource range. By analyzing the computing resource boundaries of each node, the specific solution domain of each sub-problem is determined, so that all sub-problems can run independently within the resource constraints without interfering with each other. Through the above steps, multiple sub-optimization problems are finally obtained, and each sub-problem is solved on its corresponding edge node, while meeting the global optimization objectives and constraint requirements.
[0031] S3, perform local linearization and convex optimization decomposition on the initial task allocation plan to obtain the optimal task decomposition plan;
[0032] It should be noted that the initial task allocation scheme is substituted into the continuous convex approximation model for non-convex component identification. Non-convex components are parts of the objective function that are difficult to optimize directly due to high-order nonlinearity or complex constraints. These components will significantly increase the complexity of the solution. Through the continuous convex approximation technology, the non-convex components are locally analyzed and locally linearly expanded using methods such as Taylor expansion, and the complex non-convex problem is approximated as a set of initial linearized subproblems. The initial linearized subproblems are divided into convex regions. The entire task allocation space is divided into several convex subregions, each of which is described by a simpler mathematical model. Based on the two key factors of task processing delay matrix and resource utilization matrix, a piecewise convex optimization model is constructed. The task processing delay matrix represents the time consumption of each edge node when processing a specific task, while the resource utilization matrix reflects the resource consumption of the node. By using these two types of matrices as the basis for division, a set of convex subproblem sequences are generated, each of which corresponds to a specific task allocation and resource utilization scenario. The convex subproblem sequence is input into the iterative optimizer for solution. The iterative optimizer uses gradient descent operation to optimize the task allocation scheme for each subproblem. During the gradient descent process, each iteration adjusts the task allocation scheme according to the gradient information of the objective function to gradually approach the local optimal solution. The advantage of gradient descent is that it is computationally efficient and easy to implement, but since it is essentially a local optimization algorithm, in order to ensure the global optimality of the solution, the convergence of the solution is finely controlled in combination with the dynamic step size strategy. After each gradient descent iteration, the generated solution is verified for resource constraints to ensure that it is feasible in the actual computing environment. The main basis for resource constraint verification includes the task queue length and node computing capacity, which determine whether the edge computing node can complete the assigned tasks on time. By screening the feasibility of iterative solutions, solutions that do not meet resource constraints are eliminated to obtain a preliminary feasible solution set. The solutions in the feasible solution set need to further evaluate their convergence while meeting the basic constraints. The distance between adjacent iterative solutions is calculated using the Euclidean distance criterion, and the adjustment range of the solution is gradually narrowed through the dynamic step size strategy to ensure the convergence of the final solution. After this process, a set of convergent solution sequences is formed. Substitute the converged solution sequence into the objective function for the final optimization selection. The global optimal solution is selected from the converged solution sequence by sorting the objective function values and combining the constraint satisfaction evaluation. The sorting of the objective function is mainly based on the performance of the solution in terms of task processing efficiency and resource consumption, while the constraint satisfaction evaluation ensures that the optimal solution meets the actual constraints of all resources and task allocation while optimizing performance. The optimal solution after sorting and screening is established as the optimal task decomposition solution.
[0033] S4, constructs a state transition probability matrix based on the optimal task decomposition scheme and generates a dynamic task scheduling strategy;
[0034] Specifically, the core feature information in the optimal task decomposition scheme is converted into a vector form for describing the system state. The computing load, remaining resources and task queue length of each edge node in the optimal task decomposition scheme are abstracted into a state description vector, which together constitute the Markov state space of the system. In the state space, each state describes the resource allocation and task execution of the system at a certain moment. The tasks in the Markov state space are prioritized, and resource quotas are divided according to the importance and urgency of different tasks. Combined with the characteristics of alarm data, such as the processing delay requirements of tasks, alarm levels and possible impact ranges, a set of optional task scheduling action sequences are generated to form an action candidate space. Each scheduling action in the action candidate space represents a possible task allocation scheme under the current state. By reasonably designing the action sequence, it is ensured that more resources are allocated to high-priority tasks during the dynamic scheduling process while maintaining resource utilization efficiency. The scheduling scheme in the action candidate space is applied to the historical task processing data, and the statistical characteristics of the historical data are used to calculate the statistics and transition condition probabilities of state transitions. By analyzing the impact of different scheduling actions on state transitions, a state transition matrix is generated. Each element in the state transition matrix represents the probability of the system transferring from the current state to another state. The matrix can reflect the impact of task scheduling on the system state and describe the possible paths of state evolution. The state transition matrix is reward modeled so that the pros and cons of each state transition path can be measured through the immediate reward function. The design of the immediate reward function is based on key factors such as task delay, resource utilization, and load balancing indicators. For example, reducing task delay and improving resource utilization will get positive rewards, while causing load imbalance will result in penalty values. By weighing these factors, the reward function accurately reflects the effects of different state transition schemes. According to the results of the reward function calculation, a state value mapping table is generated to describe the long-term benefit evaluation of different states. The state value mapping table is input into the policy optimization module, and the state-action pairs are evaluated and updated by the value iteration method. In the value iteration process, the value estimate of each state-action pair is recursively updated based on the Bellman equation, and gradually converges to the optimal strategy. The value iteration method can effectively handle problems with large state spaces and find the global optimal solution through iterative optimization. In this process, each iteration combines the current reward and the possible future benefits to achieve global optimization of the dynamic task scheduling strategy. Perform dynamic programming search on the optimal action strategy to calculate the optimal task scheduling sequence under the current state. The goal of dynamic programming search is to generate a specific scheduling plan based on the optimal action strategy given the current state. Through dynamic programming search, a dynamic task scheduling strategy is finally obtained, which can respond to changes in task and resource status in the edge computing environment in real time, ensuring the efficiency of task processing and balanced utilization of system resources.
[0035] S5, distributes the dynamic task scheduling strategy to each edge computing node, performs time series analysis and spatial clustering on multi-source data, and obtains security situation assessment results.
[0036] Among them, the dynamic task scheduling strategy is sent to multiple edge computing nodes. When assigning tasks, parallel processing tasks and specific execution parameters are set for each node according to its computing power and resource status to form a distributed task execution sequence. A time sliding window is set for the multi-source data in the distributed task execution sequence. The window can extract data features within a predefined time range to form the time series change characteristics of various alarm events. Through the dynamic adjustment of the time sliding window, the change trend of the alarm event is monitored in real time, and a time series feature sequence is generated to reflect the occurrence frequency and time interval of each alarm event, and provide the time correlation and periodicity pattern of the event occurrence for subsequent analysis. The time series feature sequence is topologically mapped according to the attack chain propagation path, and the event propagation network and node association relationship are constructed. By analyzing the logical association and triggering sequence between different alarm events, a topological network structure is formed, in which nodes represent various alarm events and edges represent the association relationship between events. By constructing this event propagation network, the propagation path of the attack chain and its impact range are characterized, and a spatial distribution model is generated. The spatial distribution model can describe the distribution of attack events between different nodes and regions. In order to explore the regularity and pattern characteristics of attack behavior, hierarchical clustering operations are performed on the spatial distribution model. Hierarchical clustering classifies attack events with similar propagation patterns into the same family according to the topological structure and node association characteristics of the event propagation network, and divides attack event families and threat scenario patterns. These clustering results help to discover potential security risks and typical attack patterns in the system and form alarm event clustering sequences. The alarm event clustering sequence is input into the association mining module, and the causal chain between events is extracted through deep analysis. In the association mining module, the causal reasoning method is used to derive the causal chain between events according to the occurrence sequence and association characteristics of the events, and form the measurement value of the threat situation. The threat situation measurement value is a quantitative expression of the attack intensity, impact range and association complexity. The threat situation measurement values are combined into a multidimensional feature vector, and the overall security risk score is generated through fusion calculation. The fusion calculation comprehensively considers the time characteristics, spatial distribution characteristics and risk factors in the causal chain, and calculates the overall risk assessment value through weighted summation or machine learning model, and finally obtains the security situation assessment result.
[0037] In one example, the original security alarm data collected by multiple edge computing devices is standardized to obtain standardized feature data, including:
[0038] The original security alarm data is collected through multiple edge computing devices, and outlier detection and redundant data elimination are performed on the alarm type, source IP address, target IP address, port number and alarm level of the original security alarm data to obtain preprocessed data;
[0039] Perform time alignment on the timestamp information of the preprocessed data to obtain time synchronization data, and perform structured conversion on the alarm feature attributes in the time synchronization data to obtain structure standard data;
[0040] Perform feature classification encoding on the structural standard data, establish a feature vector template according to the alarm type, alarm level and network address information, obtain a coding feature sequence, and perform correlation analysis on the coding feature sequence to obtain associated feature data;
[0041] Perform data dimension reduction and standardization on the associated feature data to obtain standardized feature data.
[0042] In this example, the original security alarm data obtained from the edge computing device contains multiple attributes, such as alarm type, source IP address, destination IP address, port number, and alarm level. These data may contain outliers or duplicate records in their original state, and outlier detection and redundant data removal are performed. Outlier detection uses statistical methods, such as the triple standard deviation method, to define the valid data range. Suppose the value of an attribute of the alarm record is , whose mean is , the standard deviation is , then the effective range is defined as When the data value exceeds this range, it is considered an abnormal value. The data is quickly deduplicated through hash mapping technology, and a hash value is generated for each alarm record. , if the hash value of a record already exists in the hash table, the record is considered redundant data and removed. After this step, the cleaned preprocessed data is obtained. The timestamp information in the preprocessed data is aligned to unify the time base of each edge computing device. Due to the time synchronization error of the device, the network time protocol (NTP) is used to correct the device time. For the original time of a device , the reference time is , the corrected time is expressed as . In this way, the timestamps of all devices are aligned to a unified time base to generate time synchronization data. The data is structured during time alignment to convert the unstructured alarm log into structured data. The original log is recorded in free text form, such as a string containing information such as time, alarm level, and attack type. By parsing the text content and extracting field information, it is converted into tabular structured standard data, including timestamps, alarm types, IP addresses, and specific values of other attributes. Feature classification and encoding of structured standard data. Convert alarm type, alarm level, and network address information into numerical features acceptable to machine learning algorithms. For example, define a feature vector template to represent the features of each record. Assume that the alarm type set is , the alarm level set is , the network address information set is , then each record is represented as a vector .in, It is a binary variable indicating whether a record contains the corresponding type, level or network address feature. Correlation analysis is performed on the coded feature sequence to identify the potential relationship between features. The correlation analysis is completed by calculating the Pearson correlation coefficient, and the formula is:
[0043] ;
[0044] in, and are vectors of two characteristic variables, represents its covariance, and is its standard deviation. By calculating the correlation, the degree of association between different features is quantified, which in turn guides the decision of feature selection and feature merging. Features with high correlation represent the same potential pattern, and after analysis, they are appropriately merged to reduce redundant features. The associated feature data is reduced in dimension and standardized to generate standardized feature data. The dimensionality reduction method uses principal component analysis to extract the most important feature dimensions through eigenvalue decomposition. Assume that the feature matrix is , whose covariance matrix is , the projection matrix is obtained by eigenvalue decomposition , the feature after dimension reduction is expressed as By retaining the main components, the dimension of the data is reduced and the computational complexity is reduced. The features after dimensionality reduction are standardized, for example, using the zero mean standardization formula:
[0045] ;
[0046] in, is the standardized eigenvalue, is the original value, and are the mean and standard deviation respectively. Standardization ensures that different features have the same scale in subsequent analysis, avoiding the influence of feature value range differences on the analysis results. After the above steps, standardized feature data is finally generated.
[0047] In one example, an objective function including processing efficiency and resource consumption is constructed based on standardized feature data, and the objective function is decoupled into multiple sub-optimization problems of edge computing nodes using Lagrange multipliers to obtain an initial task allocation solution, including:
[0048] The task complexity is calculated for the standardized feature data, and a task processing efficiency evaluation matrix is constructed based on the alarm data volume, feature dimension, and processing time requirements to obtain the processing efficiency item;
[0049] Detect the resource status of edge computing nodes, and build a resource consumption evaluation matrix based on computing load, memory occupancy, and energy consumption indicators to obtain resource consumption items;
[0050] An objective function is constructed based on processing efficiency items and resource consumption items, and task completion time constraints and resource utilization constraints are set for the objective function to obtain a constrained optimization model;
[0051] Lagrange multipliers are introduced into the constrained optimization model, and Lagrange functions are constructed according to the dual decomposition principle to obtain the Lagrange relaxation problem.
[0052] The variables of the Lagrangian relaxation problem are separated, and the global optimization problem is decomposed into node-level sub-problems according to the computing power of each edge computing node to obtain multiple sub-optimization problems;
[0053] Multiple sub-optimization problems are solved, and task allocation matching calculations are performed according to the resource status and task characteristics of the edge nodes to obtain the initial task allocation plan.
[0054] In this example, the task complexity is calculated for the standardized feature data. The task complexity depends on factors such as the amount of alarm data, feature dimensions, and processing time requirements. Assume that the total amount of alarm data is , the feature dimension is , the processing time requirement is , then the task complexity is expressed as a function:
[0055] ;
[0056] in, , , is the weight coefficient, which measures the impact of data volume, feature dimension and time requirement on task complexity. By calculating the complexity of each task, the task processing efficiency evaluation matrix is generated, which is recorded as For the matrix elements , which means that at the edge node Processing tasks The efficiency is defined as:
[0057] ;
[0058] in, It's a task The complexity of Is a node Processing tasks The time required. Through this matrix, the processing efficiency of each task on different edge nodes is quantified. At the same time, the resource status of the edge computing node is detected, mainly including real-time monitoring of computing load, memory usage and energy consumption indicators. The current computational load is , the memory usage is , the energy consumption is , then the resource consumption is expressed as:
[0059] ;
[0060] in, , , is the weight coefficient, which is used to adjust the impact of load, memory and energy consumption on resource consumption. The resource consumption calculation results form a resource consumption evaluation matrix, which is recorded as , where the matrix elements Representation Node Execute the task Based on the processing efficiency term and the resource consumption term, an objective function is constructed to optimize the task allocation. The design of the objective function requires a trade-off between task completion time and resource consumption, which is defined as:
[0061] ;
[0062] in, is the total number of edge nodes, is the total number of tasks, is the allocation decision variable. If the task Assign to Node ,but ,otherwise . Weight and Represent the importance of efficiency term and resource consumption term respectively. In order to constrain task completion time and resource utilization, the objective function needs additional constraints. For example, the task completion time constraint is:
[0063] ;
[0064] The resource utilization constraints are:
[0065] ;
[0066] in, is the maximum completion time of the task, is the resource capacity of the node. To solve the above constrained optimization problem, Lagrange multipliers are introduced to transform the constraints into part of the objective function and construct the Lagrange function:
[0067] ;
[0068] in, and are Lagrange multipliers, which represent the degree of relaxation of task completion time and resource constraints respectively. Through the dual decomposition principle, this global optimization problem is decomposed into multiple node-level sub-problems. When separating the variables of the sub-problems, they are divided according to the computing power of the nodes. For example, for the node , we only need to optimize the set of tasks it processes, which constitutes the sub-problem:
[0069] ;
[0070] in, is assigned to the node Each sub-problem is solved independently, and the common method is gradient descent. After combining the solution results of the sub-problems, an initial task allocation scheme is formed. This scheme optimizes the balance between computing efficiency and resource consumption while satisfying the constraints.
[0071] In one example, the Lagrangian relaxation problem is variable separated, and the global optimization problem is decomposed into node-level sub-problems according to the computing power of each edge computing node, resulting in multiple sub-optimization problems, including:
[0072] The variable correlation matrix is constructed for the Lagrangian relaxation problem, and the variable correlation analysis is performed based on the coupling relationship between task processing delay and resource consumption to obtain the variable grouping results;
[0073] Based on the variable grouping results, the main variables and the subordinate variables are extracted, and the variable separation criterion matrix is constructed according to the dependency relationship between the variables to obtain the variable separation scheme;
[0074] The task processing capability index of edge nodes is calculated for the variable separation scheme, and the node capability scoring vector is constructed based on processing delay, resource utilization and load balancing to obtain the node classification result;
[0075] The global optimization problem is divided into blocks according to the node classification results, and a sub-problem mapping table is constructed based on the computing resource capacity and task processing requirements to obtain a problem decomposition framework;
[0076] Assign constraints to the problem decomposition framework, and establish constraint assignment rules based on the resource sharing relationship and task dependency relationship between nodes to obtain the sub-problem constraint set;
[0077] The decomposed optimization problem is bounded based on the sub-problem constraint set, and the sub-problem solution domain is determined according to the computing resource boundary of each node to obtain multiple sub-optimization problems.
[0078] In this example, in the Lagrangian relaxation problem, each variable represents the decision of task allocation and resource utilization, so there is a coupling relationship between these variables. In order to analyze this relationship, a variable correlation matrix is constructed. Let the variable set be ,in Representation Task Is it assigned to a node? Variable relevance is defined by the coupling of task processing latency and resource consumption. For example, task At the node The processing delay on , resource consumption is , then the elements of the variable correlation matrix are defined as:
[0079] ;
[0080] in, is the correlation between latency and resource consumption, calculated using the Pearson correlation coefficient. Each element of Representation Task At the node and tasks At the node The coupling strength of the matrix Perform analysis to identify the correlation between variables and group the variables with strong correlation. Based on the variable grouping results, extract the main variable and the dependent variable in each group. The main variable refers to the variable that has a strong influence on other variables in the group, while the dependent variable depends on the change of the main variable. Include The main variable is determined by calculating the influence of the main variable on other variables in the group. For example, the influence of the main variable on the subordinate variable is defined as:
[0081] ;
[0082] in, Representation variables The variable with the greatest influence is selected as the main variable. According to the dependency relationship between the main variable and the subordinate variable, a variable separation criterion matrix is constructed. The elements of the criterion matrix define the constraint rules of the main variable on the subordinate variable, which is used to guide the variable separation process and obtain a complete variable separation scheme. After obtaining the variable separation scheme, the task processing capacity index of the edge node is calculated to evaluate the load carrying capacity of each node. The task processing capability of the node is comprehensively measured by processing latency, resource utilization, and load balancing. The task set is , then its processing delay index is:
[0083] ;
[0084] in, Is a node Processing tasks The delay, Is a node The total number of tasks on the system. The resource utilization metric is defined as:
[0085] ;
[0086] in, It's a task At the node resource consumption on Is a node The load balancing degree is calculated by the standard deviation of node resources:
[0087] ;
[0088] in, is the average resource utilization of all nodes, is the total number of nodes. Combining the above indicators, we construct a node capability scoring vector:
[0089] ;
[0090] in, , , is the weight of each indicator. According to the scoring vector Nodes are graded to obtain node classification results. Based on the node classification results, the global optimization problem is divided into blocks and tasks are assigned to different node groups. and task processing requirements , construct a sub-problem mapping table to determine the set of tasks that each node is responsible for. The mapping table clarifies the correspondence between nodes and tasks, forming a problem decomposition framework. Based on the problem decomposition framework, constraints are assigned. Constraint assignment rules are established based on the resource sharing relationship and task dependency relationship between nodes. For example, if the task Dependent tasks The completion of
[0091] ;
[0092] Constraint rules ensure the rationality of task scheduling and ultimately generate subproblem constraint sets. Based on the subproblem constraint sets, the decomposed optimization problem is bounded. The goal of bounded bounding is to determine the solution domain of each subproblem, such as limiting the number of task assignments or the range of resource consumption. A collection of tasks , whose boundary constraints are defined as:
[0093] ;
[0094] in, Is a node The maximum number of tasks that can be processed. Based on the decomposed sub-problem optimization model, a distributed solution algorithm, such as gradient descent or reinforcement learning model, is used to calculate the optimal task allocation plan for each node. Through this process, complex global problems are converted into multiple independent sub-problems and solved efficiently in a distributed manner, achieving optimal resource utilization and efficient task scheduling.
[0095] In one example, the initial task allocation scheme is subjected to local linearization and convex optimization decomposition to obtain an optimal task decomposition scheme, including:
[0096] Substitute the initial task allocation scheme into the continuous convex approximation model to identify non-convex components, perform local linear expansion on the non-convex components, and obtain the initial linearized subproblem;
[0097] The initial linearized subproblem is divided into convex regions, and a piecewise convex optimization model is established based on the task processing delay matrix and resource utilization matrix to obtain a convex subproblem sequence.
[0098] The convex subproblem sequence is input into the iterative optimizer, and the task allocation scheme of each subproblem is subjected to gradient descent operation to obtain the target iterative sequence;
[0099] Verify resource constraints for each solution in the target iteration sequence, use task queue length and node computing capacity to screen solutions for feasibility, obtain feasible solution set, apply dynamic step size strategy to the feasible solution set, and use Euclidean distance criterion to calculate the convergence degree of adjacent iterative solutions to obtain converged solution sequence;
[0100] The converged solution sequence is substituted into the objective function, and the optimal solution is selected through the objective function value sorting and constraint satisfaction evaluation to obtain the optimal task decomposition scheme.
[0101] In this example, the initial task allocation scheme is determined by the task and resource allocation strategy in the edge computing environment. Let the task set be , the edge node set is ,variable Indicates the task Is it assigned to a node? The initial task allocation scheme uses the matrix Indicates that The goal is to optimize task processing efficiency and resource utilization while satisfying node capacity constraints and task completion time constraints. In the continuous convex approximation model, the objective function contains non-convex components, such as high-order nonlinear terms or discontinuous indicator functions. Assume that the original objective function is , whose non-convex components are Represented by continuous convex approximation technology, It can be expressed approximately as follows using its first-order Taylor expansion:
[0102] ;
[0103] in, is the current iteration point, yes exist This process locally linearizes the non-convex problem and obtains the initial linearized subproblem, which is recorded as:
[0104] ;
[0105] in, , represents the feasible domain of the linearized problem. Based on the initial linearized subproblem, convex region partitioning is performed. Convex region partitioning is to divide the entire task allocation space into several convex regions, and the optimization problem in each region is solved by standard convex optimization technology. Let the task processing delay matrix be ,in Indicates the task At the node The processing delay on the resource utilization matrix is ,in Indicates the task At the node By combining task processing latency and resource utilization, a piecewise convex optimization model is established:
[0106] ;
[0107] in, and is a weight parameter used to balance latency and resource consumption. The piecewise convex optimization model divides the problem into multiple convex subproblems, each of which corresponds to a local convex region. These convex subproblem sequences are input into the iterative optimizer, and the task allocation scheme for each subproblem is optimized by the gradient descent method. Let the current solution be , the gradient is , the update rule is:
[0108] ;
[0109] in, is a dynamic step size, which accelerates convergence by adjusting the step size. In each iteration, the solution is gradually updated according to the gradient information to generate the target iteration sequence For each solution in the target iteration sequence, resource constraint verification is performed to ensure the feasibility of the solution in the actual computing environment. The capacity is ,Task The resource requirement is , then the constraints are:
[0110] ;
[0111] If a solution satisfies all constraints, it is added to the feasible solution set. In order to select the optimal solution, a dynamic step size strategy is applied to the feasible solution set, and the convergence degree between adjacent solutions is calculated using the Euclidean distance criterion. Suppose the solution of the two iterations is and , and its Euclidean distance is:
[0112] ;
[0113] when When it is less than the set threshold, the solution is considered to have converged and a converged solution sequence is generated. Substitute the converged solution sequence into the objective function, and select the optimal solution by sorting the objective function values and evaluating the constraint satisfaction. Suppose the objective function is , then find the solution corresponding to the minimum value by sorting At the same time, the constraint satisfaction of the solution is evaluated, such as the computing resource utilization and the satisfaction ratio of the task completion time. The final optimal task decomposition solution is selected based on the comprehensive sorting results and constraint satisfaction.
[0114] In one example, a state transition probability matrix is constructed according to the optimal task decomposition scheme, and a dynamic task scheduling strategy is generated, including:
[0115] The edge node computing load, remaining resources and task queue length in the optimal task decomposition scheme are constructed into a state description vector to obtain a Markov state space;
[0116] Prioritize and allocate resources to tasks in the Markov state space, construct optional task scheduling action sequences, and obtain the action candidate space;
[0117] Apply the scheduling scheme in the action candidate space to the historical task processing data, calculate the state transition statistics and transition condition probability, and obtain the state transition matrix;
[0118] The state transfer matrix is rewarded and modeled. The immediate reward function is constructed using task delay, resource utilization and load balancing indicators to obtain a state value mapping table. The state value mapping table is input into the strategy optimization module. The state-action pairs are evaluated and updated through the value iteration method to obtain the optimal action strategy.
[0119] Perform dynamic programming search on the optimal action strategy, calculate the optimal task scheduling sequence under the current state, and obtain the dynamic task scheduling strategy.
[0120] In this example, the computational load, remaining resources, and task queue length of the edge nodes are extracted from the optimal task decomposition scheme to construct the state description vector s. Let the set of edge nodes be , each node The current computational load is The remaining resources are , the task queue length is . Then the state description vector is defined as:
[0121] ;
[0122] The state description vector s reflects the resource usage and task allocation status of the current edge computing system. After constructing the Markov state space, the tasks in each state are prioritized and resource quotas are divided to generate optional task scheduling action sequences. Task priority Calculated based on the alert level, the urgency of the task and the processing time requirements, for example:
[0123] ;
[0124] in, represents the computational load of the task, Indicates the processing time requirement of the task, Indicates the length of the task queue. , , is the weight coefficient. By prioritizing tasks and Divide resource quotas and generate feasible task scheduling action sequences . Every action Indicates that some tasks are assigned to specific edge nodes. The scheduling scheme in the generated action candidate space is applied to the historical task processing data to simulate the changes in system state under different scheduling strategies. The next state The frequency of state transition is calculated and the conditional probability of transition is estimated:
[0125] ;
[0126] Transition probability matrix Each element of Indicates taking action in state s The probability of the transition probability matrix is the core of the Markov decision process, describing the dynamic evolution of the system during task scheduling. Based on the transition probability matrix, the state and action are rewarded to guide strategy optimization. Instant reward function Defined by task latency, resource utilization, and load balancing indicators, for example:
[0127] ;
[0128] in, is the average latency of all tasks, is the standard deviation of resource utilization of each node, indicating the degree of load balancing, is the overall resource utilization of the system, , , is the weight coefficient. The immediate reward function reflects the effect of the current scheduling strategy. The reward function calculation result is mapped to the state value to generate a state value mapping table . The state-action pair is evaluated and updated through the value iteration method to find the optimal strategy. The value iteration method is based on the Bellman optimal equation:
[0129] ;
[0130] in, Yes Status The value of is a discount factor that weighs immediate rewards and future rewards. Through multiple iterations, the state value and the optimal strategy Gradually converge. After obtaining the optimal strategy After that, the optimal task scheduling sequence is generated in the current state through dynamic programming search. Dynamic programming uses recursion to gradually derive the optimal solution sequence for subsequent states based on the current state and the optimal action. The final dynamic task scheduling strategy is expressed as a function:
[0131] ;
[0132] This strategy dynamically adjusts the task allocation scheme according to the real-time state description vector s, so that the system performance indicators are optimized.
[0133] In one example, a dynamic task scheduling strategy is assigned to each edge computing node, and time series analysis and spatial clustering are performed on multi-source data to obtain security situation assessment results, including:
[0134] The dynamic task scheduling strategy is sent to multiple edge computing nodes, and parallel processing tasks and execution parameters are set on each edge computing node to obtain a distributed task execution sequence.
[0135] A time sliding window is set for the multi-source data in the distributed task execution sequence to extract the time series change characteristics of various alarm events and obtain the time series feature sequence;
[0136] The time series feature sequence is topologically mapped according to the attack chain propagation path, the event propagation network and node association relationship are constructed, and the spatial distribution model is obtained. The spatial distribution model is then hierarchically clustered to divide the attack event families and threat scenario patterns to obtain the alarm event cluster sequence.
[0137] The clustering sequence of alarm events is input into the association mining module, and the causal chain between events is extracted through deep analysis to obtain the threat situation measurement value. The threat situation measurement value is combined into a multi-dimensional feature vector, and the overall security risk score is generated through fusion calculation to obtain the security situation assessment result.
[0138] In this example, the dynamic task scheduling strategy is sent to multiple edge computing nodes. The dynamic task scheduling strategy sets the execution parameters of parallel tasks for each edge node based on real-time computing resources and task requirements, including task priority, resource allocation, and processing time limit. Suppose the node set is , the task set is . Task Allocation Matrix Indicates that Indicates the task Assign to Node ,otherwise The task allocation satisfies the following constraints:
[0139] ;
[0140] ;
[0141] in, It's a task At the node The resource requirements on Is a node The distributed task execution sequence is executed in parallel by multiple nodes to ensure efficient task processing. During the task execution process, the multi-source data in the distributed task execution sequence is analyzed in time series. In order to capture the time characteristics of the alarm event, a time sliding window is set for the multi-source data. Assume that the time window length is , the first The time window is Through the sliding window, the time series change characteristics of each type of alarm event are extracted to construct the time series feature sequence ,in Represents a time window Neidi The frequency of occurrence of similar alarm events. The sliding window operation dynamically captures the time pattern of alarm behavior by updating the event frequency in real time. The generated time series feature sequence is topologically mapped according to the attack chain propagation path to construct the event propagation network and node association relationship. Suppose the alarm event set is , the node association relationship is represented by the adjacency matrix Indicates that Indicates an event and There is a direct correlation, Indicates no association. The event propagation network is formed by analyzing the occurrence sequence and triggering relationship of alarm events, and its propagation path is represented as a directed graph The spatial distribution model of the event propagation network describes the diffusion characteristics of attack behaviors between different nodes and regions. Based on the spatial distribution model, a hierarchical clustering operation is performed on the event propagation network. Hierarchical clustering divides attack event families and threat scenario patterns by recursively merging similar event nodes. Suppose the similarity measure between nodes is , defined based on the propagation path and node attributes. For example, the Euclidean distance between nodes is:
[0142] ;
[0143] in, and Is a node and The feature vector of . Through hierarchical clustering, events with high similarity are merged to form a cluster sequence of alarm events. The cluster sequence of alarm events is input into the association mining module for in-depth analysis to extract the causal chain between events. The causal chain is constructed by analyzing the time series and propagation relationship of events. Assume that the strength of the causal relationship is , calculated by time-delayed correlation or Granger causality test:
[0144] ;
[0145] in, Indicates an event Is it an event? The causal chain is constructed by constructing a correlation matrix Represents, revealing the dynamic triggering relationship between events. The extracted causal chain generates threat situation measurement values, which are generated by quantitative calculation of attack intensity, diffusion range and complexity. Let the measurement value vector be , where each Represents a threat situation indicator. Combining these metrics forms a multidimensional feature vector The overall security risk score is generated by fusion calculation, and the fusion model is assumed to be , the overall risk score is:
[0146] ;
[0147] in, is the first feature vector The weight of the item indicates the degree of influence of the indicator on the overall risk. The fusion model is optimized using a linear weighting method or a nonlinear machine learning model (such as a neural network). The security situation assessment results are presented in the form of risk scores, which provide a decision-making basis for the security protection and strategy optimization of the system. For example, if the risk score within a certain time window is , triggering high-level security warnings and dynamically adjusting resource allocation strategies.
[0148] Reference Figure 2 , this embodiment provides an edge data processing device of an information security monitoring and early warning AI platform, including:
[0149] Standardization module 1, used to standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data;
[0150] Construction module 2 is used to construct an objective function including processing efficiency terms and resource consumption terms based on the standardized feature data, and use Lagrange multipliers to decouple the objective function into multiple sub-optimization problems of edge computing nodes to obtain an initial task allocation plan;
[0151] Decomposition module 3, used for performing local linearization and convex optimization decomposition on the initial task allocation scheme to obtain the optimal task decomposition scheme;
[0152] A generation module 4 is used to construct a state transition probability matrix according to the optimal task decomposition scheme and generate a dynamic task scheduling strategy;
[0153] The allocation module 5 is used to allocate the dynamic task scheduling strategy to each edge computing node, perform time series analysis and spatial clustering on multi-source data, and obtain security situation assessment results.
[0154] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be repeated here.
[0155] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0156] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0157] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0158] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0159] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0160] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An edge data processing method for an information security monitoring and early warning AI platform, characterized in that: The following steps are involved: Standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data; Based on the standardized feature data, an objective function including a processing efficiency item and a resource consumption item is constructed, and the objective function is decoupled into multiple sub-optimization problems of edge computing nodes by using Lagrange multipliers to obtain an initial task allocation plan; specifically comprising: performing task complexity calculation on the standardized feature data, and constructing a task processing efficiency evaluation matrix according to the alarm data volume, feature dimension and processing time requirements to obtain a processing efficiency item; performing resource status detection on the edge computing node, and constructing a resource consumption evaluation matrix according to computing load, memory occupancy and energy consumption indicators to obtain a resource consumption item; based on the processing efficiency item and the The objective function is constructed based on the resource consumption item, and the task completion time constraint and resource utilization constraint are set for the objective function to obtain a constrained optimization model; Lagrange multipliers are introduced into the constrained optimization model, and Lagrange functions are constructed according to the dual decomposition principle to obtain a Lagrange relaxation problem; variables are separated for the Lagrange relaxation problem, and the global optimization problem is decomposed into node-level sub-problems according to the computing power of each edge computing node to obtain multiple sub-optimization problems; the multiple sub-optimization problems are solved, and task allocation matching calculations are performed according to the resource status and task characteristics of the edge nodes to obtain an initial task allocation plan; The initial task allocation scheme is locally linearized and convexly optimized to obtain an optimal task decomposition scheme; specifically, the method includes: substituting the initial task allocation scheme into a continuous convex approximation model to identify non-convex components, and performing local linear expansion on the non-convex components to obtain an initial linearized subproblem; dividing the initial linearized subproblem into convex regions, and establishing a piecewise convex optimization model based on a task processing delay matrix and a resource utilization matrix to obtain a convex subproblem sequence; inputting the convex subproblem sequence into an iterative optimizer, performing a gradient descent operation on the task allocation scheme of each subproblem to obtain a target iteration sequence; performing resource constraint verification on each solution in the target iteration sequence, screening the feasibility of solutions using the task queue length and the node computing capacity to obtain a feasible solution set, applying a dynamic step size strategy to the feasible solution set, and calculating the convergence degree of adjacent iterative solutions using the Euclidean distance criterion to obtain a converged solution sequence; substituting the converged solution sequence into the objective function, selecting the optimal solution by sorting the objective function values and evaluating the constraint satisfaction, and obtaining an optimal task decomposition scheme; Constructing a state transition probability matrix according to the optimal task decomposition scheme and generating a dynamic task scheduling strategy; The dynamic task scheduling strategy is distributed to each edge computing node, and time series analysis and spatial clustering are performed on multi-source data to obtain a security situation assessment result; specifically, the dynamic task scheduling strategy is sent to multiple edge computing nodes, and parallel processing tasks and execution parameters are set on each edge computing node to obtain a distributed task execution sequence; a time sliding window is set for the multi-source data in the distributed task execution sequence, and the time series change characteristics of various alarm events are extracted to obtain a time series feature sequence; the time series feature sequence is topologically mapped according to the attack chain propagation path, an event propagation network and node association relationships are constructed to obtain a spatial distribution model, and a hierarchical clustering operation is performed on the spatial distribution model to divide the attack event family and the threat scenario mode to obtain an alarm event cluster sequence; the alarm event cluster sequence is input into an association mining module, and the causal chain between events is extracted through deep analysis to obtain a threat situation measurement value, and the threat situation measurement value is combined into a multi-dimensional feature vector, and an overall security risk score is generated through fusion calculation to obtain a security situation assessment result.
2. The edge data processing method of the information security monitoring and early warning AI platform according to claim 1 is characterized in that: The standardization process of the original security warning data collected by the multiple edge computing devices to obtain the standardized feature data includes: Collecting original security alarm data through multiple edge computing devices, and performing outlier detection and redundant data elimination on the alarm type, source IP address, target IP address, port number and alarm level of the original security alarm data to obtain preprocessed data; Performing time sequence alignment on the timestamp information of the preprocessed data to obtain time synchronization data, and performing structured conversion on the alarm characteristic attributes in the time synchronization data to obtain structure standard data; Performing feature classification encoding on the structural standard data, establishing a feature vector template according to the alarm type, alarm level and network address information to obtain a coding feature sequence, and performing correlation analysis on the coding feature sequence to obtain associated feature data; The associated feature data is subjected to data dimension reduction and standardization processing to obtain standardized feature data.
3. The edge data processing method of the information security monitoring and early warning AI platform according to claim 1 is characterized in that: The Lagrangian relaxation problem is subjected to variable separation, and the global optimization problem is decomposed into node-level sub-problems according to the computing power of each edge computing node, to obtain multiple sub-optimization problems, including: Constructing a variable correlation matrix for the Lagrangian relaxation problem, and performing variable correlation analysis based on the coupling relationship between task processing delay and resource consumption to obtain variable grouping results; Extracting the main variable and the subordinate variable based on the variable grouping result, and constructing a variable separation criterion matrix according to the dependency relationship between the variables to obtain a variable separation scheme; Calculate the task processing capability index of the edge node for the variable separation scheme, and construct a node capability scoring vector according to processing delay, resource utilization and load balancing to obtain a node classification result; The global optimization problem is processed in blocks according to the node classification results, and a sub-problem mapping table is constructed based on computing resource capacity and task processing requirements to obtain a problem decomposition framework; Performing constraint allocation on the problem decomposition framework, and establishing constraint allocation rules according to resource sharing relationships and task dependencies between nodes to obtain a sub-problem constraint set; The decomposed optimization problem is bounded based on the sub-problem constraint set, and the sub-problem solution domain is determined according to the computing resource boundary of each node to obtain multiple sub-optimization problems.
4. The edge data processing method of the information security monitoring and early warning AI platform according to claim 1 is characterized in that: The step of constructing a state transition probability matrix according to the optimal task decomposition scheme and generating a dynamic task scheduling strategy includes: The edge node computing load, the remaining resource amount and the task queue length in the optimal task decomposition scheme are constructed into a state description vector to obtain a Markov state space; Prioritizing and allocating resource quotas for tasks in the Markov state space, constructing optional task scheduling action sequences, and obtaining an action candidate space; Applying the scheduling scheme in the action candidate space to the historical task processing data, calculating the state transition statistics and the transition condition probability, and obtaining the state transition matrix; Perform reward modeling on the state transfer matrix, construct an immediate reward function using task delay, resource utilization and load balancing indicators, obtain a state value mapping table, input the state value mapping table into a strategy optimization module, evaluate and update the state-action pair through a value iteration method, and obtain an optimal action strategy; A dynamic programming search is performed on the optimal action strategy to calculate the optimal task scheduling sequence under the current state to obtain a dynamic task scheduling strategy.
5. An edge data processing device for an information security monitoring and early warning AI platform, characterized in that: For implementing the steps of the method according to any one of claims 1 to 4, the device comprises: A standardization module is used to standardize the original security alarm data collected by multiple edge computing devices to obtain standardized feature data; A construction module, configured to construct an objective function including a processing efficiency term and a resource consumption term based on the standardized feature data, and decouple the objective function into a plurality of sub-optimization problems of edge computing nodes using Lagrange multipliers to obtain an initial task allocation scheme; A decomposition module, used for performing local linearization processing and convex optimization decomposition on the initial task allocation scheme to obtain an optimal task decomposition scheme; A generation module, used to construct a state transition probability matrix according to the optimal task decomposition scheme and generate a dynamic task scheduling strategy; The allocation module is used to allocate the dynamic task scheduling strategy to each edge computing node, perform time series analysis and spatial clustering on multi-source data, and obtain security situation assessment results.
6. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Multi-carrier non-orthogonal multiple access-mobile edge computing system resource allocation method
CN116390234A
Dynamic network heterogeneous resource management and task allocation method, system and equipment
CN119172334A