Software terminal big data risk prediction model construction method and system
By constructing implicit differential equations and codebook compression methods on the software terminal, the problems of insufficient fusion of multi-source heterogeneous data and static solidification of model parameters are solved, efficient detection and real-time defense of APT attacks are achieved, and the accuracy and dynamic adaptability of risk prediction are improved.
Patent Information
- Application Number
- CN202510913929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have problems in risk prediction for software terminals, such as insufficient fusion of multi-source heterogeneous data, weak attack path topology modeling capabilities, and static solidification of model parameters, which lead to delayed APT attack detection. In addition, traditional solutions are difficult to achieve efficient and real-time risk prediction and defense when terminal device resources are limited.
By collecting terminal behavior data, performance indicators and security metadata, mapping them to a high-dimensional manifold space, performing topological manifold decomposition, extracting risk-related topological feature vectors, constructing implicit differential equations, optimizing terminal risk propagation dynamics parameters, and performing codebook compression, a lightweight model is generated, and the model parameters are dynamically updated by combining error feedback.
It achieves accurate capture of cross-modal features, improves the detection rate of APT attack chains, enhances the dynamic adaptability of defense strategies, and can respond to new attack threats in real time on terminal devices.
Smart Images

Figure CN120705557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and in particular to a method and system for constructing a software terminal big data risk prediction model. Background Art
[0002] With the widespread adoption of technologies like the Industrial Internet and edge computing, the volume of real-time business data carried by software terminals has surpassed the processing limits of traditional security defense systems. APT attacks spread laterally through covert pathways like multi-stage infiltration and supply chain contamination, exploiting vulnerabilities in communication topologies between terminal devices. Traditional risk prediction models based on single-point log analysis can only capture local anomalies and are limited in their ability to model cross-device and cross-process attack chain behaviors. This can lead to delays of hours or even days in responding to critical vulnerabilities.
[0003] Existing technologies face significant bottlenecks in the deep integration of multi-source heterogeneous data and dynamic risk propagation modeling. On the one hand, differences in the temporal granularity and semantic dimensions of data such as API call sequences, process performance indicators, and network traffic packets make it difficult for traditional feature engineering methods to extract cross-modal correlation features and restore the global behavioral logic of the attack chain. On the other hand, static risk prediction models rely on fixed rules or historical threat feature libraries and lack the ability to track the dynamic evolution of attack paths in real time. When attackers adaptively adjust their attack strategies, the false positive and false negative rates increase dramatically.
[0004] Even more serious is the growing conflict between the resource-constrained nature of terminal devices and the need for real-time defense. The high computing power requirements of traditional models struggle to accommodate the low power consumption and low storage requirements of terminal devices, while cloud-based collaborative solutions, due to transmission latency, cannot meet the millisecond-level risk blocking requirements. This conflict often forces existing solutions to sacrifice model accuracy or response speed in practical deployments, making it difficult to achieve a balance between risk prediction and dynamic defense in complex attack scenarios. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a method and system for constructing a software terminal big data risk prediction model, which solves the problems of insufficient fusion of multi-source heterogeneous data, weak attack path topology modeling capabilities and static solidification of model parameters in existing software terminal risk prediction technology, resulting in delayed APT attack detection.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for constructing a software terminal big data risk prediction model, comprising the following steps:
[0007] S1. Collect terminal behavior data, performance indicators, and security metadata, and perform data cleaning;
[0008] S2, mapping the cleaned data to a high-dimensional manifold space to generate a manifold point set;
[0009] S3. Perform topological manifold decomposition on the manifold point set to extract risk-related topological feature vectors
[0010] S4, based on the topological feature vector Construct implicit differential equations to optimize the terminal risk propagation dynamics parameters;
[0011] S5. Topological feature vector The optimized terminal risk propagation dynamics parameters are used for codebook compression to generate a lightweight model and deploy it to the terminal device;
[0012] S6. Monitor the prediction error and dynamically update the implicit differential equation parameters and codebook compression parameters based on the error feedback.
[0013] Preferably, the step S1 includes:
[0014] S1.1. Collect terminal behavior data, performance metrics, and security metadata. The terminal behavior data includes API call sequences, process creation events, and network connection logs. The performance metrics include CPU usage, memory utilization, and disk I / O rate. The security metadata includes known vulnerability information and system configuration baselines.
[0015] S1.2. For missing values in the time series data of the performance indicators, Lagrange interpolation method is used to fill them;
[0016] S1.3, digitize and normalize the vulnerability chain features in the security metadata and map them to the interval [0,1], as a basis for constructing the risk indication hyperplane in step S3 and finally extracting the topological feature vector It provides the basis for quantifying the geometric risk of terminal states.
[0017] Preferably, the step S2 includes:
[0018] S2.1. Encode the API call sequence using a bidirectional LSTM to capture the bidirectional temporal dependencies of the API calls and generate a temporal feature vector that characterizes the dynamic behavior of the sequence.
[0019] S2.2. Reduce the dimensionality of the performance indicator data to extract key low-dimensional feature vectors while removing redundant information;
[0020] S2.3. Concatenate and fuse the time series feature vector of the API call sequence with the low-dimensional feature vector of the performance indicator to form a set of manifold points in the high-dimensional manifold space, where each manifold point represents the comprehensive state of the terminal within a specific time window.
[0021] Preferably, the step S3 includes:
[0022] S3.1. For the manifold point set, the persistent homology theory is used to calculate the Betti number of the manifold point set at different scales, thereby obtaining the birth and death point pairs (b i ,d i );
[0023] S3.2. Based on the normalized vulnerability chain features from step S1.3, construct one or more risk-indicating hyperplanes, defining the region of the risk-indicating hyperplane associated with known high-risk patterns in the manifold space, thereby providing a geometric basis for subsequent risk quantification of the terminal state.
[0024] S3.3. Using the risk-indicating hyperplane to slice the manifold point set, and combining the topological structures indicated by the birth-death point pairs obtained in step S3.1, identify submanifold regions that have abnormal topological structures and are adjacent to or cross the risk-indicating hyperplane as high-risk candidate regions;
[0025] S3.4. Extract and combine the topological feature vector from the topological invariants corresponding to the birth and death point pairs in the high-risk candidate area and the geometric relationship between the manifold point set and the risk indication hyperplane.
[0026] Preferably, in step S4, the implicit differential equation is:
[0027]
[0028] Among them, t is the time variable, is the topological feature vector, A(t) is the time-varying state transfer matrix, B(t) is the time-varying topological feature weight matrix, y is the state variable, and y=y(t) is the state variable that characterizes the potential risk state of the terminal at time t. is the instantaneous rate of change of the state variable, is the functional form of the implicit differential equation;
[0029] After constructing the implicit differential equation, the terminal risk propagation dynamics parameters are optimized through equation definition, adjoint variable calculation and parameter update based on the adjoint variable λ(t). The adjoint variable λ(t) satisfies:
[0030]
[0031] in, is the time derivative of the adjoint variable, is the partial derivative of function F with respect to the state variable y, is the Jacobian matrix transpose, λ(tN ) is the accompanying variable at the terminal time t N The initial conditions of , L is the loss function.
[0032] Preferably, the step S5 includes:
[0033] S5.1. Topological feature vectors extracted in step S3 Perform sparse binary encoding and convert it into sparse binary codewords with low storage requirements and easy indexing to adapt to the resource limitations of terminal devices;
[0034] S5.2. Based on the terminal risk propagation dynamics parameters optimized in step S4 and the sparse binary codewords encoded in step S5.1, pre-calculate the solution y(t+Δt) of the implicit differential equation in one or more future time steps under different initial risk states and different sparse binary codeword inputs, and construct a hash lookup table, which is the lightweight model.
[0035] Preferably, the step S6 includes:
[0036] S6.1. On the terminal device, use the deployed lightweight model to predict the risk of the real-time collected data and obtain the predicted output y pred (t), and compared with the real risk status y obtained through the side channel or delayed verification true (t) and calculate the prediction error e(t) between the two.
[0037] e(t)=||y pred (t)-y true (t)||, where y pred (t) is the predicted output of the model at time t, y true (t) is the actual risk status at time t;
[0038] S6.2. Set an error threshold γ. If the monitored prediction error e(t)>γ, activate the feedback adjustment mechanism. The feedback adjustment mechanism includes:
[0039] Adjust the parameters of the sparse binary coding described in step S5.1, specifically adjusting the sparsity of the codebook and the threshold of the binary coding to optimize the feature expression accuracy and compression efficiency of the lightweight model;
[0040] The prediction error e(t) is fed back to the cloud or a node with higher computing power to update the loss function L. Based on the updated loss function L and the accompanying variable λ(t), the terminal risk propagation dynamics parameters of the implicit differential equation are re-optimized through equation definition, accompanying variable calculation and parameter update, so that the model can adapt to environmental changes or new threat patterns.
[0041] The present invention also provides a software terminal big data risk prediction model construction system, which is applied to the above-mentioned software terminal big data risk prediction model construction method, comprising:
[0042] Multi-source data collection unit, used to collect terminal behavior data, performance indicators and security metadata, and output structured data through the data cleaning channel;
[0043] A manifold mapping unit, configured to receive the structured data, map it to a high-dimensional manifold space, generate a manifold point set, and output the generated point set;
[0044] a topological decomposition unit, configured to perform topological decomposition on the manifold point set, extract risk-related topological features, and output the extracted features through a feature transmission channel;
[0045] An implicit differential optimization unit constructs an implicit differential equation based on the topological features, optimizes the risk propagation dynamics parameters, and outputs the optimized model parameters;
[0046] A codebook compression and deployment unit receives the topological features and model parameters, performs codebook compression to generate a lightweight model, and transmits the model to the terminal device through a deployment interface;
[0047] The adaptive update unit is connected to the implicit differential optimization unit, the codebook compression and deployment unit and the terminal device, receives the prediction results of the terminal device in real time, calculates the prediction error, and sends a parameter reoptimization instruction to the implicit differential optimization unit if the error exceeds the threshold, and sends a codebook adjustment instruction to the codebook compression and deployment unit.
[0048] Preferably, the parameter reoptimization instruction sent by the adaptive update unit to the implicit differential optimization unit includes:
[0049] Trigger the implicit differential optimization unit to re-execute the adjoint optimization algorithm;
[0050] Carry the current error value e(t) as the optimization weight coefficient;
[0051] The instructions are transmitted via the MQTT protocol.
[0052] Preferably, the codebook adjustment instruction sent by the adaptive updating unit to the codebook compression and deployment unit includes:
[0053] Adjust codebook sparsity parameters;
[0054] Update the quantization accuracy level of the coding matrix;
[0055] The instructions are transmitted via the MQTT protocol.
[0056] The present invention provides a method and system for constructing a software terminal big data risk prediction model. This method has the following beneficial effects:
[0057] 1. This invention uses a hybrid architecture that uses bidirectional LSTM to encode API call sequences and PCA to reduce dimensionality and performance indicators, accurately capturing the temporal dependencies and statistical correlations of cross-modal features. Traditional solutions process APIs and performance indicators separately, resulting in feature redundancy. This invention improves threat identification accuracy through manifold space fusion, directly solving the information loss problem of multi-source data fragmentation modeling in existing technologies.
[0058] 2. This invention extracts topological birth-death point pairs based on persistent homology groups and combines them with hyperplane cutting of vulnerability chains to identify potential attack paths from a geometric perspective. Compared with traditional statistical methods that only analyze local feature distributions, this invention achieves the first precise analysis of risk propagation topological genes, improves the detection rate of APT attack chains, and completely breaks through the bottleneck of existing solutions in modeling complex attack paths.
[0059] 3. The present invention realizes real-time adaptive updating of model parameters through dual online collaborative tuning of implicit differential equation parameters and codebook compression strategy. Existing static models are difficult to cope with new attack threats due to the lack of dynamic feedback mechanism. The present invention is based on error feedback closed-loop driven parameter self-optimization, which significantly improves the dynamic adaptability of the defense strategy and effectively suppresses the spread and evolution of potential attack paths. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flow chart of the method of the present invention;
[0061] Figure 2 It is a process framework diagram of the system of the present invention. DETAILED DESCRIPTION
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0063] Please see the attached Figure 1 The embodiment of the present invention provides a method for constructing a software terminal big data risk prediction model, comprising the following steps:
[0064] S1. Collect terminal behavior data, performance indicators, and security metadata, and perform data cleaning;
[0065] The data collection and cleaning steps of this invention utilize a multi-source heterogeneous data processing method to achieve standardized preprocessing of terminal behavior data, performance indicators, and security metadata. This method includes two core stages: missing value completion and feature normalization. The specific implementation is as follows:
[0066] Completion of missing values of performance indicators:
[0067] The data collection end includes a log collection agent and a performance monitoring probe. The log collection agent captures the API call sequence S = {s1, s2, ..., s T}, the performance monitoring probe periodically collects CPU / memory / disk indicators M={m1,m2,...,m T When data packet loss or sampling interruption occurs, the performance index value at the missing time t′ is supplemented by Lagrange interpolation:
[0068]
[0069] Where, t i are adjacent valid sampling moments, is the index value at the corresponding moment, N is the number of adjacent valid data points used for interpolation calculation, t i and t j is the valid sampling time adjacent to the missing time t′ (i, j is the time index), is the Lagrange interpolation basis function. For example, when the CPU occupancy data is missing at t′=t2, the data at time t1 and t3 are selected for first-order interpolation:
[0070]
[0071] This algorithm can restore the continuity of time series data and ensure that the interpolation results are within the distribution range of the original data.
[0072] Normalization of vulnerability chain features:
[0073] The security metadata processing end connects to the vulnerability feature database and extracts the vulnerability chain feature set C = {c1, c2, ..., c K}, where each feature Contains vulnerability type, attack path weight and impact range parameters. Normalize C:
[0074]
[0075] Where, is the normalized vulnerability chain feature vector, c i is the original vulnerability chain feature, μ C The vulnerability chain feature set C = {c1, c2, ..., c K}, σC is the standard deviation vector of the vulnerability chain feature set C, and K is the total number of vulnerability chain features.
[0076] Normalized features Obeying the standard normal distribution can eliminate the dimensional differences between different vulnerability dimensions and provide consistent input for subsequent topological decomposition.
[0077] S2, mapping the cleaned data to a high-dimensional manifold space to generate a manifold point set;
[0078] The manifold mapping step of the present invention maps cleaned terminal behavior data, performance indicators, and security metadata into a high-dimensional manifold space using a heterogeneous feature fusion method, generating a manifold point set with a topological structure. The method includes three core stages: temporal feature encoding, performance indicator dimensionality reduction, and feature splicing. The specific implementation is as follows:
[0079] Bidirectional LSTM temporal feature encoding:
[0080] The input data is the API call sequence S = {s1, s2, ..., s T},in Represents the API call features at time t (such as call type, parameter value, execution time). The sequence is encoded using a bidirectional LSTM network:
[0081]
[0082] Where T is the sliding window length, h t is the time series feature vector at time t, and d1 is the LSTM hidden layer dimension. For example, when T = 10, n = 128, and d1 = 64, the network structure includes:
[0083] Input layer: 128-dimensional API feature vector;
[0084] Bidirectional LSTM layer: 32 hidden units in each direction, and the output concatenation is 64-dimensional;
[0085] Dropout layer: probability 0.2 to prevent overfitting.
[0086] This encoding process can capture the dependencies between API call sequences and generate feature vectors containing long-term temporal patterns.
[0087] Performance indicators PCA dimensionality reduction:
[0088] The input data is the performance index M={m1,m2,...,m T},in Including CPU usage, memory usage and disk I / O rate. tPerform dimensionality reduction:
[0089]
[0090] Where d2 is the dimension after dimensionality reduction, which is preferably determined by the cumulative contribution rate of eigenvalues. For example, when the original dimension is 3, the first two principal components are retained (d2 = 2), and the cumulative variance contribution rate is more than 90%. The low-dimensional eigenvector p after dimensionality reduction is t satisfy:
[0091] p t =V T (m t -μ M );
[0092] Where V is the PCA transformation matrix, μ M is the mean vector of performance indicators. This step can eliminate redundancy between indicators and reduce the risk of dimensional disaster in subsequent calculations.
[0093] Heterogeneous feature splicing and manifold construction:
[0094] The input data is the time series feature h t With the low-dimensional feature vector p t , generate a manifold point set through vector splicing operation
[0095]
[0096] Where, Represents a vector concatenation operation. For example, when d1=64 and d2=2, Manifold point set The geometry of is defined by the following elements:
[0097] Distance metric: Mahalanobis distance is used to calculate the similarity between points;
[0098] Local neighborhood: Construct an adjacency graph using the k-nearest neighbor (k-NN) algorithm, with k=5.
[0099] This manifold space can fuse API behavior timing patterns with dynamic changes in performance indicators, providing a high-dimensional geometric representation for topological decomposition.
[0100] S3. Perform topological manifold decomposition on the manifold point set to extract risk-related topological features;
[0101] The data collection and cleaning steps of this invention utilize a multi-source heterogeneous data processing method to achieve standardized preprocessing of terminal behavior data, performance indicators, and security metadata. This method includes two core stages: missing value completion and feature normalization. The specific implementation is as follows:
[0102] Continuous homology calculation and extraction of birth and death point pairs: the input data is the manifold point set generated in step S2 where x t is a manifold point, and d is the manifold space dimension.
[0103] Constructing a filter complex sequence: for example, constructing a Vietoris-Rips (VR) complex sequence based on a manifold point set Or Alpha complex sequence, where the filter parameter ∈ represents the scale.
[0104] Compute persistent homology groups and Betti numbers: Compute homology groups of various orders of filtered complex sequences at different scales ∈ And get the Betti numbers of various orders from it For example, β0 represents the number of connected branches, β1 represents the number of one-dimensional loops, and β2 represents the number of two-dimensional voids.
[0105] Extract birth-death point pairs: By tracking the changes in Betti numbers at different scales, the birth scale b of each topological structure (such as connected branches, rings, and cavities) is identified. i and death scale d i . Each such (b i ,d i ) is a birth-death point pair, and its persistence (persistence) p i =d i -b i It reflects the stability and significance of the corresponding topological structure.
[0106] For example, the birth and death point pair (b j ,d j ) corresponds to a ring structure in the manifold, if its duration p j A longer length indicates that the ring structure exists stably over a larger scale range, which may indicate some kind of cyclic pattern or association.
[0107] For example, when calculating the one-dimensional persistent homology group H1(∈) of a manifold point set, if a birth-death point pair (0.3, 0.8) is found, this indicates that a ring-like topological structure appears at scale ∈ = 0.3 and disappears at scale ∈ = 0.8, with a duration of 0.5. This persistent ring structure may be associated with potential risk transmission paths or abnormal behavior patterns.
[0108] Risk indication hyperplane construction: The input data is the features obtained after digitizing and normalizing the security metadata (such as vulnerability chain features) in step S1.3.
[0109] Construction goal: Construct one or more risk-indicating hyperplanes in the d-dimensional manifold space These hyperplanes are intended to geometrically bound regions associated with known high-risk patterns (indicated by security metadata).
[0110] Construction method:
[0111] If the security metadata features can be directly mapped to some dimensions of the manifold space or can be used to label the manifold points x i The risk level of , then supervised learning methods (such as support vector machine SVM) can be used to construct the hyperplane. For example, based on the manifold point x i and its associated normalized vulnerability chain features, giving each x i A risk label (such as high risk / low risk). Then, using the idea of the maximum margin classifier, we can find the optimal weight vector w. k and the bias term b k , so that the hyperplane can effectively distinguish the manifold regions associated with high risks.
[0112] Alternatively, if some normalized vulnerability chain feature v j ∈[0,1] (from step S1.3) can be directly interpreted as a risk indicator for a specific direction in the manifold space, and the hyperplane can be constructed.
[0113] For example, if a vulnerability severity feature v s Corresponding to a specific direction u of the manifold s , then we can set the hyperplane to where θ s It is a v-based s threshold.
[0114] The constructed risk-indicating hyperplane provides a key geometric reference for identifying high-risk candidate regions in the subsequent step S3.3 and quantifying the geometric risk of the terminal state in step S3.4.
[0115] Example: Assuming the manifold is a two-dimensional space, by analyzing the manifold points associated with high-risk vulnerability chains, we obtain a risk-indicating hyperplane equation: 0.8x1+0.6x2-0.5=0. This hyperplane divides the manifold space into two parts, one of which (for example, 0.8x1+0.6x2-0.5>0) is preliminarily considered to be more relevant to high-risk patterns.
[0116] Identification of high-risk candidate areas: The input data is the manifold point set generated in step S2 The birth-death point pairs extracted in step S3.1 and the risk indicator hyperplane H constructed in step S3.2 k .
[0117] Regional definition and screening:
[0118] Using one or more risk-indicating hyperplanes H k Convection manifold point set Perform preliminary geometric division.
[0119] For example, for the hyperplane H:w T x+b=0, we can satisfy w T x+b>θ risk (where θ risk The manifold point set area (which is a preset or dynamically adjusted threshold) is regarded as a potential high-risk side.
[0120] Within these preliminarily defined regions, the “abnormal” or “significant” topological structures (such as long-lasting rings, cavities, or densely connected branches) indicated by the birth-death point pairs with longer duration obtained in step S3.1 are further identified.
[0121] Identify high-risk candidate areas: The high-risk candidate areas finally identified are those submanifold regions or point sets that satisfy the following conditions simultaneously:
[0122] Located on the “high-risk side” indicated by one or more risk-indicating hyperplanes, or adjacent to / crossing these hyperplanes.
[0123] Furthermore, these regions themselves exhibit significant or unusual topological structures indicated by long-duration birth-death point pairs.
[0124] Example: Suppose that within the high-risk region bounded by the risk-indicating hyperplane, a persistent homology analysis reveals the presence of a particularly long-duration β1 loop (indicated by a birth-death point pair (b, d) with a large db). The manifold points that constitute this loop and its adjacent region are identified as a high-risk candidate region.
[0125] Topological feature vector extraction: The input data is the high-risk candidate area identified in step S3.3
[0126] Extracting topological invariants: From each high-risk candidate region Extract quantitative topological invariants from the internal and related birth-death point pairs. These invariants can include:
[0127] The Betti numbers of a certain order are calculated for the region (e.g., the number of rings, the number of cavities within the region).
[0128] The duration length and birth / death scale value of the birth / death point pairs that constitute the significant topological structure of the area.
[0129] Statistical properties of the persistence diagram, such as persistence entropy and coordinates of key point pairs.
[0130] Extracting geometric relationship features: quantifying high-risk candidate areas The geometric relationship between the risk indicator hyperplane and the risk indicator hyperplane. This can include:
[0131] The average distance, minimum / maximum distance, and distance variance of points in the region to the hyperplane.
[0132] The distribution or density of points in the region on both sides of the hyperplane.
[0133] The angle between the principal direction of the region (as computed by PCA within the region) and the normal vector of the hyperplane.
[0134] The intersection pattern between the region and the hyperplane (such as the volume, area, or number of points in the intersection).
[0135] Combining to generate topological feature vectors Combine the extracted topological invariants and geometric relationship features (for example, by concatenation, weighted summation or more complex function mapping) to form a comprehensive topological feature vector
[0136] The topological eigenvector The geometric risk of the terminal state (represented by a set of manifold points) is quantified and serves as the key input to the subsequent step S4 (constructing implicit differential equations).
[0137] S4. Construct implicit differential equations based on topological features to optimize the terminal risk propagation dynamics parameters;
[0138] The implicit differential equation construction and parameter optimization steps of the present invention achieve adaptive modeling of terminal risk propagation dynamics through the adjoint method. The method includes three stages: equation definition, adjoint variable calculation, and parameter update. The specific implementation method is as follows:
[0139] Implicit differential equation construction:
[0140] Input data is topological features and state variables Construct the implicit differential equation:
[0141]
[0142] Where:
[0143] Time-varying state transfer matrix, describing the nonlinear coupling relationship between risk states;
[0144] Time-varying topological feature weight matrix, quantifying the contribution of topological features to risk propagation;
[0145] The instantaneous rate of change of the state variable is calculated by numerical differentiation methods (such as central difference);
[0146] t is the time variable, is the topological eigenvector, y is the state variable, is the functional form of the implicit differential equation.
[0147] For example, when n=2 and m=3, the equation is expanded as follows:
[0148]
[0149] The equation can capture the nonlinear dynamic characteristics of the risk state and adapt to the dynamic environment through time-varying parameters.
[0150] Backpropagation with variables:
[0151] Accompanying variables Satisfy the inverse differential equation and terminal conditions:
[0152]
[0153] Partial derivative calculation:
[0154]
[0155] In the formula, the terminal condition λ(t N ) by the loss function gives:
[0156] λ(t N )=2(y(t N )-y true (t N ));
[0157] Where, is the time derivative of the adjoint variable, is the partial derivative of function F with respect to the state variable y, is the transpose of the Jacobian matrix, y true (t N ) is the terminal time t N The actual observed value.
[0158] The process calculates the gradient of parameters A(t) and B(t) by solving the adjoint equation in reverse.
[0159] Parameter optimization and update:
[0160] Gradient calculation: Use the accompanying variable λ(t) to calculate the gradient of the loss function with respect to the parameters:
[0161]
[0162] Parameter update rule: Adopt gradient descent method to iteratively optimize:
[0163]
[0164] Where η is the learning rate, which is preferably adjusted by an adaptive learning rate algorithm (such as Adam).
[0165] S5. Compress the topological features with codebook, generate a lightweight model and deploy it to the terminal device;
[0166] The codebook compression and deployment steps of the present invention achieve lightweight and real-time inference of risk models through sparse coding and pre-computed hashing techniques. The method includes two core stages: feature sparse coding and differential equation pre-computation. The specific implementation is as follows:
[0167] Sparse binary coding of topological features:
[0168] Input data is topological features Through the encoding matrix W e ∈{0,1} m×d Perform sparse binary compression:
[0169]
[0170] Where W e Each column must contain only k non-zero elements (sparseness constraint), and the maximum response dimension is selected by a greedy algorithm. Preferably, the sparsity parameter k is dynamically adjusted according to the terminal computing power. When m = 64, d = 16, k = 3, the non-zero position of each column of the encoding matrix is:
[0171] W e [:,1]=[1,0,…,1,0] T (non-zero indexes are 2, 5, 9);
[0172] This encoding process can compress feature dimensions to 25% of their original size while preserving key topological patterns.
[0173] Differential equation pre-calculation and hash table construction:
[0174] The input data is the optimized parameters A(t), B(t), the solution space of the implicit differential equation F=0 is pre-sampled, and a hash lookup table is generated. The key-value design is as follows:
[0175] Key: Sparsely coded With the state variable range y min ≤y≤y max Joint generation;
[0176] Value: The quantized result of the pre-calculated solution y(t+Δt)=Solve(F=0), stored as an 8-bit fixed-point number.
[0177] The hash function is defined as:
[0178]
[0179] Where, is the random projection matrix, s is the binning scale, H size is the hash table capacity. For example, when d=16, n=2, and l=8, the hash collision rate is less than 1%.
[0180] S6. Monitor the prediction error and dynamically update the implicit differential equation parameters and codebook compression parameters based on the error feedback.
[0181] The dynamic update step of the present invention achieves closed-loop optimization of model parameters and compression strategies through a prediction error feedback mechanism. The method includes two core stages: error monitoring and parameter adjustment. The specific implementation is as follows:
[0182] Forecast error calculation and judgment:
[0183] Input data is the model prediction output and the true risk status Calculate the instantaneous error:
[0184]
[0185] The accumulated error is updated by the sliding window mean:
[0186] (T is the window length);
[0187] when Parameter update is triggered when τ e is the preset error threshold. For example, when T=10, τ e = 0.1, if three consecutive windows Determine if parameters need to be adjusted.
[0188] Dynamic adjustment of codebook sparsity:
[0189] The input parameter is the current sparsity According to the error change rate Adjustment:
[0190] k new =Clip(k+Sign(Δe)·δk ,k min ,k max );
[0191] Where, δ k is the step size, Clip(·) is the range cutoff function, k new is the new sparsity parameter after dynamic adjustment, Sign(Δe) is the sign function, Δe is the error change rate, k min and k max are the minimum and maximum values of the sparsity parameter, respectively. For example, when k=3, δ k =1, Δe=0.05, update to k=4. Preferably, k min =1, k max =5, to prevent over-compression.
[0192] Implicit differential equation parameter update:
[0193] Input data is cumulative error Adjust the learning rate η through proportional-integral control:
[0194]
[0195] Where α is the attenuation coefficient, and its exemplary value is α = 0.5. The updated learning rate is used for parameter optimization in step S4:
[0196]
[0197] in, is the gradient matrix of the loss function L to the state transfer matrix A(t), It is the gradient matrix of the loss function L with respect to the topological feature weight matrix B(t).
[0198] The software terminal big data risk prediction model construction system described below and the software terminal big data risk prediction model construction method described above can be referenced to each other.
[0199] Please see the attached Figure 2 A software terminal big data risk prediction model construction system is applied to the above-mentioned software terminal big data risk prediction model construction method, comprising:
[0200] Multi-source data collection unit, used to collect terminal behavior data, performance indicators and security metadata, and output structured data through the data cleaning channel;
[0201] Manifold mapping unit, used to receive structured data, map it to a high-dimensional manifold space, generate a manifold point set and output it;
[0202] Topological decomposition unit, used to perform topological decomposition on the manifold point set, extract risk-related topological features, and output them through the feature transmission channel;
[0203] Implicit differential optimization unit, which constructs implicit differential equations based on topological features, optimizes risk propagation dynamics parameters, and outputs optimized model parameters;
[0204] The codebook compression and deployment unit receives topology features and model parameters, compresses the codebook to generate a lightweight model, and transmits it to the terminal device through the deployment interface;
[0205] The adaptive update unit is connected to the implicit differential optimization unit, the codebook compression and deployment unit and the terminal device, receives the prediction results of the terminal device in real time, calculates the prediction error, and sends a parameter reoptimization instruction to the implicit differential optimization unit if the error exceeds the threshold, and sends a codebook adjustment instruction to the codebook compression and deployment unit.
[0206] The parameter reoptimization instructions sent by the adaptive update unit to the implicit differential optimization unit include:
[0207] Trigger the implicit differential optimization unit to re-execute the adjoint optimization algorithm;
[0208] Carry the current error value e(t) as the optimization weight coefficient;
[0209] The instructions are transmitted via the MQTT protocol.
[0210] The codebook adjustment instruction sent by the adaptive updating unit to the codebook compression and deployment unit includes:
[0211] Adjust codebook sparsity parameters;
[0212] Update the quantization accuracy level of the coding matrix;
[0213] The instructions are transmitted via the MQTT protocol.
[0214] The system of this embodiment can be used to execute the above method embodiments, and its principles and technical effects are similar, so they will not be repeated here.
[0215] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a software terminal big data risk prediction model, characterized in that: The following steps are involved: S1. Collect terminal behavior data, performance indicators, and security metadata, and perform data cleaning; S2, mapping the cleaned data to a high-dimensional manifold space to generate a manifold point set; S3. Perform topological manifold decomposition on the manifold point set to extract risk-related topological feature vectors S4, based on the topological feature vector Construct implicit differential equations to optimize the terminal risk propagation dynamics parameters; S5. Topological feature vector The optimized terminal risk propagation dynamics parameters are used for codebook compression to generate a lightweight model and deploy it to the terminal device; S6. Monitor the prediction error and dynamically update the implicit differential equation parameters and codebook compression parameters based on the error feedback.
2. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: The steps of S1 include: S1.
1. Collect terminal behavior data, performance metrics, and security metadata. The terminal behavior data includes API call sequences, process creation events, and network connection logs. The performance metrics include CPU usage, memory utilization, and disk I / O rate. The security metadata includes known vulnerability information and system configuration baselines. S1.
2. For missing values in the time series data of the performance indicators, Lagrange interpolation method is used to fill them; S1.3, digitize and normalize the vulnerability chain features in the security metadata and map them to the interval [0,1], as a basis for constructing the risk indication hyperplane in step S3 and finally extracting the topological feature vector It provides the basis for quantifying the geometric risk of terminal states.
3. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: The steps of S2 include: S2.
1. Encode the API call sequence using a bidirectional LSTM to capture the bidirectional temporal dependencies of the API calls and generate a temporal feature vector that characterizes the dynamic behavior of the sequence. S2.
2. Dimensionality reduction is performed on the data of the performance indicators to extract key low-dimensional feature vectors while removing redundant information. S2.
3. Concatenate and fuse the time series feature vector of the API call sequence with the low-dimensional feature vector of the performance indicator to form a set of manifold points in the high-dimensional manifold space, where each manifold point represents the comprehensive state of the terminal within a specific time window.
4. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: The steps of S3 include: S3.
1. For the manifold point set, the persistent homology theory is used to calculate the Betti number of the manifold point set at different scales, thereby obtaining the birth and death point pairs (b i ,d i ); S3.
2. Based on the normalized vulnerability chain features from step S1.3, construct one or more risk-indicating hyperplanes, defining the region of the risk-indicating hyperplane associated with known high-risk patterns in the manifold space, thereby providing a geometric basis for subsequent risk quantification of the terminal state. S3.
3. Using the risk-indicating hyperplane to slice the manifold point set, and combining the topological structures indicated by the birth-death point pairs obtained in step S3.1, identify submanifold regions that have abnormal topological structures and are adjacent to or cross the risk-indicating hyperplane as high-risk candidate regions; S3.
4. Extract and combine the topological feature vector from the topological invariants corresponding to the birth and death point pairs in the high-risk candidate area and the geometric relationship between the manifold point set and the risk indication hyperplane.
5. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: In step S4, the implicit differential equation is: Where t is the time variable, is the topological feature vector, A(t) is the time-varying state transfer matrix, B(t) is the time-varying topological feature weight matrix, y is the state variable, and y=y(t) is the state variable that characterizes the potential risk state of the terminal at time t. is the instantaneous rate of change of the state variable, is the functional form of the implicit differential equation; After constructing the implicit differential equation, the terminal risk propagation dynamics parameters are optimized through equation definition, adjoint variable calculation and parameter update based on the adjoint variable λ(t). The adjoint variable λ(t) satisfies: in, is the time derivative of the adjoint variable, is the partial derivative of function F with respect to the state variable y, is the Jacobian matrix transpose, λ(t ) ) is the accompanying variable at the terminal time t ) The initial conditions of , L is the loss function.
6. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: The steps of S5 include: S5.
1. Topological feature vectors extracted in step S3 Perform sparse binary encoding and convert it into sparse binary codewords with low storage requirements and easy indexing to adapt to the resource limitations of terminal devices; S5.
2. Based on the terminal risk propagation dynamics parameters optimized in step S4 and the sparse binary codewords encoded in step S5.1, pre-calculate the solution y(t+Δt) of the implicit differential equation in one or more future time steps under different initial risk states and different sparse binary codeword inputs, and construct a hash lookup table, which is the lightweight model.
7. The method for constructing a software terminal big data risk prediction model according to claim 1, characterized in that: The steps of S6 include: S6.
1. On the terminal device, use the deployed lightweight model to predict the risk of the real-time collected data and obtain the prediction output y pred (t), and compared with the real risk status y obtained through the side channel or delayed verification true (t) and calculate the prediction error e(t) between the two. e(t)=||y pred (t)-y true (t)||, where y pred (t) is the predicted output of the model at time t, y true (t) is the actual risk status at time t; S6.
2. Set an error threshold γ. If the monitored prediction error e(t)>γ, activate the feedback adjustment mechanism. The feedback adjustment mechanism includes: Adjust the parameters of the sparse binary coding described in step S5.1, specifically adjusting the sparsity of the codebook and the threshold of the binary coding to optimize the feature expression accuracy and compression efficiency of the lightweight model; The prediction error e(t) is fed back to the cloud or a node with higher computing power to update the loss function L. Based on the updated loss function L and the accompanying variable λ(t), the terminal risk propagation dynamics parameters of the implicit differential equation are re-optimized through equation definition, accompanying variable calculation and parameter update, so that the model can adapt to environmental changes or new threat patterns.
8. A software terminal big data risk prediction model construction system, characterized by: A method for constructing a software terminal big data risk prediction model as described in any one of claims 1 to 7 above, comprising: Multi-source data collection unit, used to collect terminal behavior data, performance indicators and security metadata, and output structured data through the data cleaning channel; A manifold mapping unit, configured to receive the structured data, map it to a high-dimensional manifold space, generate a manifold point set, and output the generated point set; a topological decomposition unit, configured to perform topological decomposition on the manifold point set, extract risk-related topological features, and output the extracted features through a feature transmission channel; An implicit differential optimization unit constructs an implicit differential equation based on the topological features, optimizes the risk propagation dynamics parameters, and outputs the optimized model parameters; A codebook compression and deployment unit receives the topological features and model parameters, performs codebook compression to generate a lightweight model, and transmits the model to the terminal device through a deployment interface; The adaptive update unit is connected to the implicit differential optimization unit, the codebook compression and deployment unit and the terminal device, receives the prediction results of the terminal device in real time, calculates the prediction error, and sends a parameter reoptimization instruction to the implicit differential optimization unit if the error exceeds the threshold, and sends a codebook adjustment instruction to the codebook compression and deployment unit.
9. A software terminal big data risk prediction model construction system according to claim 8, characterized in that: The parameter reoptimization instruction sent by the adaptive update unit to the implicit differential optimization unit includes: Trigger the implicit differential optimization unit to re-execute the adjoint optimization algorithm; Carry the current error value e(t) as the optimization weight coefficient; The instructions are transmitted via the MQTT protocol.
10. A software terminal big data risk prediction model construction system according to claim 8, characterized in that: The codebook adjustment instruction sent by the adaptive updating unit to the codebook compression and deployment unit includes: Adjust codebook sparsity parameters; Update the quantization accuracy level of the coding matrix; The instructions are transmitted via the MQTT protocol.