Small program abnormal behavior management system fused with artificial intelligence

By designing a mini program abnormal behavior management system that integrates artificial intelligence, it solves the problem that existing systems are difficult to comprehensively monitor and identify abnormal behaviors, and realizes intelligent monitoring and control of abnormal behaviors of mini program, improving the refinement and efficiency of abnormal management.

CN120012078AActive Publication Date: 2025-05-16青岛深空软件科技有限公司

Patent Information

Application Number
CN202510494719.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing mini-program abnormal behavior management system is difficult to comprehensively monitor and identify abnormal behaviors, lack adaptability and intelligence, and cannot effectively deal with the evolution and spread risks of abnormal behaviors.

Method used

A mini-program abnormal behavior management system integrating artificial intelligence is designed to collect multi-dimensional data in real time through the data acquisition module, the behavior modeling module builds a dynamic behavior baseline model, the preliminary matching module generates real-time abnormal confidence coefficient, the graph construction module builds an abnormal propagation probability graph, and the comprehensive management and control module calculates the comprehensive risk entropy value and triggers a multi-level management and control strategy.

Benefits of technology

It realizes comprehensive monitoring and intelligent control of abnormal behavior of mini programs, can accurately identify abnormal deviations, deeply explore the internal correlation of abnormal behavior, evaluate the range of risk transmission, and dynamically adjust the management and control strategies, improving the refinement and efficiency of abnormal management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012078A_ABST
    Figure CN120012078A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of applets, and discloses an applet abnormal behavior management system fused with artificial intelligence. Comprising the steps of collecting a multi-dimensional data stream in real time, performing space-time slicing on the multi-dimensional data stream, extracting a behavior feature vector in each time window, and constructing a dynamic behavior baseline model bound with a user identity; calculating a deviation degree between the behavior feature vector and a dynamic behavior baseline model, and generating a real-time anomaly confidence coefficient in combination with a similarity matching degree with a preset historical anomaly event library; the hidden relevance of a cross-interface calling link is recognized, then an abnormal propagation probability graph is constructed, the weights of the nodes in the abnormal propagation probability graph are dynamically corrected through a real-time abnormal confidence coefficient, and the corrected weights of the nodes in the abnormal propagation probability graph are obtained; and calculating a comprehensive risk entropy value, and when the comprehensive risk entropy value exceeds a risk threshold value, dynamically triggering a multi-level management and control strategy, so that complex applet abnormal behaviors can be comprehensively detected and managed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mini-program technology, and more specifically, to a mini-program abnormal behavior management system integrating artificial intelligence. Background Art

[0002] With the rapid development of mobile Internet, mini programs have gradually emerged as a new form of application and have been widely used. Mini programs do not require installation and are within easy reach, providing users with an excellent user experience. However, with the popularization of mini programs, their security issues have become increasingly prominent. The openness of mini programs exposes them to various potential security threats, such as malicious code injection, sensitive information theft, resource abuse, cheating, etc. Once a mini program is attacked or exhibits abnormal behavior, it will not only affect the user experience, but may also lead to serious consequences such as privacy leakage and economic losses.

[0003] However, existing technologies often only focus on data of a single dimension, such as user operation events or API call frequency, and it is difficult to fully grasp the multiple aspects of abnormal behavior, which limits the accuracy and coverage of anomaly detection. Taking an e-commerce mini program as an example, if only user click behavior is monitored while data such as network requests and resource consumption are ignored, some potential anomalies such as order-brushing behavior or program vulnerabilities may be missed. Secondly, traditional methods are mostly based on static rules or thresholds for anomaly judgment, lacking adaptability and intelligence, and cannot effectively respond to the continuous evolution of abnormal behavior. Furthermore, it is difficult for existing systems to explore and utilize the inherent correlation between abnormal behaviors, and it is impossible to comprehensively assess the risk of anomaly propagation and the scope of impact. In addition, after discovering anomalies, traditional systems usually take a single processing measure and lack refined, multi-level anomaly control strategies. Finally, existing systems lack the ability to trace the causes and root causes of anomalies, making it difficult to accurately locate abnormal code areas and take targeted repair measures. If the specific occurrence path of the anomaly cannot be traced, it is difficult to efficiently repair related vulnerabilities or defects.

[0004] In view of this, the present invention proposes a small program abnormal behavior management system integrating artificial intelligence to solve the above problems. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solutions: A mini-program abnormal behavior management system integrating artificial intelligence, comprising: a data acquisition module, for real-time acquisition of multi-dimensional data streams during the operation of the mini-program, wherein the multi-dimensional data streams include user operation event sequences, API call frequency matrices, resource consumption fluctuation curves, and network request response graphs; A behavior modeling module, used to slice the multi-dimensional data stream into time and space according to preset time windows, extract the behavior feature vectors in each time window, and build a dynamic behavior baseline model bound to the user identity; A preliminary matching module is used to calculate the deviation between the behavior feature vector and the dynamic behavior baseline model, and generate a real-time anomaly confidence coefficient by combining the similarity matching degree with a preset historical anomaly event library; A graph construction module is used to identify the hidden correlation of cross-interface call links according to the topological structure changes of the network request response graph, and then construct an abnormal propagation probability graph, and dynamically correct the weights of the nodes in the abnormal propagation probability graph through the real-time abnormal confidence coefficient to obtain the corrected weights of the nodes in the abnormal propagation probability graph; The comprehensive management and control module is used to integrate the deviation, real-time anomaly confidence coefficient and the modified weights of the nodes in the anomaly propagation probability graph, calculate the comprehensive risk entropy value, preset the risk threshold, and dynamically trigger the multi-level management and control strategy when the comprehensive risk entropy value exceeds the risk threshold: locate the code hot zone and inject behavior interception probes according to the business module where the abnormal peak in the API call frequency matrix is ​​located, and generate an encrypted audit log containing the resource consumption fluctuation traceability path; each module is connected by wired and / or wireless means.

[0006] Furthermore, the method for acquiring the user operation event sequence includes: Inject a lightweight event capture hook into the mini-program framework layer to obtain comprehensive data of user operation events; the comprehensive data includes type, occurrence timestamp, event source component ID, and event parameters, and serialize these comprehensive data into binary format or compressed format to obtain a user event information sequence; In the view layer of the mini program, a front-end framework that supports virtual DOM trees is used. After the view layer is rendered, the new virtual DOM tree is compared with the old virtual DOM tree to obtain only the changed DOM node paths, and these DOM node paths are associated with the corresponding user event information sequences. These changed DOM node paths and the corresponding user event information sequences are serialized and uploaded to obtain the user operation event sequence.

[0007] Furthermore, the method for obtaining the API call frequency matrix includes: An API call monitoring agent is implanted in the mini-program runtime environment. By modifying the context object of the mini-program runtime, all API methods that need to be monitored are rewritten. In the rewritten API method, the timestamp and call parameters of the API call are first recorded; then the logic of the original API method is executed, and the return result is recorded; a unique ID is assigned to each API, and the API call timestamp, call parameters, and return result are serialized into a binary format to obtain a call record; using the API ID as the key, the corresponding call record is stored in the ring buffer in the memory; Maintain a configuration table to store the importance score of each API and dynamically calculate a sampling rate benchmark value for each API ;in, and is the weight parameter; is the current CPU usage, is the current memory usage; Preset an importance threshold. For APIs with an importance score greater than the importance threshold, multiply their sampling rate base value by a random number in the interval (1, 2] to obtain the corresponding sampling rate. For APIs with an importance score less than or equal to the importance threshold, use the sampling rate base value as the corresponding sampling rate. Generate a random floating point number between 0 and 1. If the sampling rate corresponding to the API is greater than or equal to the random floating point number, record the call record of this API. If the sampling rate corresponding to the API is less than the random floating point number, do not record the call record of this API and directly execute the API logic. With a fixed time window as the period, the ring buffer in the memory is scanned, the number of calls of each API in the time window is counted, and the API call frequency matrix is ​​constructed.

[0008] Furthermore, the network request response graph is obtained in the following manner: By modifying the context object of the network library when the applet is running, rewriting all methods of sending network requests and receiving responses, in the rewritten methods, intercepting the original data packets of each HTTP / HTTPS request and response, recording them as request packets and response packets; parsing the intercepted request packets, identifying the basic information of the request, including URL path, query parameters, and request body; Pre-train a semantic classification model, use the semantic classification model to classify the basic information of the request, and mark the type of the request; for each request data packet, construct the basic information into a request vector; for each response data packet, extract the response status code and response body, and construct them into a response vector; Use the pre-trained BERT semantic model to embed the request vector and response vector, map each request vector / response vector to an m-dimensional semantic vector, calculate the cosine similarity between the request vector / response vector and the semantic vector as the corresponding semantic score, i.e., the request semantic score and the response semantic score; set the request semantic score threshold t1 and the response semantic score threshold t2, if the request semantic score is greater than or equal to t1, create a request node, if the response semantic score is greater than or equal to t2, create a response node, the node attributes of the response node and the request node include the request URL, request body hash, response status code and the corresponding semantic vector; if the initiation time of the request node A is before the arrival time of the response node B, and the URL of A is the same as the URL of B, then create a directed edge from A to B, and the attributes of the directed edge include the response delay; calculate the cosine similarity between the semantic vectors of the request node A and the response node B, recorded as the semantic similarity; set the relevance threshold t3, if the semantic similarity is greater than or equal to t3, then create a directed edge from A to B, and the attributes of the directed edge include this cosine similarity; Set the path compression threshold t4 and the similarity compression threshold t5. If the path length from request node A to response node B is greater than t4, and the product of the semantic similarity from request node to response node in the path length is greater than t5, then directly create a dependency edge from A to B and mark it as a compressed edge, and obtain the network request response graph.

[0009] Furthermore, the behavior feature vector is extracted in the following manner: Perform sliding window statistics on the user operation event sequence within the time window, calculate the event type distribution entropy value and the standard deviation of the adjacent operation interval; perform singular value decomposition on the API call frequency matrix, extract the first k principal components to form the call pattern characteristics; perform wavelet packet transform on the resource consumption fluctuation curve, and extract the three frequency band coefficients with the highest energy share; normalize and splice the event type distribution entropy value, the standard deviation of the adjacent operation interval, the call pattern characteristics and the frequency band coefficients to form a dimensionally configurable feature vector, which is the behavior feature vector.

[0010] Furthermore, the dynamic behavior baseline model is constructed by: The user's unique identifier is obtained through the login status and device identification of the mini program; the user's behavioral feature vector is collected for a period of time, which is recorded as the historical behavioral feature vector. The historical behavioral feature vector is reversely decomposed in the order of normalization and splicing to obtain the corresponding vector elements, and each vector element is used as a corresponding feature view; for each feature view, different unsupervised encoding methods are used to perform feature transformation in the corresponding encoding feature space; in the encoding feature space of each feature view, different clustering algorithms are used to obtain the clustering results under each feature view, the results of different clustering algorithms under the same feature view are fused to obtain the final clustering result, and the main clustering center in the clustering result is used as the dynamic behavior baseline model of the corresponding user.

[0011] Furthermore, the method of fusing the results of different clustering algorithms under the same feature view includes: sampling a certain number of instance pairs from each feature view, the instance pairs including similar positive example pairs and dissimilar negative example pairs; constructing a A triplet of , For example, Label the similarity; build a deep metric network, which consists of two sub-networks, each of which inputs an instance of a feature view, and the weights of the two sub-networks are shared; set a contrast loss function ; ; in, is the set of positive example pairs, is the set of counterexample pairs; for The index of the instance in , for The index of the instance in , Both index The instances in the index Examples in For the depth metric network instance The embedding map of For the depth metric network instance The embedding map of For example and examples The attention weight coefficient between For example and The attention weight coefficient between For the depth metric network, The embedding map of is the Gaussian kernel function, is the polynomial kernel function, is the L2 norm, is the regularization coefficient, is the comprehensive regularization term; is the set of all attention weights; Comprehensive regularization term The calculation method is: The state transition probability matrix PL is constructed based on the similarity annotations between instances. The elements in the state transition probability matrix ;in, For example and Examples The similarity between them is marked; then ;in, is the Frobenius norm; The trained deep metric network is used to generate compact embedding representations of instances of feature views. The mean distance between the compact embedding representations of instance pairs under the same feature view is calculated as the internal correlation measure of the feature view. The mean distance between the compact embedding representations of instance pairs under different feature views is calculated as the correlation measure between different feature views. The internal correlation measure of the feature view is normalized into the intra-view weight, and the correlation measure between different feature views is normalized into the inter-view weight. Then, the final view weight of the feature view is defined as intra-view weight × inter-view weight. For the same feature view, the results obtained by different clustering algorithms are weighted averaged, and the weight is the corresponding view weight.

[0012] Furthermore, the method of generating the real-time abnormality confidence coefficient includes: The dynamic behavior baseline model is modeled as a Gaussian mixture model, and the probability density value of the current behavior feature vector under each sub-model in the Gaussian mixture model is calculated; the probability density value is logarithmically transformed and weighted summed to obtain the deviation; the dynamic time warping distance between the current behavior feature vector and each record in the historical abnormal event library is calculated, and the smallest dynamic time warping distance is recorded as the similarity matching degree. The similarity matching degree and the deviation degree are weightedly fused to generate a real-time abnormality confidence coefficient; The construction method of the abnormal propagation probability graph includes: In time window t, obtain a snapshot G(t) of the network request response graph. In the subsequent time window t+1, obtain a new snapshot G(t+1) of the network request response graph. Compare G(t) and G(t+1) to extract the changed nodes and edges, including additions, deletions, and modifications. For newly added nodes, according to their semantic vectors, based on the relevance threshold t3, similar nodes are searched in G(t) and directed edges are established; for modified nodes, their semantic similarity is recalculated, and the association relationship is updated according to the semantic similarity, that is, directed edges are connected to other nodes, that is, the hidden association of cross-interface call links is identified; based on the topological structure of G(t+1), an initial anomaly propagation probability graph PU is constructed, and for each pair of associated request node A and response node B, the shortest path length d from A to B is calculated, and the initial propagation probability of the directed edge from A to B is set to ; Record the initial propagation probabilities of all edges in PU, and then obtain the final abnormal propagation probability graph.

[0013] Furthermore, the method of dynamically correcting the weights of nodes in the abnormal propagation probability graph includes: If the node is a request node, and the API corresponding to its URL has an abnormal peak in the API call frequency matrix, the real-time abnormal confidence coefficient of the API is defined as the preliminary adjustment coefficient of the corresponding node; if the node is a response node, and the associated request node has defined a preliminary adjustment coefficient, the response node inherits the value of the same preliminary adjustment coefficient; otherwise, the preliminary adjustment coefficient of the response node is calculated according to the abnormality of the response status code = f(abnormality of response status code × Ct_gin), where f() is a preset linear function and Ct_gin is a preset benchmark preliminary adjustment coefficient; For each node, its initial weight is defined as the normalized value of its API importance score, based on the initial weight of each node x , calculate its initial revised weight ;in, is the gain coefficient, is the preset control threshold, is the preliminary adjustment coefficient corresponding to the node; the Sigmoid function is restricted to the preliminary modified weight w'_x to obtain the corresponding node The middle weight w''_x of the directed edge is then adjusted to obtain the adjusted propagation probability. The formula for adjustment is: ;in, For Node and nodes The adjusted propagation probability of the directed edges between them; For Node and nodes The initial propagation probability of the directed edge between nodes, w''_y is the node The mid-range weight, For Node Initial weights; is the difference attenuation factor, For Node The corresponding preliminary adjustment coefficient, For Node The corresponding preliminary adjustment coefficient; based on the adjusted propagation probability of the directed edge, the mid-segment weights of the adjacent nodes are back-propagated and corrected, and the update is iterated until the change value of the mid-segment weight is less than the preset change threshold, and the update is stopped to obtain the corrected weight of the final node.

[0014] Furthermore, the calculation method of the comprehensive risk entropy value includes: Construct a three-dimensional risk space, where the three coordinate axes correspond to the deviation, the real-time anomaly confidence coefficient, and the node correction weight respectively; Calculate the Shannon entropy value of each coordinate axis, and analyze the coupling relationship between the coordinate axes through the covariance matrix to obtain the coupling coefficient; use the entropy weight method to determine the weight of each coordinate axis, recorded as the dimension weight; then calculate the comprehensive risk entropy value ;in, For the coordinate axis The dimension weight of For the coordinate axis The Shannon entropy value of is the covariance matrix The determinant value of is the coupling coefficient.

[0015] The technical effects and advantages of the applet abnormal behavior management system integrating artificial intelligence of the present invention are as follows: The present invention can comprehensively monitor and intelligently control complex abnormal behaviors to ensure the safe and reliable operation of mini-programs; it integrates multi-source heterogeneous data, automatically learns users' normal behavior patterns through artificial intelligence, accurately identifies abnormal deviations, and deeply explores the intrinsic correlation between abnormal behaviors to assess the scope of risk propagation; according to the comprehensive risk score, the level of the control strategy can be dynamically adjusted to achieve refined and hierarchical abnormal management; it can not only detect anomalies in a timely manner, but also automatically track the root causes of anomalies, generate traceability path audit logs, provide direct clues for abnormal diagnosis and repair, and improve response efficiency; at the same time, it can accurately locate abnormal code areas, inject behavior interception mechanisms, actively block potential threats, curb the spread of abnormal behaviors at the source, and maximize the protection of the healthy operation of mini-programs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of an abnormal behavior management system for small programs integrating artificial intelligence according to the present invention; Figure 2 This is a schematic diagram of a method for managing abnormal behavior of mini-programs that integrates artificial intelligence according to the present invention. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] Example 1 See also Figure 1 As shown, the mini-program abnormal behavior management system integrating artificial intelligence described in this embodiment includes: a data acquisition module, which is used to collect multi-dimensional data streams during the operation of the mini-program in real time, and the multi-dimensional data streams include user operation event sequences, API call frequency matrices, resource consumption fluctuation curves, and network request response graphs; The behavior modeling module is used to slice the multi-dimensional data stream into time and space according to the preset time window, extract the behavior feature vector in each time window, and build a dynamic behavior baseline model bound to the user identity; The preliminary matching module is used to calculate the deviation between the behavior feature vector and the dynamic behavior baseline model, and generate a real-time anomaly confidence coefficient by combining the similarity matching with the preset historical anomaly event library; The graph construction module is used to identify the hidden correlation of cross-interface call links according to the topological structure changes of the network request response graph, and then construct an anomaly propagation probability graph. The weights of the nodes in the anomaly propagation probability graph are dynamically corrected through the real-time anomaly confidence coefficient to obtain the corrected weights of the nodes in the anomaly propagation probability graph. The comprehensive management and control module is used to integrate the deviation degree, real-time anomaly confidence coefficient and the modified weights of the nodes in the anomaly propagation probability graph, calculate the comprehensive risk entropy value, preset the risk threshold, and dynamically trigger the multi-level management and control strategy when the comprehensive risk entropy value exceeds the risk threshold: according to the business module where the anomaly peak is located in the API call frequency matrix, locate the code hot zone and inject behavior interception probes, and generate an encrypted audit log containing the resource consumption fluctuation traceability path; each module is connected by wired and / or wireless means to realize data transmission between modules.

[0019] The methods for obtaining user operation event sequences include: A lightweight event capture hook is injected into the mini-program framework layer to obtain comprehensive data of user operation events; the comprehensive data includes type (click, slide, input, etc.), occurrence timestamp, event source component ID and event parameters, and these comprehensive data are serialized into an efficient binary format or compressed format to reduce the amount of data transmission and obtain a user event information sequence.

[0020] In the view layer of the mini program, use a front-end framework that supports virtual DOM trees, such as React, Vue, etc.; after the view layer is rendered, compare the differences between the new virtual DOM tree and the old virtual DOM tree to obtain only the changed DOM node paths, associate these DOM node paths with the corresponding user event information sequences, serialize and upload these changed DOM node paths and the corresponding user event information sequences without transmitting the data of the entire page, and then obtain the user operation event sequence; The methods for obtaining the API call frequency matrix include: An API call monitoring agent is implanted in the mini-program runtime environment. By modifying the context object (such as wx object) of the mini-program runtime, all API methods that need to be monitored are rewritten. In the rewritten API method, the timestamp and call parameters of the API call are first recorded; then the logic of the original API method is executed, and the return result is recorded. In this way, the call status of each API can be intercepted and recorded non-invasively.

[0021] Assign a unique ID to each API, serialize the API call timestamp, call parameters, and return results into an efficient binary format to obtain a call record; use the API ID as the key and store the corresponding call record into a ring buffer in memory.

[0022] Maintain a configuration table to store the importance score of each API. Specifically, assign an importance score of 0-10 to each API based on factors such as the API category, function, and corresponding business value. For example, the importance score of network request-related APIs is 9; the importance score of data storage-related APIs is 8; the importance score of media-related APIs is 7; the importance score of UI rendering-related APIs is 6; and the importance score of other tool APIs is 4.

[0023] Dynamically calculate a sampling rate baseline for each API ;in, and is the weight parameter, which is adjusted according to the actual situation, and the sum of the two is 1. is the current CPU usage, The current memory usage.

[0024] Preset an importance threshold. For APIs with an importance score greater than the importance threshold, multiply their sampling rate base value by a random number in the interval (1, 2] to obtain the corresponding sampling rate. For APIs with an importance score less than or equal to the importance threshold, use the sampling rate base value as the corresponding sampling rate. Generate a random floating point number between 0 and 1. If the sampling rate corresponding to the API is greater than or equal to the random floating point number, record the call record of this API. If the sampling rate corresponding to the API is less than the random floating point number, do not record the call record of this API and directly execute the API logic. With a fixed time window (such as 1 minute) as the cycle, scan the circular buffer in the memory, count the number of calls of each API within the time window, and build an API call frequency matrix for subsequent analysis; accurately obtain the call frequency of each API and dynamically adjust the sampling granularity to reduce performance overhead while ensuring data quality.

[0025] A lightweight resource collection agent is deployed in the container of the mini program. The resource collection agent periodically collects comprehensive indicators of the current mini program through the interface provided by the operating system, including CPU usage, memory usage, and network traffic. The collected comprehensive indicators are serialized by time, and a fluctuation trend chart of the indicators is drawn, which is the resource consumption fluctuation curve.

[0026] The methods for obtaining the network request response graph include: By modifying the context object of the network library when the mini-program is running, all methods of sending network requests and receiving responses are rewritten. In the rewritten methods, the original data packets of each HTTP / HTTPS request and response are intercepted and recorded as request packets and response packets; the intercepted request packets are parsed to identify the basic information of the request, including the URL path, query parameters, and request body.

[0027] Pre-train a semantic classification model, use it to classify the basic information of the request by semantic content, and mark the type of request (login, query, payment, etc.); for each request data packet, construct the basic information into a request vector (vector cascade), and for each response data packet, extract the response status code and response body and construct them into a response vector. The dimensions of the request vector and response vector are both n, which is configurable and usually takes a value of 20-50.

[0028] Use pre-trained semantic models such as BERT to embed request vectors and response vectors, map each request vector / response vector to an m-dimensional semantic vector, where m is usually 512-768, and calculate the cosine similarity between the request vector / response vector and the semantic vector as the corresponding semantic score, namely the request semantic score and the response semantic score.

[0029] Set the request semantic score threshold t1 and the response semantic score threshold t2. If the request semantic score is greater than or equal to t1, create a request node. If the response semantic score is greater than or equal to t2, create a response node. The node attributes of the response node and the request node include the request URL, request body hash, response status code, and the corresponding semantic vector. If the initiation time of the request node A is before the arrival time of the response node B, and the URL (target) of A is the same as the URL (source) of B, create a directed edge from A to B. The attributes of the directed edge include the response delay.

[0030] Calculate the cosine similarity between the semantic vectors of the request node A and the response node B, recorded as semantic similarity; set the relevance threshold t3, if the semantic similarity is greater than or equal to t3, create a directed edge from A to B, and the attributes of the directed edge include this cosine similarity.

[0031] Set the path compression threshold t4 and the similarity compression threshold t5. If the path length (the number of intermediate nodes) from request node A to response node B is greater than t4, and the product of the semantic similarity from request node to response node in the path length is greater than t5, then directly create a dependency edge from A to B and mark it as a compressed edge. That is, the network request response graph is obtained. Use the graph database Neo4j to store the constructed network request response graph, create node and edge indexes, speed up subsequent query operations, and periodically take snapshots of the graph for time retrospective analysis.

[0032] The methods for performing time-space slicing include: Set a fixed time window size, such as 1 minute, 5 minutes, etc. The time window size can be dynamically adjusted according to the real-time requirements of the data and the computing resources. Starting from the start time of the multi-dimensional data stream, generate a series of non-overlapping time slices according to the set time window size. Each time slice is represented by a tuple (start time, end time). Distribute the original multi-dimensional data stream, collect all the data in each tuple, and store the data in the corresponding data storage nodes, such as distributed file systems, databases, etc.

[0033] The methods for extracting behavioral feature vectors include: Perform sliding window statistics on the user operation event sequence within the time window, calculate the event type distribution entropy value and the standard deviation of the adjacent operation interval; perform singular value decomposition on the API call frequency matrix, extract the first k principal components to form the call pattern characteristics; perform wavelet packet transform on the resource consumption fluctuation curve, and extract the three frequency band coefficients with the highest energy share; normalize and splice the event type distribution entropy value, the standard deviation of the adjacent operation interval, the call pattern characteristics and the frequency band coefficients to form a dimensionally configurable feature vector, which is the behavior feature vector.

[0034] The dynamic behavior baseline model is constructed by: The user's unique identifier is obtained through the login status and device identification of the mini program, which can be obtained by quantifying the login status and the device identification and adding them together; the user's behavior feature vector is collected for a period of time (for example, 2 weeks), which is then recorded as a historical behavior feature vector. The historical behavior feature vector is decomposed in reverse according to the order of normalization and splicing to obtain the corresponding vector elements, and each vector element is used as a corresponding feature view; for example, event type distribution view (event type distribution entropy value), API call pattern view (call pattern feature), etc.

[0035] For each feature view, different unsupervised encoding methods are used to perform feature transformation in the corresponding encoding feature space. For example, the event type distribution view can use Word2Vec encoding; the API call mode view can use autoencoder encoding, etc.

[0036] In the encoded feature space of each feature view, different clustering algorithms (such as Gaussian mixture model, DBSCAN, spectral clustering, etc.) are used to obtain the clustering results under each feature view. The results of different clustering algorithms under the same feature view are fused to obtain the final clustering result. The main cluster center in the clustering result is used as the dynamic behavior baseline model of the corresponding user; the main cluster center is evaluated by the silhouette coefficient, and clusters with higher silhouette coefficients are usually more representative.

[0037] Specifically, a certain number of instance pairs are sampled from each feature view, and the instance pairs include similar positive example pairs and dissimilar negative example pairs; construct a A triplet of , For example, is a similarity label (a random number between 0 and 1); it should be noted that for each feature view (such as event type distribution view, API call pattern view, etc.), some sample instances are randomly selected from the historical data, and these sample instances are paired two by two to form instance pairs; for each instance pair, their similarity is calculated according to a certain similarity metric (such as Euclidean distance, cosine similarity, etc.); according to the similarity, the instance pairs are divided into similar positive example pairs and dissimilar negative example pairs.

[0038] Construct a deep metric network, which consists of two sub-networks. Each sub-network inputs an instance of a feature view. The weights of the two sub-networks are shared to ensure symmetry. The sub-network can be a convolutional network, a recursive network, etc., depending on the characteristics of the view data. Set a contrast loss function ; ; in, is the set of positive example pairs, is the set of counterexample pairs; for The index of the instance in , for The index of the instance in , Both index The instances in the index Examples in For the depth metric network, The embedding map of For the depth metric network, The embedding map of For example and examples The attention weight coefficient between For example and The attention weight coefficients between can all be randomly generated random numbers. For the depth metric network, The embedding map of is the Gaussian kernel function, is the polynomial kernel function, is the L2 norm, is the regularization coefficient, which controls the weight of the regularization term, is the comprehensive regularization term; is the set of all attention weights.

[0039] Comprehensive regularization term The calculation method is: The state transition probability matrix PL is constructed based on the similarity annotations between instances. The elements in the state transition probability matrix ;in, For example and Examples The similarity between them is marked; then ;in, is the Frobenius norm; The goal of the contrastive loss function is to minimize the distance between positive pairs and maximize the distance between negative pairs. The deep metric network is iteratively trained using instance pairs and the parameters of the deep metric network are optimized using an optimizer until the preset maximum number of iterations is reached to obtain a trained deep metric network.

[0040] The trained deep metric network can be used to generate compact embedding representations of instances of feature views, and the mean distance between the compact embedding representations of instance pairs under the same feature view is calculated as the internal correlation measure of the feature view; the mean distance between the compact embedding representations of instance pairs under different feature views is calculated as the correlation measure between different feature views.

[0041] Normalize the internal correlation measure of the feature view to the intra-view weight, and normalize the correlation measure between different feature views to the inter-view weight; then define the final view weight of the feature view = intra-view weight × inter-view weight; high-weight views have a greater impact on the results, and low-weight views have a smaller impact; for the same feature view, take a weighted average of the results obtained by different clustering algorithms, and the weight is the corresponding view weight.

[0042] Deep metric learning is used to automatically capture the correlation between views and adaptively learn the weight of each view based on the correlation; within the same view, the results of different clustering algorithms are weighted and fused according to the quality of the algorithm. Finally, we obtain a robust integrated clustering model that combines the advantages of multiple views and multiple algorithms, which can accurately characterize the normal behavior patterns of users.

[0043] The dynamic behavior baseline model is modeled as a Gaussian mixture model, and the probability density value of the current behavior feature vector under each sub-model in the Gaussian mixture model is calculated; the probability density value is logarithmically transformed and weighted summed to obtain the deviation; the dynamic time warping distance between the current behavior feature vector and each record in the historical abnormal event library is calculated, and the smallest dynamic time warping distance is recorded as the similarity matching degree. The similarity matching degree and the deviation degree are weightedly fused to generate a real-time anomaly confidence coefficient.

[0044] During the running of the mini-program, different network interface calls often have inherent associations and dependencies, but these associations may be implicit and not easy to observe directly. By analyzing the topological changes in the network request response graph, these hidden associations can be discovered and mined.

[0045] Specifically, a cross-interface call link means that the response of an interface A triggers a call to another interface B, and other interfaces may be cascaded in the middle; in the network request response graph, this situation may be manifested as: There is a path between A's response node and B's request node, but there is no direct edge connecting A's response node and B's request node. It is necessary to identify this indirect call relationship and construct an association edge of A response->B request to reveal the hidden call link association between them.

[0046] In time window t, obtain a snapshot G(t) of the network request response graph. In the subsequent time window t+1, obtain a new snapshot G(t+1) of the network request response graph, compare G(t) and G(t+1), extract the changed nodes (request nodes, response nodes) and edges (directed edges, compressed edges), including additions, deletions and modifications; for newly added nodes, according to their semantic vectors, based on the relevance threshold t3, find similar nodes in G(t) and establish directed edges; for modified nodes, recalculate their semantic similarity, update the association relationship according to the semantic similarity, that is, connect directed edges with other nodes, that is, identify the hidden association of cross-interface call links; the adjustment process is the same as the construction method of the network request response graph.

[0047] Based on the topological structure of G(t+1), an initial anomaly propagation probability graph PU is constructed. For each pair of associated request nodes A and response nodes B, the shortest path length d from A to B is calculated, and the initial propagation probability of the directed edge from A to B is set to ; The initial propagation probabilities of all edges are recorded in the PU, and then the final anomaly propagation probability graph is obtained; through the anomaly propagation probability graph, the propagation path and impact range of the anomaly starting from any node in the entire request response process can be simulated, laying the foundation for subsequent risk warning and precise control.

[0048] The dynamic weight correction adopts the following steps: If the node is a request node, and the API corresponding to its URL has an abnormal peak in the API call frequency matrix, the real-time abnormal confidence coefficient of the API is defined as the preliminary adjustment coefficient of the corresponding node; if the node is a response node, and its associated request node has defined a preliminary adjustment coefficient, the response node inherits the value of the same preliminary adjustment coefficient; otherwise, the preliminary adjustment coefficient of the response node is calculated according to the abnormality of the response status code (such as 5xx error) = f(abnormality of response status code × Ct_gin), where f() is a preset linear function, such as a linear function or a quadratic function; Ct_gin is a preset benchmark preliminary adjustment coefficient, usually 0.8.

[0049] It should be noted that the way to quantify the abnormality of the response status code is as follows: Categorize the status codes: 1xx: Informational status code; 2xx: Success status code; 3xx: redirect status code; 4xx: client error status code; 5xx: server error status code; Each type of status code is assigned a basic anomaly score, which is a quantitative value of the abnormality of the response status code; for example, 1xx: 0.1; 2xx: 0; 3xx: 0.3; 4xx: 0.7; 5xx:1.0.

[0050] For each node, its initial weight is defined as the normalized value of its API importance score, based on the initial weight of each node x , calculate its initial revised weight ;in, is the gain coefficient (default 0.5), which controls the correction amplitude. is the preset control threshold, is the preliminary adjustment coefficient corresponding to the node; the Sigmoid function is restricted to the preliminary modified weight w'_x to obtain the corresponding node The mid-segment weight w''_x.

[0051] Then adjust the initial propagation probability of the directed edge to obtain the adjusted propagation probability; the adjustment formula is: ;in, For Node and nodes The adjusted propagation probability of the directed edges between them; For Node and nodes The initial propagation probability of the directed edge between nodes, w''_y is the node The mid-range weight, For Node Initial weights; is the difference attenuation factor (default is 0.9). The greater the difference between the initial adjustment coefficients of the two end nodes, the more significant the propagation probability attenuation. For Node The corresponding preliminary adjustment coefficient, For Node The corresponding preliminary adjustment coefficient; Based on the adjusted propagation probability of the directed edge, the mid-segment weights of the adjacent nodes are back-propagated and corrected, and the update is iterated until the change value of the mid-segment weight is less than the preset change threshold. The update is stopped to obtain the corrected weight of the final node.

[0052] Construct a three-dimensional risk space, where the three coordinate axes correspond to the deviation, real-time anomaly confidence coefficient, and node correction weight respectively; calculate the Shannon entropy value of each coordinate axis, and analyze the coupling relationship between the coordinate axes through the covariance matrix to obtain the coupling coefficient (such as the average absolute correlation coefficient), which usually ranges from 0 to 1; correspondingly create an M×3 matrix, where each row represents a sample and each column represents a dimension; the covariance matrix is ​​constructed as the covariance of this M×3 matrix; The entropy weight method is used to determine the weight of each coordinate axis, which is recorded as the dimension weight; then the comprehensive risk entropy value is calculated ;in, For the coordinate axis The dimension weight of For the coordinate axis The Shannon entropy value of is the covariance matrix The determinant value of is the coupling coefficient; Analyze the API call frequency matrix to find out the APIs with abnormal peaks. Specifically, you can establish a normal call frequency baseline by counting its historical call frequencies. The baseline can be a fixed value or a range, depending on whether the API call pattern is periodic or otherwise regular. Exceeding the baseline means an abnormal peak.

[0053] According to the mapping relationship between API and business modules, determine the business module where the exception is located, use the code coverage tool to analyze the code execution of the abnormal business module, find the code block with the highest execution frequency, define it as the code hot zone, and insert lightweight behavior interception probes at key locations in the code hot zone. The code of the behavior interception probe contains behavior detection and interception logic, which can limit the frequency of API calls, resource usage, etc.; collect resource consumption fluctuation curves, and use time series data mining algorithms such as DTW (dynamic time warping) to analyze and obtain the resource consumption fluctuation traceability path, integrate the resource consumption fluctuation traceability path, abnormal API call information, and the location of the code hot zone, and use an asymmetric encryption algorithm to encrypt the audit log.

[0054] The multi-level control strategy is as follows: Level 1: Increase monitoring frequency and shorten data collection cycle; Level 2: Limit the frequency of abnormal API calls; Level 3: Temporarily disable some non-core functions; Level 4: Force users to log in again for authorization; Level 5: Temporary ban of user account; Continuously monitor the effectiveness of management and control, and dynamically adjust the level of management and control based on the anomaly mitigation situation (comprehensive risk entropy value and threshold distance assessment). If the anomaly persists, gradually increase the intensity of the management and control level. After the anomaly is eliminated, gradually relax the control measures.

[0055] This embodiment can comprehensively monitor and intelligently control complex abnormal behaviors to ensure the safe and reliable operation of the mini program; it integrates multi-source heterogeneous data, automatically learns users' normal behavior patterns through artificial intelligence, accurately identifies abnormal deviations, and deeply explores the inherent correlation between abnormal behaviors to assess the scope of risk propagation; according to the comprehensive risk score, the level of the control strategy can be dynamically adjusted to achieve refined and hierarchical abnormal management; it can not only detect anomalies in a timely manner, but also automatically track the root causes of anomalies and generate traceability path audit logs, providing direct clues for abnormal diagnosis and repair, and improving response efficiency; at the same time, it can accurately locate the abnormal code area, inject behavior interception mechanism, actively block potential threats, curb the spread of abnormal behavior from the source, and maximize the healthy operation of the mini program.

[0056] Example 2 See also Figure 2 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a method for managing abnormal behavior of mini-programs integrating artificial intelligence, including: Step 1: Collect multi-dimensional data streams in real time during the operation of the mini program, including user operation event sequences, API call frequency matrix, resource consumption fluctuation curve, and network request response graph; Step 2: Slice the multi-dimensional data stream in time and space according to the preset time window, extract the behavior feature vector in each time window, and build a dynamic behavior baseline model bound to the user identity; Step 3: Calculate the deviation between the behavior feature vector and the dynamic behavior baseline model, combine it with the similarity matching degree with the preset historical abnormal event library, and generate a real-time abnormal confidence coefficient; Step 4: According to the topological structure changes of the network request response graph, the hidden correlation of the cross-interface call link is identified, and then the anomaly propagation probability graph is constructed. The weights of the nodes in the anomaly propagation probability graph are dynamically corrected by the real-time anomaly confidence coefficient to obtain the corrected weights of the nodes in the anomaly propagation probability graph; Step 5: Integrate the deviation degree, real-time anomaly confidence coefficient, and the modified weights of the nodes in the anomaly propagation probability graph to calculate the comprehensive risk entropy value and preset the risk threshold. When the comprehensive risk entropy value exceeds the risk threshold, dynamically trigger the multi-level management and control strategy: According to the business module where the anomaly peak is located in the API call frequency matrix, locate the code hot zone and inject behavior interception probes, and generate an encrypted audit log containing the traceability path of resource consumption fluctuations.

[0057] Example 3 This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the method for managing abnormal behavior of a small program integrating artificial intelligence provided above is implemented.

[0058] Since the electronic device introduced in this embodiment is an electronic device used to implement a method for managing abnormal behavior of a small program that integrates artificial intelligence in the embodiment of this application, based on the method for managing abnormal behavior of a small program that integrates artificial intelligence introduced in the embodiment of this application, a person skilled in the art can understand the specific implementation method of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application is not described in detail here. As long as a person skilled in the art implements an electronic device used in a method for managing abnormal behavior of a small program that integrates artificial intelligence in the embodiment of this application, it belongs to the scope of protection of this application.

[0059] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0060] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A small program abnormal behavior management system integrating artificial intelligence, characterized in that: include: A data collection module is used to collect multi-dimensional data streams in real time during the operation of the mini-program, wherein the multi-dimensional data streams include user operation event sequences, API call frequency matrices, resource consumption fluctuation curves, and network request response graphs; A behavior modeling module, used to slice the multi-dimensional data stream into time and space according to preset time windows, extract the behavior feature vectors in each time window, and build a dynamic behavior baseline model bound to the user identity; A preliminary matching module is used to calculate the deviation between the behavior feature vector and the dynamic behavior baseline model, and generate a real-time anomaly confidence coefficient by combining the similarity matching degree with a preset historical anomaly event library; A graph construction module is used to identify the hidden correlation of cross-interface call links according to the topological structure changes of the network request response graph, and then construct an abnormal propagation probability graph, and dynamically correct the weights of the nodes in the abnormal propagation probability graph through the real-time abnormal confidence coefficient to obtain the corrected weights of the nodes in the abnormal propagation probability graph; The comprehensive management and control module is used to integrate the deviation, real-time anomaly confidence coefficient and the modified weights of the nodes in the anomaly propagation probability graph, calculate the comprehensive risk entropy value, preset the risk threshold, and dynamically trigger the multi-level management and control strategy when the comprehensive risk entropy value exceeds the risk threshold: locate the code hot zone and inject behavior interception probes according to the business module where the abnormal peak in the API call frequency matrix is ​​located, and generate an encrypted audit log containing the resource consumption fluctuation traceability path; each module is connected by wired and / or wireless means.

2. According to the artificial intelligence-integrated small program abnormal behavior management system according to claim 1, it is characterized in that: The method for acquiring the user operation event sequence includes: Inject a lightweight event capture hook into the mini-program framework layer to obtain comprehensive data of user operation events; the comprehensive data includes type, occurrence timestamp, event source component ID, and event parameters, and serialize these comprehensive data into binary format or compressed format to obtain a user event information sequence; In the view layer of the mini program, a front-end framework that supports virtual DOM trees is used. After the view layer is rendered, the new virtual DOM tree is compared with the old virtual DOM tree to obtain only the changed DOM node paths, and these DOM node paths are associated with the corresponding user event information sequences. These changed DOM node paths and the corresponding user event information sequences are serialized and uploaded to obtain the user operation event sequence.

3. According to the artificial intelligence-integrated mini-program abnormal behavior management system of claim 2, it is characterized in that: The method of obtaining the API call frequency matrix includes: An API call monitoring agent is implanted in the mini-program runtime environment. By modifying the context object of the mini-program runtime, all API methods that need to be monitored are rewritten. In the rewritten API method, the timestamp and call parameters of the API call are first recorded; then the logic of the original API method is executed, and the return result is recorded; a unique ID is assigned to each API, and the API call timestamp, call parameters, and return result are serialized into a binary format to obtain a call record; using the API ID as the key, the corresponding call record is stored in the ring buffer in the memory; Maintain a configuration table to store the importance score of each API and dynamically calculate a sampling rate benchmark value for each API ;in, and is the weight parameter; is the current CPU usage, is the current memory usage; Preset an importance threshold. For APIs whose importance scores are greater than the importance threshold, multiply their sampling rate base value by a random number in the interval (1, 2] to obtain the corresponding sampling rate. For APIs whose importance scores are less than or equal to the importance threshold, use the sampling rate base value as the corresponding sampling rate. Generate a random floating point number between 0 and 1. If the sampling rate corresponding to the API is greater than or equal to the random floating point number, the call record of this API is recorded. If the sampling rate corresponding to the API is less than the random floating point number, the call record of this API is not recorded and the API logic is executed directly. With a fixed time window as the period, the ring buffer in the memory is scanned, the number of calls of each API in the time window is counted, and the API call frequency matrix is ​​constructed.

4. According to the method of claim 3, the abnormal behavior management system of small programs integrating artificial intelligence is characterized in that: The method of obtaining the network request response graph includes: By modifying the context object of the network library when the applet is running, rewriting all methods of sending network requests and receiving responses, in the rewritten methods, intercepting the original data packets of each HTTP / HTTPS request and response, recording them as request packets and response packets; parsing the intercepted request packets, identifying the basic information of the request, including URL path, query parameters, and request body; Pre-train a semantic classification model, use the semantic classification model to classify the basic information of the request, and mark the type of the request; for each request data packet, construct the basic information into a request vector; for each response data packet, extract the response status code and response body, and construct them into a response vector; Use the pre-trained BERT semantic model to embed the request vector and response vector, map each request vector / response vector to an m-dimensional semantic vector, calculate the cosine similarity between the request vector / response vector and the semantic vector as the corresponding semantic score, i.e., the request semantic score and the response semantic score; set the request semantic score threshold t1 and the response semantic score threshold t2, if the request semantic score is greater than or equal to t1, create a request node, if the response semantic score is greater than or equal to t2, create a response node, the node attributes of the response node and the request node include the request URL, request body hash, response status code and the corresponding semantic vector; if the initiation time of the request node A is before the arrival time of the response node B, and the URL of A is the same as the URL of B, then create a directed edge from A to B, and the attributes of the directed edge include the response delay; calculate the cosine similarity between the semantic vectors of the request node A and the response node B, recorded as the semantic similarity; set the relevance threshold t3, if the semantic similarity is greater than or equal to t3, then create a directed edge from A to B, and the attributes of the directed edge include this cosine similarity; Set the path compression threshold t4 and the similarity compression threshold t5. If the path length from request node A to response node B is greater than t4, and the product of the semantic similarity from request node to response node in the path length is greater than t5, then directly create a dependency edge from A to B and mark it as a compressed edge, and obtain the network request response graph.

5. According to the method of claim 4, the abnormal behavior management system of small programs integrating artificial intelligence is characterized in that: The method of extracting the behavior feature vector includes: Perform sliding window statistics on the user operation event sequence within the time window, calculate the event type distribution entropy value and the standard deviation of the adjacent operation interval; perform singular value decomposition on the API call frequency matrix, extract the first k principal components to form the call pattern characteristics; perform wavelet packet transform on the resource consumption fluctuation curve, and extract the three frequency band coefficients with the highest energy share; normalize and splice the event type distribution entropy value, the standard deviation of the adjacent operation interval, the call pattern characteristics and the frequency band coefficients to form a dimensionally configurable feature vector, which is the behavior feature vector.

6. According to the artificial intelligence-integrated small program abnormal behavior management system of claim 5, it is characterized in that: The dynamic behavior baseline model is constructed by: The user's unique identifier is obtained through the login status and device identification of the mini program; the user's behavioral feature vector is collected for a period of time, which is recorded as the historical behavioral feature vector. The historical behavioral feature vector is reversely decomposed in the order of normalization and splicing to obtain the corresponding vector elements, and each vector element is used as a corresponding feature view; for each feature view, different unsupervised encoding methods are used to perform feature transformation in the corresponding encoding feature space; in the encoding feature space of each feature view, different clustering algorithms are used to obtain the clustering results under each feature view, the results of different clustering algorithms under the same feature view are fused to obtain the final clustering result, and the main clustering center in the clustering result is used as the dynamic behavior baseline model of the corresponding user.

7. According to the artificial intelligence-integrated small program abnormal behavior management system of claim 6, it is characterized in that: The method of fusing the results of different clustering algorithms under the same feature view includes: sampling a certain number of instance pairs from each feature view, the instance pairs including similar positive example pairs and dissimilar negative example pairs; constructing a A triplet of , For example, Label the similarity; build a deep metric network, which consists of two sub-networks, each of which inputs an instance of a feature view, and the weights of the two sub-networks are shared; set a contrast loss function ; ; in, is the set of positive example pairs, is the set of counterexample pairs; for The index of the instance in , for The index of the instance in , Both index The instances in the index Examples in; For the depth metric network instance The embedding map of For the depth metric network instance The embedding map of For example and examples The attention weight coefficient between For example and The attention weight coefficient between For the depth metric network instance The embedding map of is the Gaussian kernel function, is the polynomial kernel function, is the L2 norm, is the regularization coefficient, is the comprehensive regularization term; is the set of all attention weights; Comprehensive regularization term The calculation method is: The state transition probability matrix PL is constructed based on the similarity annotations between instances. The elements in the state transition probability matrix ;in, For example and Examples The similarity between them is marked; then ;in, is the Frobenius norm; The trained deep metric network is used to generate compact embedding representations of instances of feature views. The mean distance between the compact embedding representations of instance pairs under the same feature view is calculated as the internal correlation measure of the feature view. The mean distance between the compact embedding representations of instance pairs under different feature views is calculated as the correlation measure between different feature views. The internal correlation measure of the feature view is normalized into the intra-view weight, and the correlation measure between different feature views is normalized into the inter-view weight. Then, the final view weight of the feature view is defined as intra-view weight × inter-view weight. For the same feature view, the results obtained by different clustering algorithms are weighted averaged, and the weight is the corresponding view weight.

8. According to the artificial intelligence-integrated mini-program abnormal behavior management system of claim 7, it is characterized in that: The method of generating the real-time abnormality confidence coefficient includes: The dynamic behavior baseline model is modeled as a Gaussian mixture model, and the probability density value of the current behavior feature vector under each sub-model in the Gaussian mixture model is calculated; the probability density value is logarithmically transformed and weighted summed to obtain the deviation; the dynamic time warping distance between the current behavior feature vector and each record in the historical abnormal event library is calculated, and the smallest dynamic time warping distance is recorded as the similarity matching degree. The similarity matching degree and the deviation degree are weightedly fused to generate a real-time abnormality confidence coefficient; The construction method of the abnormal propagation probability graph includes: In time window t, obtain a snapshot G(t) of the network request response graph. In the subsequent time window t+1, obtain a new snapshot G(t+1) of the network request response graph. Compare G(t) and G(t+1) to extract the changed nodes and edges, including additions, deletions, and modifications. For newly added nodes, according to their semantic vectors, based on the relevance threshold t3, similar nodes are searched in G(t) and directed edges are established; for modified nodes, their semantic similarity is recalculated, and the association relationship is updated according to the semantic similarity, that is, directed edges are connected to other nodes, that is, the hidden association of cross-interface call links is identified; based on the topological structure of G(t+1), an initial anomaly propagation probability graph PU is constructed, and for each pair of associated request node A and response node B, the shortest path length d from A to B is calculated, and the initial propagation probability of the directed edge from A to B is set to ; Record the initial propagation probabilities of all edges in PU, and then obtain the final anomaly propagation probability graph.

9. According to the method of claim 8, the abnormal behavior management system of small programs integrating artificial intelligence is characterized in that: The method of dynamically correcting the weights of nodes in the abnormal propagation probability graph includes: If the node is a request node, and the API corresponding to its URL has an abnormal peak in the API call frequency matrix, the real-time abnormal confidence coefficient of the API is defined as the preliminary adjustment coefficient of the corresponding node; if the node is a response node, and the associated request node has defined a preliminary adjustment coefficient, the response node inherits the value of the same preliminary adjustment coefficient; otherwise, the preliminary adjustment coefficient of the response node is calculated according to the abnormality of the response status code = f(abnormality of response status code × Ct_gin), where f() is a preset linear function and Ct_gin is a preset benchmark preliminary adjustment coefficient; For each node, its initial weight is defined as the normalized value of its API importance score, based on the initial weight of each node x , calculate its initial revised weight ;in, is the gain coefficient, is the preset control threshold, is the preliminary adjustment coefficient corresponding to the node; the Sigmoid function is restricted to the preliminary modified weight w'_x to obtain the corresponding node The middle weight w''_x of the directed edge is then adjusted to obtain the adjusted propagation probability. The formula for adjustment is: ;in, For Node and nodes The adjusted propagation probability of the directed edges between them; For Node and nodes The initial propagation probability of the directed edge between nodes, w''_y is the node The mid-segment weight of the node Initial weights; is the difference attenuation factor, For Node The corresponding preliminary adjustment coefficient, For Node The corresponding preliminary adjustment coefficient; based on the adjusted propagation probability of the directed edge, the mid-segment weights of the adjacent nodes are back-propagated and corrected, and the update is iterated until the change value of the mid-segment weight is less than the preset change threshold, and the update is stopped to obtain the corrected weight of the final node.

10. According to the artificial intelligence-integrated small program abnormal behavior management system of claim 9, it is characterized in that: The calculation method of the comprehensive risk entropy value includes: Construct a three-dimensional risk space, where the three coordinate axes correspond to the deviation, the real-time anomaly confidence coefficient, and the node correction weight respectively; Calculate the Shannon entropy value of each coordinate axis, and analyze the coupling relationship between the coordinate axes through the covariance matrix to obtain the coupling coefficient; use the entropy weight method to determine the weight of each coordinate axis, recorded as the dimension weight; then calculate the comprehensive risk entropy value ;in, For the coordinate axis The dimension weight of For the coordinate axis The Shannon entropy value of is the covariance matrix The determinant value of is the coupling coefficient.

Citation Information

Patent Citations

  • Fraud-related APP detection method and system, electronic equipment and medium

    CN119691742A

  • Risk prediction method, system, equipment and storage medium based on hydropower units

    CN119740127A

  • Software supply chain risk detection protection method and system

    CN119808082A

  • Protection method, device and equipment of car networking service platform and storage medium

    CN119835068A

  • Method for temporal knowledge graph reasoning based on distributed attention

    US20230401466A1

Cited By

  • API interface security protection method based on anomaly detection

    CN120200850A

  • An API interface security protection method based on anomaly detection

    CN120200850B

  • Network security early warning method and system based on network port data

    CN120434057A

  • Multi-AI algorithm collaborative intelligent management system based on large model

    CN120455569A

  • Power grid data acquisition and analysis system based on big data

    CN120498050A