Anomaly detection method based on feature distribution adaptive update

By constructing a data ecosystem turbulence index and combining it with a multi-dimensional fusion model, the passive updating problem of anomaly detection methods in existing technologies is solved, enabling forward-looking insights into the data environment and improving security.

CN122634413APending Publication Date: 2026-08-25STORAGEX TECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610370274.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing anomaly detection methods are passively updated after data distribution drift, resulting in a high false alarm rate. Furthermore, they lack comprehensive consideration of internal system malfunctions and sudden changes in the external environment, posing security risks.

Method used

By constructing a data ecosystem turbulence index, and using state transition Markov entropy, semantic space Wasserstein distance, and model prediction uncertainty entropy for multi-dimensional fusion, the future stability of the data environment is predicted, the stability assessment threshold is dynamically adjusted, and the model is adaptively updated.

Benefits of technology

It enables early prediction of data distribution drift, reduces the risk of model contamination, improves the robustness and accuracy of the anomaly detection system in complex environments, and ensures security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122634413A_ABST
    Figure CN122634413A_ABST
Patent Text Reader

Abstract

The application provides an anomaly detection method based on feature distribution adaptive updating, and relates to the technical field of artificial intelligence and computer vision, and the specific steps comprise: obtaining to-be-detected data, and extracting corresponding to-be-detected features through digital fingerprint collection; according to a currently maintained feature consistency model, calculating the anomaly deviation degree of the to-be-detected features by using a consistency baseline verification mechanism; based on the anomaly deviation degree, filtering out normal candidate samples by using a trust threshold; based on group mode resonance evaluation, stably evaluating the feature distribution of the normal candidate samples; and updating the feature consistency model based on the normal candidate samples by using a cognitive core self-evolution mechanism; the application can dynamically adjust the updating strategy and even suspend the updating in an extreme risk condition by constructing a data ecological turbulence index which fuses the system internal state, external environment information and model self-uncertainty, and actively predicting the data environment risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and computer vision technology, specifically to an anomaly detection method based on adaptive update of feature distribution. Background Technology

[0002] With artificial intelligence and computer vision technologies deeply empowering various industries, anomaly detection, as a key technology for ensuring the safe, stable, and reliable operation of systems, is becoming increasingly important. Whether in industrial IoT condition monitoring, intrusion detection in cybersecurity, or environmental perception in autonomous driving systems, high-precision anomaly detection methods are needed to identify events that deviate from normal patterns in real time. As application scenarios become increasingly complex and dynamic, how to make consistency baseline verification mechanisms adaptive to cope with the constantly changing data distribution in the real world has become a key technological focus in this field.

[0003] Existing technologies have made significant strides in pursuing model adaptability, but there is still room for further improvement. Specifically, the following challenges exist: Existing methods typically only passively trigger model updates or adjustments after a significant shift in data distribution has been detected, resulting in model update decisions lagging by an average of 3-5 time windows compared to the actual time of the drift. This approach inherently introduces a time delay; in the early stages of a drift, the system still relies on the old model for judgments, leading to a 15%-30% increase in false alarm rates, making it difficult to meet the demands of scenarios with stringent real-time requirements.

[0004] Some existing technologies, such as the scheme disclosed in CN115879505A, while considering adaptive updates based on the correlation between data, rely primarily on the statistical characteristics of the data stream itself for decision-making. They lack a comprehensive consideration of the synergistic effects of internal system malfunctions, sudden changes in the external semantic environment, and blind spots in model cognition. Experiments show that when the system itself is unstable or the external environment is drastically volatile, relying solely on the data stream for update decisions results in a model contamination probability exceeding 40%, posing serious security risks. Summary of the Invention

[0005] The purpose of this invention is to provide an anomaly detection method based on adaptive feature distribution updates to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: An anomaly detection method based on adaptive feature distribution updating includes the following steps: Acquire the data to be detected and extract the corresponding features to be detected through digital fingerprint acquisition; Based on the currently maintained feature consistency model, the abnormal deviation of the feature to be detected is calculated using the consistency baseline verification mechanism; Based on the degree of abnormal deviation, a trust threshold is used to filter and select normal candidate samples. Based on the population pattern resonance assessment, the stability of the feature distribution of normal candidate samples is evaluated. When the feature distribution meets the preset stability conditions, the feature consistency model is updated based on normal candidate samples using the cognitive core self-evolution mechanism.

[0007] Furthermore, the features to be detected are in the form of multi-dimensional vectors, which are used to reflect the structural, semantic, and statistical information of the data to be detected.

[0008] Furthermore, the anomaly deviation is used to quantify the degree of abnormal deviation of the data to be detected relative to the normal feature distribution of the feature consistency model. The anomaly deviation is calculated based on feature distance, feature similarity difference, or feature reconstruction error.

[0009] Furthermore, the trust threshold filtering is based on at least two of the following conditions: anomaly deviation threshold, temporal continuity, spatial consistency, and distribution density, to jointly screen the data to be detected.

[0010] Furthermore, the group pattern resonance assessment is used to calculate the data ecology turbulence index, which characterizes the future instability risk of the data flow environment; based on the data ecology turbulence index, the preset stability conditions are dynamically adjusted.

[0011] Furthermore, the calculation process for stability assessment is as follows: At the start of the preset computation cycle, three independent computation tasks are launched in parallel to calculate the normalized state transition Markov entropy, semantic space Wasserstein distance, and model prediction uncertainty entropy within the current cycle. The first element is the Markov entropy of the state transition, the second element is the Wasserstein distance of the semantic space, and the third element is the model prediction uncertainty entropy. These are combined into a three-dimensional input vector. The input vectors of the past N periods are taken out from the queue used to store historical input vectors and together with the vector of the current period, they form a time series matrix. The time series matrix is ​​fed into a pre-loaded, offline-trained gated recurrent unit network model. After processing the time series matrix, the output layer of the gated recurrent unit network model generates a single scalar value between zero and one. This scalar value is defined as the data ecological turbulence index for this period. Obtain the first assessment threshold as a benchmark in the group pattern resonance assessment; combine the data ecological turbulence index to calculate the second assessment threshold that is actually effective in the current cycle; The second evaluation threshold is output to the population pattern resonance evaluation; the population pattern resonance evaluation will use the second evaluation threshold as the judgment criterion when conducting subsequent stability evaluation.

[0012] Furthermore, the steps for calculating the Markov entropy of the state transition include: obtaining the internal state transition sequence of the system within a preset time window; constructing the first-order Markov transition probability matrix of the state transition sequence; and calculating the information entropy based on the transition probability matrix.

[0013] Furthermore, the steps for calculating the Wasserstein distance in the semantic space include: acquiring external unstructured text data related to the system's application domain through a network interface; converting the text data into semantic vectors using a pre-trained large language model; and calculating the Wasserstein distance between the empirical distribution of the semantic vectors within the current time window and the preset long-term baseline distribution.

[0014] An anomaly detection system based on adaptive feature distribution updates includes a processor and a computer program stored in memory and capable of running on the processor.

[0015] A computer-readable storage medium storing a computer program.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention no longer bases model update decisions on the analysis of the current data stream, but instead constructs a multi-dimensional, predictive data ecosystem turbulence index. Through an advanced time-series fusion model, it deeply integrates three orthogonal and complementary data dimensions from three sources, thereby gaining insight into the risks of future data quality instability. By calculating the state transition Markov entropy, it quantifies the degree of disorder in the internal operating state of the monitored system, perceiving its own health status. By calculating the Wasserstein distance in the semantic space, it quantifies the semantic drift intensity of related topics in the external public information domain, perceiving potential and new external threats or environmental changes. By calculating the model prediction uncertainty entropy, it quantifies the confusion of the consistency baseline verification mechanism itself with the output results, perceiving whether the current data is in the model's cognitive blind spot. This invention can predict that the entire data ecosystem is about to enter a turbulent state before the data quality actually deteriorates; based on predictive indices, a two-layer dynamic risk control mechanism is further designed; the stability assessment threshold used to determine whether to update the model is adjusted in real time and with fine precision to achieve adaptive adjustment that matches the risk level; and an absolute turbulence alarm threshold is set so that when extreme risks are predicted, the update operation can be decisively shut down, providing the final guarantee for safety. By constructing a data ecosystem turbulence index, this invention enables early prediction of data distribution drift, triggering defensive strategies 2-3 time windows before the actual drift occurs; by dynamically weighting three heterogeneous data sources through an adaptive nonlinear fusion network, the prediction accuracy is improved compared to the simple weighted average method; through circuit breaking and rollback mechanisms, the risk of model contamination can be reduced in extreme environments, while the risk of contamination in such scenarios by existing technologies is usually greater than that. In summary, this invention, by constructing and utilizing the data ecosystem turbulence index, enables the anomaly detection system to possess, for the first time, a forward-looking insight into the future stability of the data environment. This allows the system to adapt nimbly to benign environmental changes while fundamentally avoiding the risk of data contamination damaging the model during periods of drastic environmental change. Consequently, without human intervention, the robustness, accuracy, and security of the anomaly detection system in long-term, dynamic, and complex real-world environments are improved. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the simulation framework of the present invention; Figure 2 This is a schematic diagram of the overall method flow of the present invention; Figure 3 This is a schematic diagram of the calculation process for the stability assessment of this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly. Example

[0020] Please see Figures 1 to 3 This invention provides a technical solution: an anomaly detection method based on adaptive update of feature distribution, comprising the following steps: Acquire the data to be detected and extract the corresponding features to be detected through digital fingerprint acquisition; The features to be detected are in the form of multi-dimensional vectors, used to reflect the structural, semantic, and statistical information of the data to be detected; Structural information describes the form and composition of data: focusing on the format, size, and complexity of the data itself; Specifically, this includes request / response size, URL / path complexity, data format complexity, and protocol header information; Semantic information describes the intent and meaning of data: focusing on the deeper meaning of data content is key to understanding its behavioral intent; Specifically, this includes semantic vectors of text content, quantization methods, and API endpoint category encoding; Statistical information describes the behavioral patterns of data over time and space; it focuses on the correlation between data points and historical behavior and the surrounding environment. Specifically, this includes frequency and rate, time-series behavior indicators, and historical pattern entropy values.

[0021] Based on the currently maintained feature consistency model, the abnormal deviation of the feature to be detected is calculated using the consistency baseline verification mechanism; Feature consistency models can be represented using feature centers, feature distribution ranges, or feature similarity models; Anomaly deviation is used to quantify the degree of abnormal deviation of the data to be detected relative to the normal feature distribution of the feature consistency model; Anomaly deviation is calculated based on feature distance, feature similarity difference, or feature reconstruction error; The calculation logic based on feature distance: Define the normal point as the center point, calculate the geometric distance from the current data point to the center point, and the farther the distance, the more abnormal it is.

[0022] The normal representation is the feature center vector, which is obtained by averaging the feature vectors of a large number of historical normal samples. The calculation logic is as follows: obtain the feature vector of the current data to be detected; obtain the pre-calculated feature center vector; calculate the Euclidean distance or other distance metric between the feature vector and the feature center vector (in this embodiment, the other distance metric is Mahalanobis distance); calculate the sum of squares of the differences between the feature vector and the feature center vector in each dimension, and then take the square root, which intuitively represents the straight-line distance in multidimensional space; suitable for scenarios where the normal data distribution is relatively concentrated and presents a spherical or ellipsoidal distribution.

[0023] Calculation logic based on feature similarity difference: We don't care about absolute distance, but rather the consistency of direction between two vectors; if the feature vector of the current data point points in a direction that differs greatly from that of the normal vector, then it is an anomaly. A normal representation is a feature center vector, or a set of representative normal sample vectors; The calculation logic is as follows: obtain the feature vector of the current data to be detected; obtain the normal feature center vector; calculate the cosine similarity between the feature vector and the feature center vector; the cosine similarity measures the cosine value of the angle between the feature vector and the feature center vector, which is in the range of [-1, 1]. The closer the value is to 1, the more consistent the directions are. Subtracting the similarity from 1 transforms the similarity index (higher is more normal) into a deviation index (higher is more abnormal). When the directions are completely consistent, the deviation is 0, and when the directions are completely opposite, the deviation is 2. This method is very effective when the length (size) of the feature vector is not important, but the direction (pattern) is more important. The calculation logic based on feature reconstruction error is as follows: Train a model to only learn how to understand and reproduce normal data; when abnormal data is input, the model will not be able to reproduce it well, and the difference between the original data and the reproduced data will be large. The normal representation is an autoencoder neural network; after being trained on a large number of normal samples, the autoencoder neural network learns to compress the input data into a low-dimensional representation, and then decompress it to reconstruct the original input. The calculation logic is as follows: The process involves: acquiring the feature vector of the current data to be detected; inputting the feature vector into a pre-trained autoencoder model; the autoencoder model outputting a reconstructed vector; calculating the difference between the feature vector and the reconstructed vector, typically using mean squared error (MSE); MSE directly quantifies the degree of failure of the autoencoder model in reconstructing the current data; for normal data familiar to the autoencoder model, the error will be small; for unfamiliar anomalous data, the error will be large; it is suitable for non-linear data distributions and can capture deeper relationships between features.

[0024] Based on the degree of abnormal deviation, a trust threshold is used to filter and select normal candidate samples. Trust threshold filtering is based on at least two of the following conditions: anomaly deviation threshold, temporal continuity, spatial consistency, and distribution density, which are used to jointly screen the data to be detected. Data meeting the following criteria are selected as normal candidate samples: the abnormal deviation is below the abnormal deviation threshold; the deviation changes little over multiple consecutive detection periods; and the data exhibits continuity in the spatial or temporal dimensions. The abnormal deviation threshold is set to 0.15. Before performing conditional checks, define the following configurable system parameters: In the current testing cycle The normalized deviation is calculated, ranging from [0,1]. The time stability evaluation window is used to evaluate the number of consecutive periods of deviation change; the set value is 5 periods. The variance threshold is used to determine whether the deviation change is small; the set value is 0.001. The spatial neighbor set acquisition function takes the current data point as input and returns a set of spatial neighbor data points. The key attribute extraction function is used to extract the key feature values ​​used to determine continuity in the data points. The spatial continuity tolerance is used to determine whether the key attributes are similar; the set value is 0.1. A precise description of the filtering logic: The data points to be detected in the current detection period A sample is selected as a normal candidate sample if and only if all three of the following conditions are met: Condition 1, Instantaneous Deviation Review: The abnormal deviation of the current data point must be lower than the set abnormal deviation threshold; Condition 2, Time Series Stability Review: Within the continuous period of the time stability assessment window, the change in abnormal deviation must be sufficiently small; obtain the deviation sequence within the time window; calculate the variance of the deviation sequence; Condition 3, Spatial / Temporal Consistency Check: Data points must demonstrate consistency with their neighbors or the data from the previous moment; Calculation logic: Time consistency: The difference between the current deviation and the deviation of the previous period should not exceed half of the abnormal deviation to avoid sudden changes. Spatial consistency requires the existence of at least one spatial neighbor such that the relative rate of change of the key attributes of the current data point compared with the same attributes of its neighbors is lower than the spatial continuity tolerance. A loop is needed to check all neighbors; the check passes as long as one of them meets the criteria.

[0025] Based on the population pattern resonance assessment, the stability of the feature distribution of normal candidate samples is evaluated. The group pattern resonance assessment is used to calculate the data ecology turbulence index, which characterizes the future instability risk of the data flow environment; based on the data ecology turbulence index, the preset stability conditions are dynamically adjusted. The Data Ecosystem Turbulence Index is a predictive indicator obtained by time-series fusion of at least two heterogeneous data sources using a pre-defined time-series fusion model. The heterogeneous data sources include a first data source, a second data source, and a third data source. The first data source represents the state transition Markov entropy, indicating the degree of disorder in the system's internal operating state. The second data source represents the semantic space Wasserstein distance of changes in the external technological environment in which the system exists. The third data source model represents the predictive uncertainty entropy, indicating the confidence level of the quantified consistency baseline verification mechanism in the prediction results. The time-series fusion model uses a gated recurrent unit network (GRU) at each time step. The step-t receives a vector of state transition Markov entropy, semantic space Wasserstein distance, and model prediction uncertainty entropy as input, and outputs a scalar value in the range [0,1] as the data ecosystem turbulence index. By fusing information from the system's internal state and external environment, a virtual sensor for predicting the future stability of the data flow is constructed. Instead of passively assessing the stability of the current sample, it can predict whether the entire data ecosystem is about to enter a turbulent state. This allows the system to take conservative strategies in advance before the data quality declines, enhancing the safety of model updates and fundamentally avoiding the risk of the model being contaminated by transient, pseudo-stable data during periods of drastic environmental change.

[0026] The steps for calculating the Markov entropy of the state transition include: obtaining the internal state transition sequence of the system within a preset time window; constructing the first-order Markov transition probability matrix of the state transition sequence; and calculating the information entropy based on the transition probability matrix. The steps for calculating the Wasserstein distance in semantic space include: acquiring external unstructured text data related to the system's application domain through a network interface; converting the text data into semantic vectors using a pre-trained large language model; and calculating the Wasserstein distance between the empirical distribution of semantic vectors within the current time window and the preset long-term baseline distribution. The steps for dynamically adjusting the preset stability conditions based on the data ecological turbulence index are as follows: The first evaluation threshold in the preset stability conditions is set as an adjustment function positively correlated with the data ecological turbulence index, used as input to the data ecological turbulence index, to generate the second evaluation threshold. The initial value of the first evaluation threshold is set to 0.05, and the second evaluation threshold is generated using the adjustment function, which is set as follows: ,in It is a monotonically increasing function. For the data ecosystem turbulence index, This is the second evaluation threshold; This embodiment is specifically set As a linear function, the second assessment threshold increases by a fixed percentage for every unit increase in the data ecological turbulence index; the response is stable and predictable; suitable for maintaining a constant ratio between adaptability and the rate of environmental change. Piecewise functions are set with different response sensitivities under different data ecological turbulence indices, increasing slowly in the low turbulence region and rapidly in the high turbulence region; this is suitable for complex scenarios that require more refined and nonlinear responses to specific environmental conditions. The logarithmic / exponential function is highly sensitive to the initial response (the threshold rises rapidly when turbulence just begins to increase), but the growth then slows down (it no longer overreacts to extremely high turbulence); the initial response is flat, but the threshold rises rapidly as turbulence increases; it is used to simulate adaptive strategies with diminishing marginal effects or accelerated risk accumulation. The Sigmoid function exhibits a gradual threshold change in regions with low and extremely high turbulence exponents; however, in a certain intermediate critical region, the threshold undergoes a sharp, jump-like increase. It is suitable for simulating phase transition processes from steady to dynamic states and can make the most rapid and decisive strategy adjustments at key inflection points.

[0027] When the data ecosystem turbulence index exceeds the preset turbulence alarm threshold, the cognitive core self-evolution mechanism is configured to suspend the execution of update operations and maintain the current feature consistency model unchanged. The turbulence alarm threshold ranges from [0,1], with a specific value of 0.9.

[0028] The parameters involved in this embodiment are defined as follows: The parameter sign of the Markov entropy for state transition is set to It is used to quantify the disorder or chaos of the state transition of the data to be detected within a preset time window; a high entropy value indicates that the system state transition is frequent and irregular, which is a sign of instability; based on Shannon entropy theory in information theory, it is applied to the first-order Markov chain model constructed by analyzing the system operation log; the state transition process of the system is regarded as a random process, and information entropy is used to measure uncertainty; The calculation logic is as follows: A finite set of state spaces is defined, containing all key operational states of the monitored object that can be identified from logs. In this embodiment, for a network server application, the state space is defined as six states: {Idle Waiting, Data Receiving, Business Processing, Data Sending, Connection Retry, Internal Error}. A first time window is set to count the frequency of state transitions, and the parameter symbol is set to... The first time window is used to define the characteristic distribution range of normal candidate samples used to calculate entropy. In this embodiment, the first time window is set to ten minutes. Within the first time window, the running log is traversed to count the number of times a transition occurs from one state to another, and a transition frequency matrix is ​​constructed. Each element in the transition frequency matrix is ​​divided by the sum of all elements in its row to normalize the frequency matrix and obtain the transition probability matrix. The elements in the transition probability matrix are the conditional probabilities of transitioning from one state to another. The probability distribution of the system in each state within the same time window is obtained by calculating the proportion of the total dwell time of each state to the total time. Based on the calculation principle of Shannon entropy, for each state, the occurrence probability is multiplied by the calculated value, which is the conditional entropy calculated based on the outgoing transition probabilities of the state. Specifically, for each outgoing transition probability, the base-2 logarithm of the outgoing transition probability is taken and then multiplied by the outgoing transition probability itself. The sum of all outgoing transition results is taken as a negative value to obtain the conditional entropy. The sum of the calculation results of all states is used to obtain the final state transition Markov entropy. To facilitate subsequent integration, the calculated entropy value is normalized. By analyzing a large amount of historical data offline, the typical entropy value range of the system under normal operation is determined. The lower limit of the entropy value is subtracted from the currently calculated entropy value, and then the difference is divided by the difference between the upper and lower limits of the entropy value. The entropy value is then mapped to the interval between zero and one. The parameter sign of the semantic space Wasserstein distance is set to This is used to quantify the degree of semantic drift of focus in the external public information domain related to the application domain within a preset time window; a high distance value indicates that the external technical environment is undergoing drastic changes. In this embodiment, the disclosure of new vulnerabilities, the discussion of new attack methods, etc., indicate potential and new external threats; the calculation relies on a deep language model pre-trained on a large-scale corpus, and the Wasserstein distance calculation method based on optimal transport theory; this embodiment uses the BERT-base-uncased model as a tool to convert text into high-dimensional semantic vectors; it can deeply understand the contextual information of the text, and the generated vectors can more accurately reflect the true semantics of the text; The calculation logic is as follows: The latest text data is obtained from multiple preset high-quality information sources (in this embodiment, several technical forums in a specific field, official security bulletin websites, and relevant open-source project code commit logs) to form a recent text corpus; the second time window parameter for the acquisition operation is set to... In this embodiment, the second time window is set to 24 hours to capture the main changes in public opinion within a day; each piece of text data in the recent text corpus is preprocessed, including removing irrelevant characters and converting to lowercase; each preprocessed piece of text data is input into the loaded BERT-base-uncased model to extract the corresponding 768-dimensional semantic vector; the vectors together constitute the recent semantic distribution; the pre-stored baseline semantic distribution is loaded; the baseline semantic distribution is a semantic vector obtained by performing the same processing on text data obtained over a relatively long period of time (in this embodiment, it is set to 30 days), and the semantic vector represents the topic distribution under normal conditions; the Wasserstein distance between the recent semantic distribution and the baseline semantic distribution is calculated; the calculated distance value is normalized. The parameter sign of the model prediction uncertainty entropy is set to ; Used to quantify the confidence of the consistency baseline verification mechanism in the prediction results; a high entropy value indicates that the consistency baseline verification mechanism is confused by the currently received data, and the output prediction results tend to be random or uniformly distributed, indicating that the current data belongs to a gray area that the consistency baseline verification mechanism has never seen before and is difficult to judge; the calculation is also based on Shannon entropy theory. The calculation logic is as follows: Obtain the predicted probability distribution vector output by the consistency baseline verification mechanism after processing the data to be detected. In this embodiment, if the consistency baseline verification mechanism divides the data into three categories: "normal," "suspected abnormal," and "confirmed abnormal," the output is a vector containing three probability values, the sum of which is one. Apply the Shannon entropy calculation logic to the predicted probability distribution vector. Specifically: for each probability value in the predicted probability distribution vector, take the base-2 logarithm of the probability value, and then multiply it by the probability value itself. Summate all these calculation results and take the negative value to obtain the prediction uncertainty entropy of a single sample. Set a third time window for smoothing the entropy value, with the parameter symbol [symbol missing]. In this embodiment, the time window is set to one minute; the third time window aims to smooth out noise from a single sample; it reflects the average uncertainty over a period of time; the arithmetic mean of the prediction uncertainty entropy of all samples within the third time window is calculated as the final model prediction uncertainty entropy. Normalize the calculated average entropy value; the theoretical upper limit is obtained when the probability is uniformly distributed. In this embodiment, the upper limit for the three-class classification problem is... The lower bound is zero; normalization is performed using the theoretical lower bound, mapping to the interval between zero and one; The calculation process for stability assessment is as follows: At the start of a preset computation cycle, three independent computation tasks are launched in parallel to calculate the normalized state transition Markov entropy for the current cycle. Semantic space Wasserstein distance and model prediction uncertainty entropy ; Transition the state to Markov entropy Semantic space Wasserstein distance and model prediction uncertainty entropy In this embodiment, following a preset order, the state transition Markov entropy is set as follows: As the first element, the Wasserstein distance in the semantic space As the second element, the model predicts the uncertainty entropy. The third element is used to form a three-dimensional input vector. The input vectors of the past N periods (in this embodiment, N=120, representing the past two hours) are taken out from the queue used to store historical input vectors and combined with the vector of the current period to form a time series matrix. The time series matrix is ​​fed as input into a pre-loaded, offline-trained gated recurrent unit (GRU) network model. After processing the time series matrix, the output layer of the GRU model generates a single scalar value between 0 and 1. This scalar value is defined as the data ecological turbulence index for the current period. ; Obtain the first evaluation threshold as a benchmark in the population pattern resonance assessment. The parameter symbol of the first evaluation threshold is: First evaluation threshold It is an empirical value calibrated under the most stable system condition; based on a preset, non-linear adjustment function, combined with the data ecological turbulence index. Calculate the second evaluation threshold that is actually effective in the current period. The parameter symbol of the second evaluation threshold is: In this embodiment, the calculation logic of the adjustment function is as follows: subtract the data ecological turbulence index from the constant "1". Square this difference; then compare the squared result with the first evaluation threshold. Multiply to obtain the second evaluation threshold. ; The second evaluation threshold The output is then sent to the swarm pattern resonance assessment; the swarm pattern resonance assessment will use the second assessment threshold in the subsequent stability assessment. It serves as a criterion for judgment, rather than a fixed static threshold.

[0029] When the feature distribution meets the preset stability conditions, the feature consistency model is updated based on normal candidate samples using the cognitive core self-evolution mechanism. The update methods for feature consistency models include weighted average update, sliding window update, or incremental update; Based on the updated feature consistency model, anomaly detection is performed on the subsequent input data to be detected, and the anomaly detection results are output. The anomaly detection results include anomaly location, anomaly severity score, or anomaly region identification information. When performing updates, the cognitive core self-evolution mechanism includes an abnormal feature masking mechanism, which is used to identify and reduce the weight of feature dimensions with large abnormal fluctuations in the feature consistency model update process. An anomaly detection system based on adaptive feature distribution updates includes a processor and a computer program stored in memory and executable on the processor. When the computer program is executed, it implements the method described above.

[0030] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0031] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, Min-Max Normalization and Z-Score standardization. The technical solution of this invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of this invention.

[0032] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0033] This embodiment provides specific algorithm parameter configurations and comparative experimental data to verify the technical effects of the present invention.

[0034] I. Adaptive Nonlinear Fusion Network Structure Parameters: The gated recurrent unit network adopts a sequence-to-sequence prediction architecture, with the input layer receiving a three-dimensional vector sequence of the past 120 periods (N=120). The network structure is as follows: First-layer GRU: 128 hidden units, tanh activation function, and 0.2 dropout rate during backpropagation; The second-layer GRU has 64 hidden units and the activation function is tanh. Fully connected layer: Contains 32 neurons, ReLU activated; Output layer: 1 neuron, sigmoid activation, output data ecological turbulence index: .

[0035] II. Specific implementation of dynamic threshold adjustment: The preset stability conditions are determined by dynamically evaluating thresholds. Achieve a fixed threshold, rather than a fixed threshold. In this embodiment, an exponentially decaying nonlinear mapping function is used: ; when When =0 (absolutely stable), =0.05; when When =0.9 (extreme turbulence), It automatically decreases to approximately 0.012, significantly increasing the rigor of stability assessment.

[0036] III. Circuit Breaker and Rollback Mechanism: Set turbulence alarm threshold =0.9.

[0037] when hour: Immediately suspend the model update operation of the cognitive core self-evolution mechanism; Roll back the feature consistency model to the previous checkpoint version that passed the stability assessment; An alarm is triggered, and the administrator is notified to intervene.

[0038] IV. Comparison of experimental data: Tests were conducted on a dataset containing 100,000 network traffic data points, simulating three environmental change scenarios: Scenario A (Gradual Drift): The data distribution changes slowly; Scenario B (Mutation Attack): A sudden external attack causes a drastic change in data distribution; Scenario C (Chaotic State): Internal system errors and external attacks occur simultaneously.

[0039] The experimental results are shown in the table below:

[0040] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An anomaly detection method based on adaptive update of feature distribution, characterized in that, The specific steps include: Acquire the data to be detected and extract the corresponding features to be detected through digital fingerprint acquisition; Based on the currently maintained feature consistency model, the abnormal deviation of the feature to be detected is calculated using the consistency baseline verification mechanism; Based on the degree of abnormal deviation, a trust threshold is used to filter and select normal candidate samples. Based on the group pattern resonance assessment, the stability of the feature distribution of normal candidate samples is evaluated. The group pattern resonance assessment is achieved by calculating the data ecological turbulence index. The data ecological turbulence index is used to perform time-series fusion of three heterogeneous data sources: internal state transition Markov entropy, external environment semantic space Wasserstein distance, and model prediction uncertainty entropy through an adaptive nonlinear fusion network. During the fusion process, the weight coefficients of each data source are dynamically adjusted according to historical prediction errors. When the feature distribution meets the preset stability conditions, the feature consistency model is updated based on normal candidate samples using the cognitive core self-evolution mechanism; when the data ecology turbulence index exceeds the preset turbulence alarm threshold, the circuit breaker mechanism is triggered to suspend the model update and the rollback mechanism is started to restore the feature consistency model to the previous stable version.

2. The anomaly detection method based on adaptive feature distribution update according to claim 1, characterized in that: The features to be detected are in the form of multi-dimensional vectors, which are used to reflect the structural, semantic and statistical information of the data to be detected.

3. The anomaly detection method based on adaptive feature distribution update according to claim 2, characterized in that: Anomaly deviation is used to quantify the degree of abnormal deviation of the data to be detected relative to the normal feature distribution of the feature consistency model. Anomaly deviation is calculated based on feature distance, feature similarity difference, or feature reconstruction error.

4. The anomaly detection method based on adaptive feature distribution update according to claim 3, characterized in that: Trust threshold filtering is based on at least two of the following conditions: anomaly deviation threshold, temporal continuity, spatial consistency, and distribution density, to jointly screen the data to be detected.

5. The anomaly detection method based on adaptive feature distribution update according to claim 4, characterized in that: The group pattern resonance assessment is used to calculate the data ecology turbulence index, which characterizes the future instability risk of the data flow environment; based on the data ecology turbulence index, the preset stability conditions are dynamically adjusted.

6. The anomaly detection method based on adaptive feature distribution update according to claim 5, characterized in that: The calculation process for stability assessment is as follows: At the start of the preset computation cycle, three independent computation tasks are launched in parallel to calculate the normalized state transition Markov entropy, semantic space Wasserstein distance, and model prediction uncertainty entropy within the current cycle. The first element is the Markov entropy of the state transition, the second element is the Wasserstein distance of the semantic space, and the third element is the model prediction uncertainty entropy. These are combined into a three-dimensional input vector. The input vectors of the past N periods are taken out from the queue used to store historical input vectors and together with the vector of the current period, they form a time series matrix. The time series matrix is ​​fed into a pre-loaded, offline-trained gated recurrent unit network model. After processing the time series matrix, the output layer of the gated recurrent unit network model generates a single scalar value between 0 and 1. This scalar value is defined as the data ecological turbulence index for the current period. Obtain the first assessment threshold as a benchmark in the group pattern resonance assessment; combine the data ecological turbulence index to calculate the second assessment threshold that is actually effective in the current cycle; The second evaluation threshold is output to the population pattern resonance evaluation; the population pattern resonance evaluation will use the second evaluation threshold as the judgment criterion when conducting subsequent stability evaluation.

7. The anomaly detection method based on adaptive feature distribution updating according to claim 6, characterized in that: The steps for calculating the Markov entropy of the state transition include: obtaining the internal state transition sequence of the system within a preset time window; constructing the first-order Markov transition probability matrix of the state transition sequence; and calculating the information entropy based on the transition probability matrix.

8. The anomaly detection method based on adaptive feature distribution update according to claim 7, characterized in that: The steps for calculating the Wasserstein distance in semantic space include: acquiring external unstructured text data related to the system's application domain through a network interface; converting the text data into semantic vectors using a pre-trained large language model; and calculating the Wasserstein distance between the empirical distribution of semantic vectors within the current time window and the preset long-term baseline distribution.

9. The anomaly detection method based on adaptive feature distribution update according to claim 1, characterized in that: The adaptive nonlinear fusion network is a gated recurrent unit network, which includes two GRU layers and a fully connected output layer. The first GRU layer contains 128 hidden units, the second GRU layer contains 64 hidden units, and the output layer outputs the data ecological turbulence index through the Sigmoid activation function.

10. The anomaly detection method based on adaptive feature distribution updating according to claim 1, characterized in that: The stability condition is achieved through a dynamic evaluation threshold, which is correlated with the data ecological turbulence index via a nonlinear mapping function. The nonlinear mapping function is: ; in, To dynamically evaluate the threshold, As the baseline threshold, The data ecology turbulence index is denoted by α and β, which are adjustment coefficients obtained by offline training based on historical data, and α∈[0.1,0.5], β∈[0.5,2.0]; The calculation of the Markov entropy of the system's internal state transition includes: slicing the system log using a sliding time window, constructing the state transition probability matrix P, and then using the formula: ; Calculate the entropy value, where Let be the steady-state probability of state i. Let be the probability of transitioning from state j to state j; for Min-Max normalization is performed to obtain normalized values. .

Citation Information

Patent Citations

  • Adaptive correlation perception unsupervised deep learning anomaly detection method

    CN115879505A