Balanced Training Data Management Method and System
Through real-time monitoring and pre-training machine learning models, we ensure that the abnormal behavior detection model is initially trained on actual data and fine-tuned on synthetic data, solving the problem of over-reliance on synthetic data, improving detection accuracy and response efficiency, and ensuring public safety.
Patent Information
- Application Number
- CN202411476942.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-22
AI Technical Summary
In the training of abnormal behavior detection model, the existing technology over-reliance on synthesis of abnormal behavior data leads to the model to misjudgment of normal behavior as abnormal behavior, increasing the risk of misreports and false alarms, and it is difficult to detect such problems in a timely manner.
By monitoring the degree of dependence of the abnormal behavior detection model on synthetic abnormal behavior data in real time, using pre-trained machine learning models to predict and optimize training strategies, ensuring that the model is initially trained on the actual abnormal behavior data, and then fine-tuning it on the synthetic abnormal behavior data to avoid over-reliance on synthetic data.
It effectively reduces the false alarm and missed rate, reduces the risk of the system triggering false emergency responses, saves emergency resources, improves the system's response efficiency in real events, and ensures public safety.
Smart Images

Figure CN119377845B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of balanced training data management, and particularly to a method and system for managing balanced training data. Background Art
[0002] A balanced training data management system is a system used to manage and optimize data distribution during the training process of a machine learning model. By analyzing the phenomenon of uneven class distribution or feature distribution in the training dataset, it uses techniques such as data augmentation, undersampling, and oversampling to adjust the proportion of the dataset, ensuring that the model can fully learn different classes or features, avoiding the model from being biased towards a certain class, and thus improving the generalization ability and prediction accuracy of the model. This system can also monitor the changes in data balance during the training process in real time and dynamically adjust the strategy to adapt to different training requirements.
[0003] In video surveillance, the amount of data for normal behaviors is much larger than that of abnormal behavior samples, which may cause the model to tend to predict normal behaviors, thereby increasing the risks of false negatives (failure to detect abnormalities) and false positives (misjudging normal behaviors as abnormalities). The balanced training data management system synthesizes abnormal behavior data (such as using generative adversarial network GAN or data augmentation techniques) to expand the quantity and diversity of abnormal samples, ensuring that the abnormal behavior detection model can more fully learn the characteristics of abnormal behaviors during training. In addition, this system can also dynamically adjust the data sampling strategy according to real-time feedback, optimize the detection threshold of the abnormal behavior monitoring model, further reduce false positives and false negatives, and thus enhance the recognition accuracy and response ability of the monitoring system to abnormal events.
[0004] When the balanced training data management system enhances the detection ability of the abnormal behavior detection model by synthesizing abnormal behavior data, if the abnormal behavior detection model overly relies on the characteristics of the synthesized abnormal behavior data for learning and ignores the diversity of actual abnormal behaviors, the existing technology usually cannot detect this problem in time. When this problem occurs, the abnormal behavior detection model may misjudge and identify normal behaviors as abnormal behaviors. For example, in banks, shopping malls or transportation hubs, ordinary customer activities may be misjudged as suspicious behaviors. Frequent false alarms may cause the system to trigger incorrect emergency responses (such as security personnel being dispatched repeatedly, equipment shutdown). Emergency resources are exhausted, and real events cannot be processed in time when they occur, further amplifying security vulnerabilities and social panic.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The objective of the present invention is to provide a method and system for managing balanced training data. By monitoring in real time the dependence of the abnormal behavior detection model on synthetic data, and using a pre-trained machine learning model to predict and optimize the training strategy, it is ensured that the abnormal behavior detection model does not overly rely on synthetic data, but rather learns more about the characteristics of actual abnormal behaviors and captures the diversity of real scenarios. A hierarchical training method is adopted, where the model is initially trained with actual abnormal behavior data and then fine-tuned with synthetic data to enhance the generalization ability and detection accuracy of the model, effectively reducing false alarms and incorrect responses, saving emergency resources, improving the response efficiency of the system in real events, and ensuring public safety. The system dynamically adjusts the data ratio and training strategy to maintain the healthy learning of the model, and corrects problems in a timely manner through real-time monitoring to enhance its robustness, enabling the model to adapt to data changes and new types of abnormal behaviors, ensuring stable operation in complex environments, and at the same time supporting long-term maintenance and upgrading to solve the problems in the above-mentioned background technology.
[0007] To achieve the above objective, the present invention provides the following technical solution: A method for managing balanced training data, comprising the following steps:
[0008] Construct a complete abnormal behavior data set, which contains all the data used for training the abnormal behavior detection model. On this basis, the data set is carefully divided into two parts: synthetic abnormal behavior data and actual abnormal behavior data;
[0009] During the training of the abnormal behavior detection model, obtain in real time the training situation of the synthetic abnormal behavior data, and extract from the real-time monitored data the key information related to the training of the synthetic abnormal behavior data;
[0010] Under the monitoring window, analyze and process the extracted key information, and input the analyzed and processed feature vectors into the pre-trained machine learning model to use the pre-trained machine learning model to predict the usage situation of the synthetic abnormal behavior data in the current training stage;
[0011] According to the prediction result of the machine learning model, classify the usage situation of the synthetic abnormal behavior data into: normal usage and over-reliance usage;
[0012] If the prediction result is normal usage, then continue to maintain the current training strategy and data usage method. At this time, the learning of the abnormal behavior detection model for the synthetic abnormal behavior data and the actual abnormal behavior data is healthy, and can effectively enhance the generalization ability and detection accuracy of the model;
[0013] For the situation determined to be over-reliance usage, readjust the composition ratio of the training data set, increase the weight of the actual abnormal behavior data, and reduce the ratio of the synthetic abnormal behavior data, so as to prompt the abnormal behavior detection model to learn more about the characteristics of the actual abnormal behaviors;
[0014] During the training process, a hierarchical training method is adopted. The abnormal behavior detection model is initially trained on actual abnormal behavior data to learn the characteristics of the real scenario. Then, it is fine-tuned on synthetic abnormal behavior data to enrich the understanding of abnormal behavior by the abnormal behavior detection model and prevent the abnormal behavior detection model from being dominated by the characteristics of synthetic abnormal behavior data.
[0015] Based on the prediction results of the machine learning model, the usage of synthetic abnormal behavior data in the current training stage is judged in real time until the abnormal behavior detection model no longer overly relies on synthetic abnormal behavior data for training.
[0016] Preferably, the data set is carefully divided into synthetic abnormal behavior data and actual abnormal behavior data. The specific steps are as follows:
[0017] Collect abnormal behavior data from multiple sources to ensure the extensiveness and coverage of the data.
[0018] Label the samples in the data set to clarify which belong to synthetic abnormal behavior data and which belong to actual abnormal behavior data.
[0019] After the data annotation is completed, in-depth feature analysis is carried out to evaluate the characteristic differences and quality between synthetic abnormal behavior data and actual abnormal behavior data.
[0020] After the data annotation and analysis are completed, according to the training requirements and data distribution, the data set is divided into a synthetic abnormal behavior data set and an actual abnormal behavior data set.
[0021] Preferably, a pre-trained machine learning model is used to predict the usage of synthetic abnormal behavior data in the current training stage. The specific steps are as follows:
[0022] During the training process of the abnormal behavior detection model, indicators related to the training situation of synthetic abnormal behavior data are obtained in real time.
[0023] From the real-time monitored data, key information related to the training of synthetic abnormal behavior data is extracted. The key information includes weight distribution, gradient information, training loss trend, and feature activation distribution.
[0024] After extracting the key information, it is input into a monitoring window for analysis and processing. The analysis and processing include feature vectorization, time series analysis, and anomaly detection.
[0025] The analyzed and processed feature vectors are input into a pre-trained machine learning model.
[0026] The pre-trained machine learning model generates a usage dependence coefficient based on the input feature vectors to quantify the dependence degree on synthetic abnormal behavior data in the current training stage.
[0027] Preferably, within the monitoring window, the usage dependence coefficient generated when predicting the usage of the synthetic abnormal behavior data in the current training stage by the machine learning model is compared with a preset reference threshold of the usage dependence coefficient, and the usage of the synthetic abnormal behavior data is classified. The specific classification process is as follows:
[0028] If the usage dependence coefficient is greater than or equal to the reference threshold of the usage dependence coefficient, the usage of the synthetic abnormal behavior data under this monitoring window is classified as over-reliance on usage;
[0029] If the usage dependence coefficient is less than the reference threshold of the usage dependence coefficient, the usage of the synthetic abnormal behavior data under this monitoring window is classified as normal usage.
[0030] Preferably, the adaptive weighted algorithm and the sample reweighting algorithm are used to readjust the composition ratio of the training data set. The specific steps are as follows:
[0031] Initialize the adaptive weighting mechanism and assign initial weights to each type of data;
[0032] During the training process, the performance of the model on the synthetic abnormal behavior data and the actual abnormal behavior data is tracked in real time through the monitoring window, and the key indicators of the abnormal behavior detection model on each type of data are collected. If the monitoring results show that the loss of the abnormal behavior detection model on the synthetic abnormal behavior data decreases rapidly, but the error or false alarm rate on the actual abnormal behavior data remains high, then the abnormal behavior data is input into a feedback loop for analyzing the contribution of each type of data to the training results of the abnormal behavior detection model and judging whether it is necessary to adjust the usage ratio of the data;
[0033] According to the feedback data collected during the training process of the abnormal behavior detection model, dynamically adjust the weights of the synthetic abnormal behavior data and the actual abnormal behavior data.
[0034] Preferably, after applying the adaptive weighting and sample reweighting strategies, continue to run the training process and perform iterative optimization according to the performance of the abnormal behavior detection model. After each round of training, evaluate the performance of the abnormal behavior detection model on the actual abnormal behavior data and the synthetic abnormal behavior data again to check whether the abnormal behavior detection model has reduced its dependence on the synthetic abnormal behavior data and improved its detection ability for the actual abnormal behavior data; if the dependence problem is alleviated, the system gradually maintains the current weights and no longer makes further adjustments; if the problem still exists, continue with a new round of weight optimization and sample reweighting.
[0035] Preferably, through the phased learning rate adjustment and weight freezing algorithms, a hierarchical training method is adopted to initially train the abnormal behavior detection model on actual abnormal behavior data and then fine-tune it on synthetic abnormal behavior data. The specific steps are as follows:
[0036] Before the start of training, initialize the abnormal behavior detection model and formulate a phased learning rate adjustment strategy;
[0037] In the first stage, the abnormal behavior detection model focuses on initially training on actual abnormal behavior data, enabling the abnormal behavior detection model to learn the key features in the real scenario. Once the abnormal behavior detection model reaches the preset performance on the actual abnormal behavior data, it enters the fine-tuning stage of the next phase;
[0038] When entering the second stage of training, adopt the weight freezing technique, that is, lock most of the trained parameters in the abnormal behavior detection model and only release the partial weights related to anomaly detection;
[0039] After the fine-tuning is completed, conduct a comprehensive performance evaluation of the abnormal behavior detection model to ensure the balanced performance of the abnormal behavior detection model on actual abnormal behavior data and synthetic abnormal behavior data.
[0040] Preferably, before the start of training, initialize the abnormal behavior detection model and formulate a phased learning rate adjustment strategy. The specific steps are as follows:
[0041] First, select the basic abnormal behavior detection model and set the initial parameters of the abnormal behavior detection model. On this basis, define two stages of training: the initial training stage and the fine-tuning stage. In each stage, adopt different learning rates to control the convergence speed of the abnormal behavior detection model, ensure that the abnormal behavior detection model is gradually optimized during training, and reduce the risk of overfitting.
[0042] Preferably, based on the prediction results of the machine learning model, the usage situation of the synthetic abnormal behavior data in the current training stage is judged in real time. The specific steps are as follows:
[0043] During the training process, generate multiple usage dependency coefficients in real time to establish an analysis set, and use the usage dependency coefficients in the analysis set to quantify the degree of dependence of the model on the synthetic abnormal behavior data;
[0044] In the constructed analysis set, calculate the mean and standard deviation of several usage dependency coefficients. The mean of the usage dependency coefficients reflects the average degree of dependence of the abnormal behavior detection model in multiple training cycles, and the standard deviation of the usage dependency coefficients is used to evaluate the volatility of the usage dependency coefficients;
[0045] The mean of the usage dependency coefficient will be compared with a preset reference threshold of the usage dependency coefficient to determine whether the abnormal behavior detection model has exceeded the allowed degree of dependency. If the mean of the usage dependency coefficient is greater than or equal to the reference threshold of the usage dependency coefficient or the standard deviation of the usage dependency coefficient is greater than or equal to the reference threshold of the standard deviation, the training strategy will be dynamically adjusted and the composition of the training data set will be re-optimized until the mean of the usage dependency coefficient is less than the reference threshold of the usage dependency coefficient and the standard deviation of the usage dependency coefficient is less than the reference threshold of the standard deviation.
[0046] A balanced training data management system, including a data set construction and division module, a real-time data monitoring module, a key information analysis and prediction module, a usage classification module, a training strategy maintenance module, a data ratio adjustment module, a hierarchical training module, and a real-time usage judgment module;
[0047] The data set construction and division module constructs a complete abnormal behavior data set, which contains all the data for training the abnormal behavior detection model. On this basis, the data set is carefully divided into two parts: synthetic abnormal behavior data and actual abnormal behavior data;
[0048] The real-time data monitoring module, during the training of the abnormal behavior monitoring model, obtains the training situation of the synthetic abnormal behavior data in real time, and extracts key information related to the training of the synthetic abnormal behavior data from the real-time monitored data;
[0049] The key information analysis and prediction module, under the monitoring window, analyzes and processes the extracted key information, and inputs the analyzed and processed feature vectors into a pre-trained machine learning model, and uses the pre-trained machine learning model to predict the usage situation of the synthetic abnormal behavior data in the current training stage;
[0050] The usage classification module classifies the usage situation of the synthetic abnormal behavior data according to the prediction results of the machine learning model: normal usage and over-reliance usage;
[0051] The training strategy maintenance module, if the prediction result is normal usage, continues to maintain the current training strategy and data usage method. At this time, the abnormal behavior detection model's learning of the synthetic abnormal behavior data and the actual abnormal behavior data is healthy, and can effectively improve the generalization ability and detection accuracy of the model;
[0052] The data ratio adjustment module, for the situation determined to be over-reliance usage, re-adjusts the composition ratio of the training data set, increases the weight of the actual abnormal behavior data, and reduces the ratio of the abnormal behavior data, so as to prompt the abnormal behavior detection model to learn more features of the actual abnormal behavior;
[0053] Hierarchical training module. During the training process, a hierarchical training method is adopted, enabling the abnormal behavior detection model to conduct preliminary training on actual abnormal behavior data to learn the basic features of the real scenario. Then, fine-tuning is performed on the synthetic abnormal behavior data to enrich the understanding of abnormal behaviors by the abnormal behavior detection model and prevent the abnormal behavior detection model from being dominated by the features of the synthetic abnormal behavior data.
[0054] Real-time usage judgment module. Based on the prediction results of the machine learning model, it judges the usage of the synthetic abnormal behavior data in the current training stage in real time until the abnormal behavior detection model no longer overly relies on the synthetic abnormal behavior data for training.
[0055] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0056] The present invention ensures that the abnormal behavior detection model does not overly rely on the synthetic abnormal behavior data by monitoring the dependence of the abnormal behavior detection model on the synthetic abnormal behavior data in real time, using the pre-trained machine learning model to predict and optimize the training strategy, enabling it to more fully learn the actual abnormal behavior features and capture the diversity of the real scenario. By adopting the hierarchical training method, first conduct preliminary training with actual abnormal behavior data and then fine-tune with synthetic abnormal behavior data to enhance the understanding of abnormal behaviors by the abnormal behavior detection model and improve its generalization ability and detection accuracy. It effectively reduces false alarms, reduces the risk of the system triggering incorrect emergency responses, saves emergency resources, and improves the response efficiency of the system in real events, ensuring public safety and stability.
[0057] The present invention dynamically adjusts the data ratio and training strategy by extracting key information related to the synthetic abnormal behavior data and using the pre-trained machine learning model for prediction, ensuring that the abnormal behavior detection model learns in a healthy state and preventing performance degradation. The real-time monitoring and optimization of the system make the training process of the abnormal behavior detection model transparent and controllable, promptly correcting potential problems and ensuring that the abnormal behavior detection model can adapt to data changes and new types of abnormal behaviors. Through iterative training and optimization, the robustness of the abnormal behavior detection model is improved, enabling it to operate stably in complex environments, not only improving the reliability and effectiveness of the system but also providing support for future maintenance and upgrades. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0059] Figure 1 It is the method flow chart of the method for managing balanced training data of the present invention.
[0060] Figure 2 This is a schematic diagram of the modules of the balance training data management system of the present invention. Detailed implementation manners
[0061] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0062] The present invention provides a Figure 1 balance training data management method as shown below, including the following steps:
[0063] Construct a complete abnormal behavior data set, which contains all the data for training the abnormal behavior detection model. On this basis, the data set is carefully divided into two parts: synthetic abnormal behavior data and actual abnormal behavior data;
[0064] The data set is carefully divided into synthetic abnormal behavior data and actual abnormal behavior data, and the specific steps are as follows:
[0065] Collect abnormal behavior data from multiple sources to ensure the extensiveness and coverage of the data;
[0066] These data may be sourced from actual surveillance videos, IoT sensors, historical records of intelligent monitoring platforms, etc. Actual abnormal behavior data may involve scenarios such as intrusion, theft, violent incidents, and abnormal operations. At the same time, data augmentation techniques or generative models (such as GAN) are used to generate synthetic abnormal behavior data to make up for the deficiencies of actual abnormal behavior data. After collection, all the data is cleaned and standardized, removing noisy data, incomplete samples, and duplicate data to ensure the consistency and reliability of the data. The purpose of this step is to lay a high-quality data foundation for subsequent division and analysis.
[0067] Label the samples in the data set to clarify which belong to synthetic abnormal behavior data and which belong to actual abnormal behavior data;
[0068] This can be done through the source information of the data, metadata analysis, and timestamps. For example, actual abnormal behavior data usually has a clear scene source (such as bank surveillance records), while synthetic abnormal behavior data will be shown as samples from data augmentation or GAN generation. Automatic analysis of data features can also be performed. For example, synthetic abnormal behavior data may have specific patterns or generated noise. After labeling, each sample is bound to its label metadata to provide a basis for subsequent classification and processing.
[0069] After the data annotation is completed, in-depth feature analysis is carried out to evaluate the characteristic differences and quality between the synthetic abnormal behavior data and the actual abnormal behavior data;
[0070] Specifically, key features of each sample can be extracted through feature engineering, such as behavior type, time dimension, spatial location, and dynamic change information of video frames, etc. By analyzing the similarities and differences of different features between the synthetic and actual abnormal behavior data, it can be ensured that the deviation between the synthetic abnormal behavior data and the actual abnormal behavior data is controlled within a reasonable range. In addition, it is also necessary to evaluate the quality, distribution, and representativeness of the data to ensure the diversity and authenticity of the actual abnormal behavior data, as well as the coverage and rationality of the synthetic abnormal behavior data. This step lays the foundation for the subsequent combination of training data.
[0071] After completing the data annotation and analysis, according to the training requirements and data distribution, the data set is divided into a synthetic abnormal behavior data set and an actual abnormal behavior data set;
[0072] A fixed ratio (such as 7:3) can be adopted or the division ratio can be dynamically determined according to the data quality to ensure that the actual abnormal behavior data occupies sufficient weight to prevent the model from over-relying on the synthetic abnormal behavior data. After the division is completed, the two parts of the data are stored in different data warehouses or partitions respectively, and a perfect data management mechanism is established, such as data version control and metadata recording. This management mechanism ensures that the model can flexibly call the synthetic and actual abnormal behavior data during training, and supports real-time update and adjustment of the division ratio, providing support for model optimization and strategy adjustment.
[0073] This division helps to clarify the degree of dependence of the model on different types of data during the training process, providing a basis for subsequent monitoring and adjustment.
[0074] During the training of the abnormal behavior detection model, the training situation of the synthetic abnormal behavior data is obtained in real time, and key information related to the training of the synthetic abnormal behavior data is extracted from the real-time monitored data;
[0075] Under the monitoring window, the extracted key information is analyzed and processed, and the analyzed feature vectors are input into the pre-trained machine learning model to predict the usage situation of the synthetic abnormal behavior data in the current training stage by using the pre-trained machine learning model;
[0076] The steps to predict the usage situation of the synthetic abnormal behavior data in the current training stage by using the pre-trained machine learning model are as follows:
[0077] During the training of the abnormal behavior detection model, indicators related to the training situation of the synthetic abnormal behavior data are obtained in real time;
[0078] The specific content to be monitored includes: the changes in the loss value, accuracy, recall, and precision of the model in each training cycle, etc. These are key performance indicators. At the same time, it is also necessary to monitor the gradient updates and training time of the model during different batches of training to ensure the transparency of the training process. By continuously tracking these indicators in the monitoring window, the system can timely detect whether the model over-focuses on the synthetic abnormal behavior data. The purpose of this step is to lay a foundation for subsequent key information extraction and dependency analysis.
[0079] Extract the key information related to the training of synthetic abnormal behavior data from the real-time monitored data. The key information includes weight distribution, gradient information, training loss trend, and feature activation distribution.
[0080] Weight distribution: Analyze whether the abnormal behavior detection model assigns abnormally high weights to certain features in the synthetic abnormal behavior data during the training process.
[0081] Gradient information: Extract the gradient changes of the abnormal behavior detection model when processing the synthetic abnormal behavior data and observe whether there is gradient explosion or disappearance.
[0082] Training loss trend: Track whether the loss function of the abnormal behavior detection model on the synthetic abnormal behavior data converges quickly, indicating that the model may be biased towards the synthetic abnormal behavior data.
[0083] Feature activation distribution: Analyze the activation of the neural network when processing the synthetic abnormal behavior data to determine whether the abnormal behavior detection model highly depends on certain specific features.
[0084] These key information provide the learning state of the model on the synthetic abnormal behavior data, ensuring that the system can accurately grasp the risk points during the training process.
[0085] After extracting the key information, input it into the monitoring window for analysis and processing. The analysis and processing include feature vectorization, time series analysis, and anomaly detection.
[0086] Feature vectorization: Convert the extracted data such as weights, gradients, and loss trends into vector format for input into subsequent machine learning models.
[0087] Time series analysis: Based on the time series data of different training batches, judge whether the learning state of the abnormal behavior detection model gradually tilts towards the synthetic abnormal behavior data over time.
[0088] Anomaly detection: Analyze whether the fluctuations of the loss value and gradient exceed the preset threshold to timely detect signs that the abnormal behavior detection model may overly depend on the synthetic abnormal behavior data.
[0089] This process ensures that the system can accurately capture potential anomalies during the training process and lays the foundation for the prediction of machine learning models.
[0090] Input the processed feature vectors into a pre-trained machine learning model (such as Random Forest, XGBoost, or neural network);
[0091] This machine learning model is specifically used to judge the degree of dependence of the model on synthetic anomaly behavior data. By analyzing the key information during training, this machine learning model can predict the dependence situation at the current stage. This step enables the system to identify problems in the training process by means of data-driven intelligent means, rather than relying solely on manual experience.
[0092] The pre-trained machine learning model generates a usage dependence coefficient based on the input feature vectors to quantify the degree of dependence on synthetic anomaly behavior data at the current training stage;
[0093] The pre-trained machine learning model (such as XGBoost, Random Forest, or neural network), after receiving the input feature vectors, based on the key information in the feature vectors (such as weight distribution, gradient change, loss trend, feature activation pattern, etc.), generates a comprehensive score by calculating the importance and contribution of each feature to the model's decision-making. These scores will go through a normalization process to map the results to the interval between 0 and 1, thus forming the usage dependence coefficient. The model uses various algorithms such as feature importance analysis, time series trend, and anomaly pattern matching to judge the degree of dependence of the model on synthetic anomaly behavior data. If the model detects that the synthetic anomaly behavior data accounts for too high a proportion in feature learning or model weights, the usage dependence coefficient will be close to 1; while if the model shows a relatively balanced use of synthetic and actual anomaly behavior data, the usage dependence coefficient will approach 0. The generation of the usage dependence coefficient can quantitatively reflect the dependence state of the model at the current training stage and provide an accurate basis for the dynamic adjustment of the training strategy.
[0094] The pre-trained machine learning model refers to a model that has been pre-trained in advance and is specifically used for prediction or judgment on new tasks or data. The key to this model is that it has completed preliminary training on a dataset highly relevant to the target task and has learned how to extract key features from the input data and perform effective classification or regression analysis. In the anomaly behavior detection task, the pre-trained machine learning model (such as Random Forest, XGBoost, Support Vector Machine, or deep neural network) is designed to analyze specific feature vectors and predict the degree of dependence of the model on synthetic anomaly behavior data during training. By pre-training on historical training data and feature data, this model can quickly generate effective results when new data is input without the need for re-training, thus accelerating the decision-making process and improving the efficiency of the system.
[0095] The advantage of this model is that it can utilize historical data and experience to learn in advance to identify abnormal patterns and trends in model training. For example, when monitoring the usage dependence of synthetic abnormal behavior data, the model will pay attention to details such as weight deviation, gradient abnormality, and loss convergence speed, and compare this information with the learned patterns. If it is found that the current training state conforms to the typical "over-dependence" pattern, the model will give a relatively high usage dependence coefficient. This pre-training mechanism ensures the real-time monitoring ability of the system, without having to wait until the training is completed to conduct performance evaluation, thereby reducing potential risks and ensuring that the model training process is always within a controllable range.
[0096] When constructing a pre-trained machine learning model, supervised learning methods are usually adopted. The model will be initially trained on a large amount of historical data (such as the past training records of the monitoring system) to learn how to identify the relevance and importance between different features. For example, when constructing a random forest model for dependence analysis, a large amount of training data with known labels (labeled as over-dependence or normal usage) will be input first. Through multiple trainings and optimizations of decision trees, the model can learn how to make accurate predictions based on input features. For the XGBoost model, the boosting method will be adopted to gradually reduce the prediction error and fully optimize the performance of the model during the training stage.
[0097] In specific applications, when the feature vectors extracted by the real-time system are input into the pre-trained model, the model will quickly complete the calculation and generate a usage dependence coefficient. The usage dependence coefficient reflects the degree of dependence on synthetic abnormal behavior data in the current model training. The value ranges from 0 to 1. A value close to 1 represents a relatively high degree of dependence, and a value close to 0 indicates a low degree of dependence. Through the usage dependence coefficient, the system can quantify and monitor potential problems during the training process, and take measures such as hybrid sampling or strategy adjustment in a timely manner to prevent the imbalance of the abnormal behavior monitoring model. The use of the pre-trained model not only improves the efficiency of monitoring and prediction, but also significantly reduces resource consumption, and enhances the system's automated processing ability and model robustness.
[0098] According to the prediction results of the machine learning model, classify the usage of synthetic abnormal behavior data into: normal usage and over-dependence usage;
[0099] Within the monitoring window, compare and analyze the usage dependence coefficient generated when predicting the usage of synthetic abnormal behavior data in the current training stage by the machine learning model with the pre-set reference threshold of the usage dependence coefficient, and classify the usage of synthetic abnormal behavior data. The specific classification process is as follows:
[0100] If the usage dependency coefficient is greater than or equal to the usage dependency coefficient reference threshold, the usage of the synthetic abnormal behavior data in this monitoring window is classified as over-reliance on usage;
[0101] If the usage dependency coefficient is less than the usage dependency coefficient reference threshold, the usage of the synthetic abnormal behavior data in this monitoring window is classified as normal usage.
[0102] Specifically, normal usage means that the abnormal behavior detection model learns synthetic abnormal behavior data and actual abnormal behavior data in a balanced manner, and no adjustment is required during the training process.
[0103] Over-reliance on usage means that the abnormal behavior detection model pays too much attention to synthetic abnormal behavior data and may ignore the characteristics of actual abnormal behavior data.
[0104] If the prediction result is normal usage, the current training strategy and data usage method are continued. At this time, the abnormal behavior detection model learns synthetic abnormal behavior data and actual abnormal behavior data in a healthy manner, which can effectively improve the generalization ability and detection accuracy of the model;
[0105] Maintaining the existing strategy can ensure the stable training of the model, avoid the disturbance caused by unnecessary adjustments, and save resources and time costs at the same time.
[0106] For the situation judged as over-reliance on usage, readjust the composition ratio of the training data set, increase the weight of actual abnormal behavior data, and reduce the proportion of synthetic abnormal behavior data, so as to prompt the abnormal behavior detection model to learn more characteristics of actual abnormal behavior;
[0107] Adopt the adaptive weighted algorithm and sample reweighting algorithm to readjust the composition ratio of the training data set. The specific steps are as follows:
[0108] Initialize the adaptive weighting mechanism and assign initial weights to each type of data (synthetic abnormal behavior data and actual abnormal behavior data);
[0109] For example, the initial weight of actual abnormal behavior data is set to W_actual = 0.5, and the initial weight of synthetic abnormal behavior data is set to W_synthetic = 0.5. These weights will be dynamically adjusted according to the feedback during the training process. In addition, a weight adjustment strategy needs to be defined, such as based on the performance of the loss function or the accuracy of the abnormal behavior detection model on actual abnormal behavior data. Through this initialization, the system establishes a benchmark point for weight adjustment, so as to gradually strengthen the influence of actual abnormal behavior data according to the learning situation of the abnormal behavior detection model.
[0110] During the training process, the performance of the model on synthetic abnormal behavior data and actual abnormal behavior data is tracked in real time through a monitoring window, and key metrics such as the loss, precision, and recall rate of the abnormal behavior monitoring model on each type of data are collected. If the monitoring results show that the loss of the abnormal behavior detection model on synthetic abnormal behavior data decreases rapidly, but the error or false negative rate on actual abnormal behavior data remains high, it indicates that the abnormal behavior detection model may be overly dependent on the features of synthetic abnormal behavior data. The system inputs this monitoring data into a feedback loop to analyze the contribution of each type of data to the training results of the abnormal behavior detection model and determine whether it is necessary to adjust the usage ratio of the data.
[0111] Dynamically adjust the weights of synthetic abnormal behavior data and actual abnormal behavior data according to the feedback data collected during the training process of the abnormal behavior detection model;
[0112] For example, if the system detects an increase in the dependence of the abnormal behavior detection model on synthetic abnormal behavior data, the weight W_actual of the actual abnormal behavior data will be increased (e.g., adjusted from 0.5 to 0.7), while the weight W_synthetic of the synthetic abnormal behavior data will be decreased (e.g., adjusted from 0.5 to 0.3). This adaptive weighting mechanism ensures that the abnormal behavior detection model gradually pays more attention to the features of actual abnormal behavior data. In addition, the system will perform sample reweighting according to the latest weight ratio, that is, dynamically select data in the training batch according to the adjusted weight ratio. This adjustment can be achieved by random sampling in each batch, so that the abnormal behavior detection model is always consistent with the latest weight distribution during training.
[0113] After applying the adaptive weighting and sample reweighting strategies, continue to run the training process and perform iterative optimization according to the performance of the abnormal behavior detection model. After each round of training, the system re-evaluates the performance of the abnormal behavior detection model on actual abnormal behavior data and synthetic abnormal behavior data to check whether the abnormal behavior detection model has reduced its dependence on synthetic abnormal behavior data and improved its detection ability for actual abnormal behavior data; if the dependence problem is alleviated, the system will gradually maintain the current weights and no longer make further adjustments; if the problem still exists, continue with a new round of weight optimization and sample reweighting;
[0114] This iterative process ensures that the abnormal behavior detection model will not be unbalanced due to short-term adjustments and gradually improves its learning ability for actual abnormal behavior features during long-term training, ensuring the generalization and detection ability of the abnormal behavior detection model.
[0115] During the training process, a hierarchical training method is adopted. The abnormal behavior detection model is initially trained on actual abnormal behavior data to learn the basic features of the real scenario. Then, it is fine-tuned on synthetic abnormal behavior data to enrich the understanding of abnormal behavior by the abnormal behavior detection model and prevent the abnormal behavior detection model from being dominated by the features of synthetic abnormal behavior data.
[0116] Through the phased learning rate adjustment and weight freezing algorithms, a hierarchical training method is adopted. The abnormal behavior detection model is initially trained on actual abnormal behavior data and then fine-tuned on synthetic abnormal behavior data. The specific steps are as follows:
[0117] Before the training starts, the abnormal behavior detection model is initialized, and a phased learning rate adjustment strategy is formulated. First, an appropriate basic abnormal behavior detection model (such as a convolutional neural network or a Transformer abnormal behavior detection model) is selected, and the initial parameters of the abnormal behavior detection model are set. On this basis, two training stages are defined: the initial training stage (training on actual abnormal data) and the fine-tuning stage (fine-tuning on synthetic abnormal behavior data). In each stage, different learning rates are used to control the convergence speed of the abnormal behavior detection model - a higher learning rate (such as 0.01 to 0.001) is used in the initial training stage to ensure that the abnormal behavior detection model can quickly learn the key features in the actual abnormal behavior data; the learning rate is reduced in the fine-tuning stage (such as 0.001 to 0.0001) to avoid the abnormal behavior detection model losing the basic features it has learned during adjustment. This strategy ensures that the abnormal behavior detection model is gradually optimized during the training process, reduces the risk of overfitting, and improves the adaptability of the abnormal behavior detection model.
[0118] In the first stage, the abnormal behavior detection model focuses on initial training on actual abnormal behavior data. The goal of this stage is to enable the abnormal behavior detection model to learn the key features in the real scenario, such as abnormal behavior patterns within a specific time, spatial location changes, or behavior trajectories in surveillance videos. With a higher learning rate, the abnormal behavior detection model can quickly capture the core patterns in the actual abnormal behavior data and achieve better optimization results on the loss function. At the same time, the system needs to monitor the performance of the abnormal behavior detection model in real time, such as the loss decline trend, precision, and recall rate, to ensure the stability of the training process of the abnormal behavior detection model on actual abnormal behavior data. Once the abnormal behavior detection model reaches the preset performance on the actual abnormal behavior data (such as the accuracy and loss value of the validation set reaching a certain threshold), it can enter the next stage of fine-tuning.
[0119] When entering the second stage of training, the weight freezing technique is adopted, that is, most of the well-trained parameters in the anomaly behavior detection model (such as the parameters of the convolutional layer or Transformer encoder) are locked, and only the partial weights related to anomaly detection (such as the last few layers or the classifier layer) are released. At this time, the learning rate is reduced to a small value (such as 0.0001) to ensure that the parameters of the anomaly behavior detection model will not deviate significantly due to fine-tuning. In this stage, the anomaly behavior detection model is fine-tuned on the synthetic anomaly behavior data. Through the diversity of the synthetic anomaly behavior data, the understanding of new or rare anomaly behaviors is further enriched. This training strategy ensures that the anomaly behavior detection model will not be dominated by the synthetic anomaly behavior data, but through the supplement of the synthetic anomaly behavior data, it can comprehensively master various anomaly patterns.
[0120] After the fine-tuning is completed, a comprehensive performance evaluation of the anomaly behavior detection model is required to ensure its balanced performance on the actual anomaly behavior data and the synthetic anomaly behavior data. This includes evaluating the precision, recall rate, F1 score of the anomaly behavior detection model on the test set and the recognition ability of the anomaly behavior detection model for different types of anomaly behaviors. If the evaluation results show that the anomaly behavior detection model performs well on the synthetic anomaly behavior data but there are false positives or false negatives on the actual anomaly behavior data, then the anomaly behavior detection model needs to be optimized by further fine-tuning or adjusting the data weights. If the performance of the anomaly behavior detection model meets the expectations, the current anomaly behavior detection model can be frozen and deployed in the actual scenario. In addition, the system can periodically retrain and fine-tune the anomaly behavior detection model to cope with new types of anomaly behaviors that appear in the actual scenario, ensuring the robustness and generalization ability of the anomaly behavior detection model in long-term use.
[0121] Through the above four steps, the hierarchical training method can ensure that the anomaly behavior detection model can not only learn the key features in the actual anomaly behavior data but also expand the understanding of diverse anomaly behaviors from the synthetic anomaly behavior data. The combination of this phased learning rate adjustment and weight freezing not only improves the detection ability and stability of the anomaly behavior detection model but also effectively prevents the anomaly behavior detection model from over-relying on a certain type of data, ensuring its long-term adaptability and reliability in complex environments.
[0122] Based on the prediction results of the machine learning model, the usage of the synthetic anomaly behavior data in the current training stage is judged in real time until the anomaly behavior detection model no longer overly relies on the synthetic anomaly behavior data for training;
[0123] Based on the prediction results of the machine learning model, the usage of the synthetic anomaly behavior data in the current training stage is judged in real time. The specific steps are as follows:
[0124] During the training process, multiple analysis sets are generated in real time using dependence coefficients. The dependence coefficients within the analysis sets are used to quantify the degree of dependence of the model on the synthetic abnormal behavior data;
[0125] After each batch of training is completed, based on the prediction results of the pre-trained machine learning model, the dependence coefficients for the current training stage are calculated (e.g., close to 1 indicates high dependence, close to 0 indicates low dependence). These dependence coefficients are collected into an analysis set and continuously updated. The construction of the analysis set ensures that the system can accurately grasp the data dependence trend of the abnormal behavior detection model during the training process based on the dependence coefficients of multiple rounds of training data. This data accumulation provides rich samples for subsequent analysis and avoids misjudgments caused by the contingency of a single training.
[0126] In the constructed analysis set, the mean and standard deviation of several dependence coefficients are calculated. The mean of the dependence coefficients reflects the average degree of dependence of the abnormal behavior detection model over multiple training cycles and can indicate whether the abnormal behavior detection model as a whole has an excessive dependence on the synthetic abnormal behavior data; while the standard deviation of the dependence coefficients is used to evaluate the volatility of the dependence coefficients, indicating the stability of the dependence of the abnormal behavior detection model on the synthetic abnormal behavior data in different training batches. If the standard deviation of the dependence coefficients is large, it means that the abnormal behavior detection model has an excessive and unstable dependence on the synthetic abnormal behavior data in some stages and the training strategy needs to be further adjusted; if the standard deviation of the dependence coefficients is small, it indicates that the dependence trend of the abnormal behavior detection model on the data is relatively consistent. This calculation step ensures that the system can accurately grasp the dependence state of the abnormal behavior detection model and discover potential problems through volatility analysis.
[0127] The mean of the dependence coefficients is compared with a preset reference threshold of the dependence coefficient to determine whether the abnormal behavior detection model has exceeded the allowed degree of dependence. At the same time, the standard deviation of the dependence coefficients is compared with a reference threshold of the standard deviation to evaluate the stability of the model's dependence. If the mean of the dependence coefficients is greater than or equal to the reference threshold of the dependence coefficient, or the standard deviation of the dependence coefficients is greater than or equal to the reference threshold of the standard deviation, it indicates that the abnormal behavior detection model still has an excessive dependence on the synthetic abnormal behavior data, and the system will dynamically adjust the training strategy, such as reducing the proportion of the synthetic abnormal behavior data, increasing the weight of the actual abnormal behavior data, or switching to a mixed sampling mode to re-optimize the composition of the training data set. The system will repeatedly execute this judgment and adjustment process until the mean of the dependence coefficients is less than the reference threshold of the dependence coefficient and the standard deviation of the dependence coefficients is less than the reference threshold of the standard deviation, indicating that the abnormal behavior detection model no longer depends on the synthetic abnormal behavior data and the training enters an ideal state. At this time, the abnormal behavior detection model has successfully achieved sufficient learning of the actual abnormal behavior characteristics and has good robustness and generalization ability.
[0128] The present invention monitors in real time the degree of dependence of the abnormal behavior detection model on the synthetic abnormal behavior data, and uses a pre-trained machine learning model to predict and adjust the training strategy, ensuring that the abnormal behavior detection model does not overly rely on the synthetic abnormal behavior data, enabling the model to more fully learn the characteristics of the actual abnormal behavior data and capture the diversity and complexity in the real scenario. By increasing the weight of the actual abnormal behavior data, the understanding of the real abnormal behavior by the model during training is strengthened, avoiding the deviation caused by overly relying on the synthetic abnormal behavior data. In training, a hierarchical training method is adopted. First, preliminary training is carried out on the actual abnormal behavior data, and then fine-tuning is performed on the synthetic abnormal behavior data to further enrich the model's understanding of abnormal behavior. This strategy improves the generalization ability and detection accuracy of the model, effectively reducing the risk of misjudging normal behavior as abnormal behavior. Reducing false alarms not only reduces the possibility of the system triggering incorrect emergency responses, saves emergency resources, but also improves the response efficiency of the system when real events occur, maintaining public safety and social stability.
[0129] The present invention extracts the key information related to the synthetic abnormal behavior data during the training process and inputs it into the pre-trained machine learning model for analysis and prediction, enabling the system to judge in real time the degree of dependence of the model on the synthetic abnormal behavior data. According to the prediction results, the composition ratio of the training data set and the training strategy are dynamically adjusted to ensure that the model is always learning in a healthy state. This dynamic monitoring and adjustment mechanism makes the model training process more transparent and controllable, promptly discovers and corrects potential problems, and prevents the degradation of the model performance. Continuous monitoring and optimization enable the abnormal behavior detection model to quickly adapt and maintain a high level of detection performance when facing data distribution changes or new types of abnormal behaviors. Through iterative training and optimization, the robustness of the model is enhanced, enabling it to operate stably in a complex and changing real environment. This not only improves the reliability and effectiveness of the system, but also provides strong support for long-term model maintenance and upgrade, ensuring that the system continues to play a role in future development.
[0130] The present invention provides a Figure 2 balanced training data management system as shown, including a data set construction and division module, a real-time data monitoring module, a key information analysis and prediction module, a usage classification module, a training strategy maintenance module, a data ratio adjustment module, a hierarchical training module, and a real-time usage judgment module;
[0131] The data set construction and division module constructs a complete abnormal behavior data set, which contains all the data for training the abnormal behavior detection model. On this basis, the data set is carefully divided into two parts: synthetic abnormal behavior data and actual abnormal behavior data;
[0132] The real-time data monitoring module obtains the training status of the synthetic abnormal behavior data in real time during the training of the abnormal behavior monitoring model, and extracts key information related to the training of the synthetic abnormal behavior data from the real-time monitored data;
[0133] The key information analysis and prediction module analyzes and processes the extracted key information under the monitoring window, and inputs the analyzed feature vectors into the pre-trained machine learning model, and uses the pre-trained machine learning model to predict the usage of the synthetic abnormal behavior data in the current training stage;
[0134] The usage classification module classifies the usage of the synthetic abnormal behavior data according to the prediction results of the machine learning model: normal usage and over-reliance usage;
[0135] The training strategy maintenance module, if the prediction result is normal usage, continues to maintain the current training strategy and data usage method. At this time, the abnormal behavior detection model's learning of the synthetic abnormal behavior data and the actual abnormal behavior data is healthy, which can effectively improve the generalization ability and detection accuracy of the model;
[0136] The data ratio adjustment module, for the situation determined to be over-reliance usage, re-adjusts the composition ratio of the training data set, increases the weight of the actual abnormal behavior data, and reduces the ratio of the synthetic abnormal behavior data, so as to prompt the abnormal behavior detection model to learn more features of the actual abnormal behavior;
[0137] The hierarchical training module, during the training process, adopts a hierarchical training method, allowing the abnormal behavior detection model to conduct preliminary training on the actual abnormal behavior data to learn the basic features of the real scenario; then, fine-tune on the synthetic abnormal behavior data to enrich the abnormal behavior detection model's understanding of abnormal behavior and prevent the abnormal behavior detection model from being dominated by the features of the synthetic abnormal behavior data;
[0138] The real-time usage judgment module judges the usage of the synthetic abnormal behavior data in the current training stage in real time based on the prediction results of the machine learning model until the abnormal behavior detection model no longer overly relies on the synthetic abnormal behavior data for training;
[0139] The balanced training data management method provided by the embodiments of the present invention is implemented through the above-mentioned balanced training data management system. The specific methods and processes of the balanced training data management system can be seen in the embodiments of the above-mentioned balanced training data management method, which will not be elaborated here.
[0140] Only some exemplary embodiments of the present invention have been described by way of illustration. Without doubt, for those of ordinary skill in the art, various different ways can be used to modify the described embodiments without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A balanced training data management method, characterized in that: The following steps are involved: Build a complete abnormal behavior dataset, which contains all the data used to train the abnormal behavior detection model. On this basis, divide the dataset into two parts: synthetic abnormal behavior data and actual abnormal behavior data. During the training of the abnormal behavior detection model, the training status of the synthetic abnormal behavior data is obtained in real time, and key information related to the training of the synthetic abnormal behavior data is extracted from the real-time monitored data, where the key information includes weight distribution, gradient information, training loss trend, and feature activation distribution; In the monitoring window, the extracted key information is analyzed and processed, and the feature vector after analysis and processing is input into the pre-trained machine learning model, and the pre-trained machine learning model is used to predict the usage of the synthetic abnormal behavior data in the current training stage; According to the prediction results of the machine learning model, the usage of synthetic abnormal behavior data is classified into normal use and over-reliance; If the prediction result is normal use, continue to maintain the current training strategy and data usage method. At this time, the abnormal behavior detection model is healthy in learning the synthetic abnormal behavior data and the actual abnormal behavior data, which can effectively improve the generalization ability and detection accuracy of the model. For cases that are judged to be over-reliant, the composition ratio of the training data set is readjusted to increase the weight of the actual abnormal behavior data and reduce the proportion of abnormal behavior data, so as to encourage the abnormal behavior detection model to learn more about the characteristics of the actual abnormal behavior; During the training process, a hierarchical training method is used to allow the abnormal behavior detection model to be initially trained on actual abnormal behavior data to learn the characteristics of real scenarios. Then, fine-tuning is performed on synthetic abnormal behavior data to enrich the abnormal behavior detection model's understanding of abnormal behavior and prevent the abnormal behavior detection model from being dominated by the characteristics of synthetic abnormal behavior data. Based on the prediction results of the machine learning model, the usage of synthetic abnormal behavior data in the current training phase is judged in real time until the abnormal behavior detection model no longer relies too much on synthetic abnormal behavior data for training.
2. The method for managing balanced training data according to claim 1, characterized in that: The data set is divided into synthetic abnormal behavior data and actual abnormal behavior data in detail. The specific steps are as follows: Collect abnormal behavior data from multiple sources to ensure the breadth and coverage of the data; Label the samples in the data set to identify which are synthetic abnormal behavior data and which are actual abnormal behavior data; After the data is labeled, in-depth feature analysis is performed to evaluate the differences and quality of the characteristics of the synthetic abnormal behavior data and the actual abnormal behavior data; After completing data annotation and analysis, the dataset is divided into synthetic abnormal behavior dataset and actual abnormal behavior dataset according to training requirements and data distribution.
3. The method for managing balanced training data according to claim 1, characterized in that: Use the pre-trained machine learning model to predict the usage of synthetic abnormal behavior data in the current training phase. The specific steps are as follows: During the training process of the abnormal behavior detection model, indicators related to the training status of synthetic abnormal behavior data are obtained in real time; Extract key information related to synthetic abnormal behavior data training from real-time monitoring data; After extracting key information, it is input into the monitoring window for analysis and processing, which includes feature vectorization, time series analysis, and anomaly detection; Input the analyzed feature vector into a pre-trained machine learning model; The pre-trained machine learning model generates a usage dependency coefficient based on the input feature vector, which is used to quantify the degree of dependence of the current training stage on the synthetic abnormal behavior data.
4. The method for managing balanced training data according to claim 3, characterized in that: In the monitoring window, the usage dependency coefficient generated when the machine learning model predicts the usage of the synthetic abnormal behavior data in the current training phase is compared and analyzed with the pre-set usage dependency coefficient reference threshold, and the usage of the synthetic abnormal behavior data is classified. The specific classification process is as follows: If the usage dependency coefficient is greater than or equal to the usage dependency coefficient reference threshold, the usage of the synthetic abnormal behavior data in the monitoring window is classified as excessive dependency usage; If the usage dependency coefficient is less than the usage dependency coefficient reference threshold, the usage of the synthetic abnormal behavior data in the monitoring window is classified as normal usage.
5. The method for managing balanced training data according to claim 1, characterized in that: Adopt adaptive weighting algorithm and sample reweighting algorithm to readjust the composition ratio of training data set. The specific steps are as follows: Initialize the adaptive weighting mechanism and assign initial weights to each type of data; During the training process, the performance of the model on the synthetic abnormal behavior data and the actual abnormal behavior data is tracked in real time through the monitoring window, and the key indicators of the abnormal behavior detection model on each type of data are collected. If the monitoring results show that the loss of the abnormal behavior detection model on the synthetic abnormal behavior data is rapidly reduced, but the error or omission rate on the actual abnormal behavior data remains high, the abnormal behavior data is input into a feedback loop to analyze the contribution of each type of data to the training results of the abnormal behavior detection model and determine whether the data usage ratio needs to be adjusted; According to the feedback data collected during the training of the abnormal behavior detection model, the weights of the synthetic abnormal behavior data and the actual abnormal behavior data are dynamically adjusted.
6. The method for managing balanced training data according to claim 5, characterized in that: After applying the adaptive weighting and sample reweighting strategies, continue to run the training process and perform iterative optimization based on the performance of the abnormal behavior detection model. After each round of training, evaluate the performance of the abnormal behavior detection model on the actual abnormal behavior data and synthetic abnormal behavior data again to check whether the abnormal behavior detection model has reduced its dependence on synthetic abnormal behavior data and improved its detection capability for actual abnormal behavior data. If the dependency problem is alleviated, the system gradually maintains the current weights without further adjustment. If the problem still exists, continue with a new round of weight optimization and sample reweighting.
7. The method for managing balanced training data according to claim 1, characterized in that: Through the phased learning rate adjustment and weight freezing algorithm, a layered training method is adopted to allow the abnormal behavior detection model to be initially trained on the actual abnormal behavior data, and then fine-tuned on the synthetic abnormal behavior data. The specific steps are as follows: Before training begins, the abnormal behavior detection model is initialized and a phased learning rate adjustment strategy is formulated; In the first stage, the abnormal behavior detection model is trained on actual abnormal behavior data to enable it to learn key features in real scenarios. Once the abnormal behavior detection model reaches the preset performance on the actual abnormal behavior data, it will enter the next stage of fine-tuning. When entering the second stage of training, the weight freezing technology is used, that is, most of the trained parameters in the abnormal behavior detection model are locked, and only some weights related to anomaly detection are released; After fine-tuning is completed, a comprehensive performance evaluation of the abnormal behavior detection model is performed to ensure that the abnormal behavior detection model performs balanced on actual abnormal behavior data and synthetic abnormal behavior data.
8. The method for managing balanced training data according to claim 7, characterized in that: Before training begins, initialize the abnormal behavior detection model and formulate a phased learning rate adjustment strategy. The specific steps are as follows: First, a basic abnormal behavior detection model is selected and the initial parameters of the abnormal behavior detection model are set. On this basis, two stages of training are defined: the preliminary training stage and the fine-tuning stage. In each stage, different learning rates are used to control the convergence speed of the abnormal behavior detection model to ensure that the abnormal behavior detection model is gradually optimized during the training process and reduce the risk of overfitting.
9. The method for managing balanced training data according to claim 4, characterized in that: Based on the prediction results of the machine learning model, the usage of synthetic abnormal behavior data in the current training phase is judged in real time. The specific steps are as follows: During the training process, multiple usage dependency coefficients are generated in real time to establish an analysis set. The usage dependency coefficients in the analysis set are used to quantify the degree of dependence of the model on the synthetic abnormal behavior data. In the constructed analysis set, the mean and standard deviation of several usage dependency coefficients are calculated. The mean of the usage dependency coefficient reflects the average degree of dependency of the abnormal behavior detection model in multiple training cycles, and the standard deviation of the usage dependency coefficient is used to evaluate the volatility of the usage dependency coefficient. The mean usage dependency coefficient is compared with the preset usage dependency coefficient reference threshold to determine whether the abnormal behavior detection model exceeds the allowed degree of dependency. If the mean usage dependency coefficient is greater than or equal to the usage dependency coefficient reference threshold or the standard deviation of the usage dependency coefficient is greater than or equal to the standard deviation reference threshold, the training strategy is dynamically adjusted and the composition of the training data set is re-optimized until the mean usage dependency coefficient is less than the usage dependency coefficient reference threshold and the standard deviation of the usage dependency coefficient is less than the standard deviation reference threshold.
10. A balanced training data management system, used to implement the balanced training data management method according to any one of claims 1 to 9, characterized in that: It includes data set construction and division module, real-time data monitoring module, key information analysis and prediction module, usage classification module, training strategy maintenance module, data ratio adjustment module, hierarchical training module and real-time usage judgment module; The dataset construction and division module builds a complete abnormal behavior dataset, which contains all the data used to train the abnormal behavior detection model. On this basis, the dataset is carefully divided into two parts: synthetic abnormal behavior data and actual abnormal behavior data. The real-time data monitoring module obtains the training status of the synthetic abnormal behavior data in real time during the training of the abnormal behavior monitoring model, and extracts key information related to the training of the synthetic abnormal behavior data from the real-time monitoring data; The key information analysis and prediction module analyzes and processes the extracted key information in the monitoring window, and inputs the analyzed and processed feature vectors into the pre-trained machine learning model, and uses the pre-trained machine learning model to predict the usage of the synthetic abnormal behavior data in the current training stage; The usage classification module classifies the usage of the synthetic abnormal behavior data into normal usage and over-reliance usage according to the prediction results of the machine learning model; The training strategy maintenance module, if the prediction result is normal use, continues to maintain the current training strategy and data usage. At this time, the abnormal behavior detection model is healthy in learning the synthetic abnormal behavior data and the actual abnormal behavior data, which can effectively improve the generalization ability and detection accuracy of the model; The data ratio adjustment module readjusts the composition ratio of the training data set for situations that are judged to be over-reliant, increases the weight of actual abnormal behavior data, and reduces the proportion of abnormal behavior data, so as to encourage the abnormal behavior detection model to learn more about the characteristics of actual abnormal behavior; Hierarchical training module: During the training process, a hierarchical training method is used to allow the abnormal behavior detection model to perform preliminary training on actual abnormal behavior data and learn the characteristics of real scenarios; Then, fine-tune the model on the synthetic abnormal behavior data to enrich the abnormal behavior detection model’s understanding of abnormal behavior and prevent the abnormal behavior detection model from being dominated by the characteristics of the synthetic abnormal behavior data. The real-time usage judgment module judges the usage of the synthetic abnormal behavior data in the current training phase in real time based on the prediction results of the machine learning model, until the abnormal behavior detection model no longer relies too much on the synthetic abnormal behavior data for training.
Citation Information
Patent Citations
Prediction method for unbalanced data set based on isolated forest learning
CN112070125A
Image model training method, electronic equipment, roadside equipment and cloud control platform
CN113420792A