Computer intelligent service management method based on network big data
Through multi-source data monitoring and feature analysis, the tolerance range for abnormal judgment is dynamically adjusted, the problem of camouflage data pollution is solved, the detection accuracy and defense capabilities of the computer intelligent service management system are improved, and the stability and security of the system are ensured.
Patent Information
- Application Number
- CN202510398597.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, it is difficult for computer intelligent service management systems to effectively identify and defend against disguised data pollution, resulting in model learning deviations, intelligent decision-making inaccurate, and affect service quality and user experience.
Through real-time monitoring of multi-source data, trend analysis of characteristic indicators and residual fluctuation perception of model, a closed-loop data quality control process is constructed, the tolerance range of abnormal judgment is dynamically adjusted, detection sensitivity is improved, and the confidence management is used to refuse to disguise data into the database.
It improves the detection accuracy and defense capabilities of camouflage data, ensures the security and stability of model training and decision-making, enhances the robustness and adaptability of the system, and prevents the long-term impact of camouflage data pollution on the system.
Smart Images

Figure CN120296319A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a computer intelligent service management method based on network big data. Background Art
[0002] "Computer intelligent service management based on network big data" refers to the process of using large-scale, multi-source, heterogeneous massive data (i.e., network big data) distributed in the Internet and various network environments, through technical means such as big data analysis, artificial intelligence, and machine learning, to perform intelligent modeling, analysis, decision-making, and optimization management on various types of information such as user behavior, business processes, resource allocation, and service quality involved in the computer service system. Its core lies in the efficient collection, analysis, and mining of network big data, automatically perceiving user needs, dynamically optimizing service strategies, improving service response speed and service quality, and at the same time being able to achieve intelligent monitoring and exception handling of the service process, ultimately realizing the intelligentization, self-adaptation, and efficient operation of the computer service system.
[0003] The prior art has the following deficiencies: In the existing computer intelligent service management process based on network big data, the system usually already has basic data security protection and anomaly detection mechanisms, which can identify and process common abnormal data and attack behaviors to a certain extent. However, with the continuous evolution of attack means, attackers can inject highly concealed and misleading camouflage data during the network data transmission process. This type of data is highly similar to normal business data in form, structure, and statistical characteristics, and can bypass the existing anomaly detection mechanisms, and then be wrongly collected by the system and used in the training and reasoning processes of intelligent algorithms. Due to the lack of self-adaptive recognition ability of the anomaly detection and defense mechanisms in the prior art for complex and dynamically changing camouflage data pollution behaviors, when the system faces such attacks, it is still prone to problems such as model learning deviation and inaccurate intelligent decision-making results, and in severe cases, it may even lead to a series of systemic consequences such as chaotic service resource scheduling, degraded service quality, and damaged user experience.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The object of the present invention is to provide a computer intelligent service management method based on network big data. By means of real-time monitoring of multi-source data, trend analysis of characteristic indicators, perception of model residual fluctuations and dynamic threshold regulation, a closed-loop data quality control process is constructed, effectively improving the detection sensitivity and discrimination accuracy of abnormal data. Through confidence management, hierarchical processing and storage control of suspicious data are realized, and it is possible to efficiently identify and dynamically defend against camouflaged data pollution, solving the problems in the prior art of lacking adaptability and being difficult to cope with mimetic and persistent camouflaged data, ensuring the security and stability of model training and decision-making, so as to solve the problems in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: A computer intelligent service management method based on network big data, comprising the following steps:
[0007] Through real-time monitoring of network data streams, business logs, data access processes and model feedback, obtain the original data information of network data characteristics;
[0008] Preprocess the collected original feature information, and construct a structured feature data set based on the preprocessed data, providing a standardized and structured input data basis for subsequent feature extraction, model training and intelligent evaluation;
[0009] Based on the data set, extract characteristic indicators reflecting data trend and persistent deviation through feature engineering technology, and comprehensively analyze the extracted characteristic indicators to complete the quantification of data camouflage characteristics;
[0010] Input the comprehensively analyzed characteristic indicators as feature vectors into a pre-trained machine learning model, and use the machine learning model to conduct intelligent risk assessment on the current network data to determine whether there is a risk of camouflaged data pollution in the current network data;
[0011] When the machine learning model determines that there is a risk of camouflaged data pollution in the current data, dynamically narrow the tolerance interval for abnormal determination, improve the detection sensitivity of abnormal data, and at the same time, reduce the confidence level of the abnormal data determined to be camouflaged. When the confidence level is lower than the preset threshold, reject the data from being stored in the database to prevent camouflaged data from polluting the subsequent model training and inference processes.
[0012] Preferably, the specific steps for obtaining the original data information of network data characteristics are as follows:
[0013] First, through real-time monitoring of the network data stream, capture network communication behavior data including the transmission path, access frequency, request time, and traffic characteristics of data packets;
[0014] Secondly, monitor the business logs of the monitoring system and extract the business activity records generated by users in the service call, resource request, and operation behavior links;
[0015] Subsequently, monitor the data access process and collect the structural information, transmission status, and access sources of the data during the access, distribution, aggregation, and storage processes;
[0016] Finally, monitor the input and output of the machine learning model in real time, record the confidence level of the model for network data, prediction deviation, and model residual feedback information, and form model response features reflecting the data deviation trend.
[0017] Preferably, based on the data set, extract feature indicators reflecting the trend and persistent deviation of the data through feature engineering techniques. The extracted indicators include the activation frequency of low-frequency patterns and the volatility of model output residuals within a time window. The activation frequency of low-frequency patterns and the volatility of model output residuals within a time window are comprehensively analyzed under the detection window to generate a reference value for low-frequency pattern activation and a reference value for model residual fluctuation respectively. The data camouflage features are quantified through the reference value for low-frequency pattern activation and the reference value for model residual fluctuation.
[0018] Preferably, the specific steps for comprehensively analyzing the activation frequency of low-frequency patterns under the detection window to generate a reference value for low-frequency pattern activation are as follows:
[0019] Within the detection window, first define the low-frequency behavior patterns. For each low-frequency behavior pattern, record the activation times within the current detection window, and at the same time find the rarity factor corresponding in the historical data, and calculate the rare activation contribution value of each low-frequency pattern. The calculation expression is as follows:
[0020]
[0021] , where c i is the rare activation contribution, representing the quantification value of the abnormal activity degree of the i-th low-frequency behavior pattern, f i is the activation frequency of the low-frequency behavior pattern, representing the actual activation times of the i-th low-frequency behavior pattern, x i is the historical rarity factor, representing the rarity of the i-th low-frequency behavior pattern during the long-term operation;
[0022] After obtaining the rare activation contribution values c i of all low-frequency behavior patterns, perform non-linear fusion on all contribution values to generate a reference value for low-frequency pattern activation, comprehensively measure the potential pollution risk of all low-frequency pattern activation behaviors, and the calculation expression is as follows:
[0023]
[0024] , where LFPA is the reference value for low-frequency mode activation, N is the total number of low-frequency behavior patterns, and θ is the dynamic adjustment parameter.
[0025] Preferably, the specific steps for comprehensively analyzing the volatility of the model output residual within the time window under the detection window to generate the reference value of the model residual volatility are as follows:
[0026] Under the detection window, first, based on the residual sequence R = {r q} = {r1, r2, r3,..., r M}, where r q is the residual value of the model at the q-th time point, M is the total number of time points, calculate the dispersion of the residual amplitude interval of the residual sequence, extract the maximum value of the absolute value of the difference between the residuals of any two points in the sequence, and form a ratio with the sum of the absolute values of the first-order differences of the sequence to obtain the dispersion of the residual amplitude interval. The calculation expression is as follows:
[0027]
[0028] , where L is the dispersion of the residual amplitude interval, r j is the residual value of the model at the j-th time point, r k is the residual value of the model at the k-th time point, r k+1 is the residual value of the model at the (k + 1)-th time point, max(|r q - r j |) is the maximum residual difference value;
[0029] Based on the same residual sequence, extract the fluctuation density. The calculation expression is as follows:
[0030]
[0031] , where D is the residual fluctuation density, ω is the residual fluctuation determination threshold, which is the determination boundary for judging whether the residual fluctuation is "significant", and δ(·) is the indicator function, which returns 1 when the condition in the parentheses is true and 0 when the condition is false;
[0032] Integrate the dispersion of the residual amplitude interval L and the residual fluctuation density D to generate the reference value of the model residual volatility, which is used as the final index to characterize the abnormal degree of the residual sequence fluctuation within the current window. The calculation expression is as follows:
[0033] MRF = L·(1 + λD)
[0034] , where MRF is the reference value of the model residual volatility, and λ is the density weighting coefficient.
[0035] Preferably, the activated reference value of the low-frequency mode and the reference value of the model residual fluctuation after comprehensive analysis are used as feature vectors and input into a pre-trained machine learning model. A data pollution risk coefficient is generated through the machine learning model, and an intelligent risk assessment of the current network data is performed through the data pollution coefficient to determine whether there is a risk of disguised data pollution in the current network data.
[0036] Preferably, when performing an intelligent risk assessment of the current network data through a pre-trained machine learning model, the generated data pollution risk coefficient is compared and analyzed with a pre-set reference threshold of the data pollution risk coefficient to determine whether there is a risk of disguised data pollution in the current network data. The judgment logic is as follows:
[0037] If the data pollution risk coefficient is greater than the pre-set reference threshold of the data pollution risk coefficient, it is determined that there is a risk of disguised data pollution in the current network data;
[0038] If the data pollution risk coefficient is less than or equal to the pre-set reference threshold of the data pollution risk coefficient, it is determined that there is no risk of disguised data pollution in the current network data.
[0039] Preferably, when the machine learning model determines that there is a risk of disguised data pollution in the current data, the tolerance interval for anomaly determination is dynamically narrowed to improve the detection sensitivity to abnormal data. At the same time, the confidence level of the abnormal data determined to be disguised is reduced. When the confidence level is lower than the preset threshold, the specific steps for rejecting data entry are as follows:
[0040] After the machine learning model evaluates the current network data, the tolerance interval for anomaly determination is adaptively adjusted according to the data pollution risk coefficient DPR to improve the detection sensitivity to abnormal data. The adjustment formula is as follows:
[0041] Ω t =Ω t-1 ·exp(-λ·DPR)
[0042] where Ω t is the tolerance interval for anomaly determination at time point t, Ω t-1 is the tolerance interval for anomaly determination at the previous moment, and λ is a sensitivity control parameter used to adjust the influence intensity of the data pollution risk coefficient on the narrowing speed of the tolerance interval;
[0043] After obtaining the tolerance interval Ω t , the credibility of the current network data is dynamically evaluated to generate a confidence score to measure whether the current data can be received. The generation formula is as follows:
[0044] C=(1-DPR) β ·Ω t
[0045] , where C is the confidence score and β is the non - linear adjustment coefficient;
[0046] Compare the confidence score C with a preset security confidence threshold to determine whether the current network data is allowed to be stored in the database. The determination rules are as follows: if In the formula, C th is the security confidence threshold. If the current data confidence C is less than the threshold C th , then the current data is determined to have a spoofing risk and is not allowed to be stored in the database, thus blocking the interference of spoofed data on the subsequent model training and inference processes, and ensuring the data quality and operation safety of the intelligent system.
[0047] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0048] By introducing real - time monitoring of multi - source data, trend analysis of characteristic indicators, intelligent perception of model residual fluctuations, and a dynamic threshold regulation mechanism, the present invention constructs a closed - loop data quality control process. It not only improves the detection sensitivity and discrimination accuracy of abnormal data, but also realizes hierarchical processing and storage control of suspicious data through a confidence management mechanism. It can efficiently identify and dynamically defend against spoofed data pollution behaviors, effectively solve the problems of lack of adaptive recognition ability in the prior art, and the model misguidance and system instability caused by highly mimetic and continuously evolving spoofed data. Thus, it effectively guarantees the purity of model training data and the reliability of the intelligent decision - making process, and significantly enhances the robustness, adaptability, and long - term operation stability of the system against spoofed attacks. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is the flowchart of the computer intelligent service management method based on network big data of the present invention.
[0051] Figure 2 It is the mind map of the computer intelligent service management method based on network big data of the present invention. Detailed Embodiments
[0052] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0053] The present invention provides a computer intelligent service management method based on network big data as shown in Figure 1 the following, including the steps of:
[0054] Obtain the original data information of network data characteristics through real-time monitoring of network data streams, service logs, data access processes, and model feedback;
[0055] By monitoring network data streams, underlying information such as user behavior characteristics, access frequencies, request times, data sources, and communication paths can be collected; by monitoring service logs, actual operation records, resource invocation situations, and exception logs of users in the system can be obtained; by monitoring the data access process, the characteristics of data transmission, distribution, and aggregation processes within the system can be recorded; by monitoring model feedback, the response of the model to input data, confidence distribution, and model deviation can be observed. The role of this step is to establish a complete set of original feature data bases covering network data behavior, source, time series, and model feedback, providing a full-volume and reliable original information source for subsequent identification of disguised data.
[0056] The steps to obtain the original data information of network data characteristics are as follows: First, through real-time monitoring of network data streams, capture network communication behavior data including the transmission path of data packets, access frequencies, request times, and traffic characteristics; second, monitor the system service logs and extract business activity records generated by users in service calls, resource requests, operation behaviors, etc.; then, monitor the data access process and collect the structural information, transmission status, and access sources of data during access, distribution, aggregation, storage, etc.; finally, monitor the input and output of the machine learning model in real time, record the confidence of the model in network data, prediction deviation, and model residual feedback information, and form model response characteristics reflecting the data deviation trend. Through the above multi-dimensional monitoring process, the system can comprehensively, real-time, and continuously obtain the original feature information of network data in terms of behavior characteristics, time series characteristics, statistical characteristics, source attribute characteristics, and model deviation characteristics, providing data support for subsequent disguised data identification and intelligent analysis.
[0057] Preprocess the collected original feature information, and construct a structured feature data set based on the preprocessed data, providing a standardized and structured input data basis for subsequent feature extraction, model training, and intelligent evaluation;
[0058] To ensure the accuracy and efficiency of subsequent feature extraction and model recognition processes, it is necessary to preprocess the originally collected feature information (such as user behavior logs, network data streams, data access information, model feedback information, etc.). The preprocessing process usually includes data cleaning (removing abnormal, duplicate, and invalid data), data standardization (unifying formats, time alignment, and numerical normalization), feature correction (filling missing values and suppressing outliers), and feature mapping (mapping original fields to feature vectors or metrics for analysis). After completing the above operations, the system builds a structured feature dataset that meets the requirements of analysis and modeling based on the preprocessing results. This dataset can be used as input in a unified format and standardized features for subsequent feature extraction, machine learning model training, and intelligent evaluation. In short, the role of this step is to convert the original and messy monitoring data into a high-quality and ordered feature dataset that can be understood and analyzed by AI models, ensuring a stable, accurate, and sustainable input basis for subsequent disguised data detection and risk assessment processes.
[0059] Based on the data set, extract feature indicators that reflect the trend and persistent deviation of the data through feature engineering techniques, and conduct a comprehensive analysis of the extracted feature indicators to complete the quantification of data disguise features;
[0060] Based on the data set, extract feature indicators that reflect the trend and persistent deviation of the data through feature engineering techniques. The extracted indicators include the activation frequency of low-frequency patterns (such as rare requests and extreme behaviors) and the volatility of model output residuals within a time window. Conduct a comprehensive analysis of the activation frequency of low-frequency patterns (such as rare requests and extreme behaviors) and the volatility of model output residuals within a time window under the detection window to generate a low-frequency pattern activation reference value and a model residual fluctuation reference value respectively, and quantify the data disguise features through the low-frequency pattern activation reference value and the model residual fluctuation reference value.
[0061] Low-frequency patterns (such as rare requests and extreme behaviors) are frequently activated, which can usually be used as one of the important signals to determine whether the current data has the risk of disguised data pollution. Normal network data flow or user behavior data usually has a stable feature distribution. Low-frequency patterns only appear occasionally in specific scenarios and show random and sparse characteristics. However, in order to evade conventional anomaly detection, disguised data often chooses to simulate low-probability features in normal business (such as occasional marginal behaviors, unpopular requests, operations in unconventional time periods, etc.) to confuse system judgment and reduce the probability of being detected by operating in low-frequency patterns. When the system detects that low-frequency patterns are abnormally activated or recurring within the sliding time window and show trending and continuous characteristics, it often indicates that there is abnormal data intervention or artificially constructed behavior characteristics, which is a typical disguised data pollution signal. This type of disguised behavior maintains normal data format, field distribution, etc. on the surface, but often shows continuous deviations in hidden features such as time series distribution and pattern activation frequency, thereby interfering with model training, decision-making and even resource scheduling, and is extremely hidden and destructive. Therefore, abnormal activation of low-frequency patterns is an important feature for detecting the risk of disguised data pollution.
[0062] The specific steps of comprehensively analyzing the low-frequency mode activation frequency in the detection window to generate the low-frequency mode activation reference value are as follows:
[0063] In the detection window, we first define low-frequency behavior patterns. Low-frequency behavior patterns are system preset or learned low-frequency behavior patterns, such as rare interface access, high-frequency operations during non-peak hours, or abnormal path calls. For each low-frequency behavior pattern, we record the number of activations in the current detection window and find the corresponding rarity factor in the historical data. The behavior factor reflects the rarity of the pattern in long-term behavior. The value range is [0, 1]. The smaller the value, the rarer the pattern. We calculate the rare activation contribution value of each low-frequency pattern. The calculation expression is as follows:
[0064]
[0065] , where c i is the rare activation contribution, which represents the quantified value of abnormal activity for the i-th low-frequency behavior pattern, and f i is the activation frequency of the low-frequency behavior mode, indicating the actual number of times the i-th low-frequency behavior mode is activated, x i is the historical rarity factor, which indicates the rarity of the i-th low-frequency behavior pattern in the long-term operation process;
[0066] By extracting the activation frequency of each low-frequency pattern within the current detection window and combining it with its historical rarity factor, a discriminative rare activation contribution value is calculated. This step can highlight those behavioral patterns that are abnormally frequently triggered in the current period and extremely rare in history, providing a highly sensitive feature basis for subsequent calculations.
[0067] After obtaining the rare activation contribution values c of all low-frequency behavioral patterns i a non-linear fusion of all contribution values is performed to generate a low-frequency pattern activation reference value, comprehensively measuring the potential contamination risk of all low-frequency pattern activation behaviors. The calculation expression is as follows:
[0068]
[0069] , where LFPA is the low-frequency pattern activation reference value, N is the total number of low-frequency behavioral patterns, and θ is a dynamic adjustment parameter used to control the sensitivity and adjustability of the exponential amplification function.
[0070] By performing non-linear fusion on the rare activation contribution values of each low-frequency pattern, a low-frequency pattern activation reference value that can comprehensively reflect the abnormal activation level of the global low-frequency pattern is generated. The low-frequency pattern activation reference can be used to accurately measure whether there is a risk of disguised data contamination caused by multiple low-frequency abnormal behaviors in the current data.
[0071] The larger the low-frequency pattern activation reference value generated after comprehensive analysis of the low-frequency pattern activation frequency under the detection window, the more "deliberately created abnormal edge behaviors" are triggered in the current system. This is often a phenomenon caused by disguised data attempting to simulate normal behaviors through non-mainstream paths and avoid detection. Therefore, it can be determined that there is a risk of disguised data contamination in the current data. Conversely, if this reference value remains near the historical baseline without abnormal fluctuations, it indicates that the low-frequency pattern activation remains at a normal level, and it can be initially considered that there are no significant signs of disguised contamination in the current data.
[0072] When the volatility of the model output residuals shows a trend of growth within a time window, it can often be regarded as an important signal indicating the risk of disguised data contamination in the current data. The residual reflects the degree of difference between the model prediction value and the actual observation value. Under normal circumstances, the model is trained based on the historical feature distribution. For inputs with stable data features, the residual fluctuation should remain within a relatively stable range. However, although disguised data highly simulates normal data in form, there are abnormalities in feature logic, combination patterns, or underlying structures. When the model processes such data, it often exhibits unstable prediction performance, resulting in abnormal residual fluctuations. When the system continuously monitors that the residual volatility shows an increasing trend within the time window, it indicates that the abnormal input is not accidental but has a certain persistence and trend. It is very likely that the attacker uses a step-by-step interference and progressive contamination strategy to try to bypass one-time detection and chronically interfere with the model learning process. This trend of residual fluctuation not only indicates that the current data has begun to deviate from the model's cognitive boundary but also implies that the disguised data is continuously affecting the system's judgment in a "subtle" way. Therefore, the potential contamination risk behind it should be highly vigilant.
[0073] The specific steps for comprehensively analyzing the volatility of the model output residuals within a time window under the detection window to generate a reference value for model residual fluctuation are as follows:
[0074] Under the detection window, first, based on the residual sequence R = {r q} = {r1, r2, r3, ……, r M}, where r q is the residual value of the model at the qth time point, and M is the total number of time points. Calculate the discreteness of the residual amplitude interval of the residual sequence, extract the maximum value of the absolute value of the difference between the residuals of any two points in the sequence, and form a ratio with the sum of the absolute values of the first-order differences of the sequence to obtain the residual amplitude interval discreteness. The calculation expression is as follows:
[0075]
[0076] , where L is the residual amplitude interval discreteness, which is used to measure the proportion of the strongest jump amplitude in the residual sequence relative to the overall fluctuation level, r j is the residual value of the model at the jth time point, r k is the residual value of the model at the kth time point, r k+1 is the residual value of the model at the (k + 1)th time point, max(|r q - r j |) is the maximum residual difference value, which represents the maximum value of the absolute value of the difference between the residual values of any two time points within the detection window;
[0077] By calculating the ratio of the maximum local fluctuation in the residual sequence to the overall fluctuation trend, it is possible to identify whether there are sudden and non-stationary abnormal changes in the current data. The above steps help to capture the intense disturbance signals caused by disguised data, providing a key basis for judging whether the data deviates from the model's cognitive boundary.
[0078] Based on the same residual sequence, the fluctuation density is extracted, that is, within the detection window, the proportion of the number of fluctuations where the amplitude change of the residual exceeds a specified threshold, and the calculation formula is as follows:
[0079]
[0080] , where D is the residual fluctuation density, indicating the frequency at which the adjacent change amplitude of the model residual value exceeds the set threshold in the detection window, ω is the residual fluctuation determination threshold, which is the determination boundary for judging whether the residual fluctuation is "significant". When the adjacent residual change amplitude |r k+1 -r k | exceeds this threshold, it is recognized as a "high-fluctuation event". δ(·) is an indicator function that returns 1 when the condition in the parentheses is true and 0 when the condition is false;
[0081] By statistically analyzing the change frequency of the fluctuation amplitude of the model residual sequence that exceeds the set threshold, the abnormal disturbance density of the current data within the detection window is quantified. This indicator can effectively reflect the latent characteristics of disguised data interfering with the model in a small-scale and frequent manner, providing key support for subsequent risk assessment.
[0082] Combining the residual amplitude interval dispersion L and the residual fluctuation density D, a model residual fluctuation reference value is generated as the final indicator for characterizing the abnormal degree of the residual sequence fluctuation within the current window. The calculation formula is as follows:
[0083] MRF = L·(1 + λD)
[0084] , where MRF is the model residual fluctuation reference value, and λ is the density weighting coefficient, which is used to control the influence weight of the residual fluctuation density D on the model residual fluctuation reference value and is a sensitivity adjustment factor.
[0085] By fusing the residual amplitude interval dispersion and the residual fluctuation density, a model residual fluctuation reference value that can simultaneously characterize local mutations and high-frequency perturbations is constructed to quantify the abnormal fluctuation characteristics of the data within the current detection window. The above steps can effectively identify the trend and persistent fluctuation risks caused by disguised data contamination, providing an accurate judgment basis for subsequent anomaly detection and dynamic defense.
[0086] The volatility of the model output residuals within the time window is comprehensively analyzed under the detection window to generate a reference value for the model residual volatility. The larger the value, the more significant the volatility of the prediction residuals of the model when processing the current data compared to the historical normal level. This is manifested as unstable or abnormal fluctuations in the model's prediction ability for the input data, which is often caused by the gradual deviation of the data feature distribution or the hidden abnormal features in the input data. Since the disguised data may still conform to the original statistical characteristics in the short term, it is prone to causing difficulties in model fitting or local anomalies on a longer time scale, ultimately reflected in the continuous increase of the reference value for residual volatility. Therefore, when the reference value for model residual volatility is larger, it can be reasonably determined that there is a risk of disguised data contamination in the current data; conversely, if this reference value is maintained within the normal volatility range for a long time, it indicates that the model's fitting of the data is stable, and the current data does not exhibit the characteristics of disguised data contamination.
[0087] The feature indicators after comprehensive analysis are used as feature vectors and input into a pre-trained machine learning model. Through the machine learning model, an intelligent risk assessment of the current network data is carried out to determine whether there is a risk of disguised data contamination in the current network data;
[0088] The reference value for low-frequency pattern activation and the reference value for model residual volatility after comprehensive analysis are used as feature vectors and input into a pre-trained machine learning model. Through the machine learning model, a data contamination risk coefficient is generated, and through the data contamination coefficient, an intelligent risk assessment of the current network data is carried out to determine whether there is a risk of disguised data contamination in the current network data.
[0089] The "pre-trained machine learning model" refers to a model constructed based on historical network data and capable of identifying data disguise characteristics. This model has completed model structure optimization and parameter learning through offline training before system deployment and has the ability to conduct risk assessment on newly input data. The training process of this model usually includes multiple stages such as data collection, feature engineering, sample annotation, model selection and training. The training data used comes from known normal data and disguised data in history, or approximately disguised samples constructed based on rules, covering various behavior patterns, data distributions, and attack methods. By training on the key features (such as the reference value for low-frequency pattern activation, the reference value for model residual volatility, etc.) extracted from these samples, the model can learn the differences and anomalies presented by the disguised data in the high-dimensional feature space. During the training process, the system can use supervised learning methods (such as decision trees, support vector machines, neural networks, etc.), or combine unsupervised methods (such as clustering, isolation forest, anomaly detection algorithms, etc.) to improve the model's generalization ability for unknown disguise patterns.
[0090] During actual deployment, the pre-trained model is no longer adjusted in terms of structure or parameters, but only serves as an inference model to quickly evaluate newly accessed data. The system will preprocess and extract features from the newly collected network data to form multi-dimensional feature vectors such as "low-frequency pattern activation reference value" and "model residual fluctuation reference value", and input them into the model. The model evaluates the current input data according to the decision boundary or probability distribution learned internally and outputs a "data pollution risk coefficient", which can be understood as the likelihood score of the current data being judged as camouflaged data. Subsequently, the system makes an intelligent decision on whether there is a risk of camouflaged data pollution in the current network data based on the magnitude of this risk coefficient, in combination with the set security policies (such as threshold judgment, confidence scoring, dynamic response mechanism, etc.). The core value of this model lies in that it can quickly and stably complete the discrimination task in the face of complex and variable camouflaged data behavior patterns, effectively improve the system's ability to identify highly concealed data pollution, and achieve closed-loop management from the data feature layer to the risk control layer.
[0091] The machine learning model is not limited here. Any machine learning model that can comprehensively analyze the low-frequency pattern activation reference value LFPA and the model residual fluctuation reference value MRF to generate a data pollution risk coefficient DPR can be used. To implement the technical solution of the present invention, the present invention provides a specific generation method:
[0092] The formula for generating the data pollution risk coefficient DPR is as follows: DPR = β a ·LFPA + β b ·MRF, where β a and β b are respectively the preset proportionality coefficients of the low-frequency pattern activation reference value LFPA and the model residual fluctuation reference value MRF, and β a and β b are both greater than 0.
[0093] The preset proportionality coefficients here (i.e., β a and β b in the formula) refer to the weight coefficients assigned to the low-frequency pattern activation reference value LFPA and the model residual fluctuation reference value MRF respectively when generating the data pollution risk coefficient DPR, which are used to control the contribution degree and influence of these two feature indicators on the final data pollution risk coefficient DPR. Since the numerical ranges, sensitivities, and importance in actual camouflaged data detection of the low-frequency pattern activation reference value LFPA and the model residual fluctuation reference value MRF are usually different, if directly added or combined, it may cause a certain indicator to dominate the result of the data pollution risk coefficient DPR, affecting the accuracy of the evaluation. Therefore, β a and β bAs a proportionality coefficient, it is possible to allocate reasonable weights to the low-frequency mode activation reference value LFPA and the model residual fluctuation reference value MRF according to actual business requirements or through experimental optimization, so that the data pollution risk coefficient DPR can more accurately reflect the combined impact of the two on the risk of camouflaged data.
[0094] For example, if experiments find that the activation of the low-frequency mode has a stronger recognition ability for camouflage detection, then β can be appropriately a set to a larger value. And if the model residual fluctuation is more sensitive to persistent camouflaged data, the weight of β can be increased. b The role of the preset proportionality coefficient is to ensure that when generating the risk coefficient, the model reasonably weighs the contributions of multiple features according to different scenarios and data distributions, so that the data pollution risk coefficient DPR is more in line with the actual situation in subsequent risk judgments, improving the stability and accuracy of camouflaged data detection. The preset proportionality coefficient can be set manually or obtained through automatic learning during model training or system optimization.
[0095] From the data pollution risk coefficient, it can be seen that the larger the low-frequency mode activation reference value generated by comprehensively analyzing the low-frequency mode activation frequency under the detection window, and the larger the model residual fluctuation reference value generated by comprehensively analyzing the volatility of the model output residuals within the time window under the detection window, the greater the data pollution risk coefficient generated when the pre-trained machine learning model conducts intelligent risk assessment on the current network data, indicating that there is a risk of camouflaged data pollution in the current data. On the contrary, it indicates that there is no risk of camouflaged data pollution in the current data.
[0096] Compare and analyze the data pollution risk coefficient generated when the pre-trained machine learning model conducts intelligent risk assessment on the current network data with the pre-set data pollution risk coefficient reference threshold to determine whether there is a risk of camouflaged data pollution in the current network data. The judgment logic is as follows:
[0097] If the data pollution risk coefficient is greater than the pre-set data pollution risk coefficient reference threshold, it is determined that there is a risk of camouflaged data pollution in the current network data;
[0098] If the data pollution risk coefficient is less than or equal to the pre-set data pollution risk coefficient reference threshold, it is determined that there is no risk of camouflaged data pollution in the current network data.
[0099] When the machine learning model determines that there is a risk of camouflaged data pollution in the current data, dynamically narrow the tolerance interval for anomaly determination, improve the detection sensitivity for abnormal data. At the same time, reduce the confidence level of the abnormal data determined to be camouflaged. When the confidence level is lower than the preset threshold, reject the data from being stored in the database to prevent camouflaged data from polluting the subsequent model training and inference processes;
[0100] When the machine learning model identifies a potential risk of camouflaged contamination in network data, it actively and intelligently enhances the monitoring intensity and filtering accuracy of abnormal data. On the one hand, by dynamically narrowing the tolerance interval for anomaly determination, that is, automatically adjusting the range of the anomaly detection threshold, the system becomes more sensitive to tiny, concealed, and trend-based offset data features, thereby enhancing the ability to identify abnormal data and preventing camouflaged data from being missed due to overly loose thresholds. On the other hand, the system reduces the confidence level of suspected camouflaged abnormal data in real-time, that is, assigns a lower trust level, and establishes a data credibility grading mechanism. When the credibility is lower than a set security threshold, such data is directly rejected from entering the subsequent data storage phase. This mechanism can ensure that truly contaminated camouflaged data is intercepted at the source stage, effectively preventing camouflaged data from further entering key links such as data storage, model training, and model inference, avoiding the long-term and implicit contamination of the model, ensuring the accuracy, reliability of subsequent intelligent decision-making results, and the safe and stable operation of the service system, reducing the systemic risks brought by data contamination risks, and improving service quality and user experience.
[0101] When the machine learning model determines that there is a risk of camouflaged data contamination in the current data, the specific steps of dynamically narrowing the tolerance interval for anomaly determination, enhancing the detection sensitivity of abnormal data, and at the same time reducing the confidence level of abnormal data determined to be camouflaged and rejecting data from being stored in the database when the confidence level is lower than the preset threshold are as follows:
[0102] After the machine learning model evaluates the current network data, according to the data contamination risk coefficient DPR, the tolerance interval for anomaly determination is adaptively adjusted to enhance the detection sensitivity of abnormal data. The adjustment formula is as follows:
[0103] Ω t =Ω t-1 ·exp(-λ·DPR)
[0104] , where Ω t is the tolerance interval for anomaly determination at time point t, Ω t-1 is the tolerance interval for anomaly determination at the previous moment, and λ is the sensitivity control parameter used to adjust the influence intensity of the data contamination risk coefficient on the narrowing speed of the tolerance interval;
[0105] When the data contamination risk coefficient, the system can exponentially converge the tolerance interval for anomaly determination, enhance the detection sensitivity of subsequent camouflaged data, and prevent camouflaged data from continuously mixing in the short term.
[0106] After obtaining the tolerance interval Ω t , the credibility of the current network data is dynamically evaluated to generate a confidence score to measure whether the current data can be received. The generation formula is as follows:
[0107] C = (1 - DPR) β ·Ω t
[0108] , where C is the confidence score, representing the degree of trust of the system in the current network data being "true and reliable data", which is a scoring-type indicator, β is the non-linear adjustment coefficient, controlling the influence intensity of the data pollution risk coefficient on the confidence, and amplifying or reducing the weight of the risk through the exponential function;
[0109] In the above formula, when the data pollution risk coefficient is relatively high, even if the tolerance interval has not completely narrowed, the confidence can still drop rapidly; for low-risk data, there is still a chance to maintain a relatively high confidence under the influence of the change of the tolerance interval, so as to achieve intelligent hierarchical management driven by risk.
[0110] Compare the confidence score C with the preset safety confidence threshold to determine whether the current network data is allowed to be stored in the database. The judgment rule is as follows: if where C th is the safety confidence threshold. If the current data confidence C is less than the threshold C th , then the current data is determined to have the risk of being disguised and is not allowed to be stored in the database, thus blocking the interference of the disguised data on the subsequent model training and inference processes, and ensuring the data quality and operation safety of the intelligent system.
[0111] By comparing the data confidence with the safety confidence threshold, it can accurately judge whether there is a risk of disguised data pollution in the current network data. It can effectively block the high-risk data with low confidence from being stored in the database and prevent the disguised data from polluting the subsequent model training and inference processes.
[0112] The present invention constructs a set of closed-loop data quality control processes by introducing real-time monitoring of multi-source data, trend analysis of characteristic indicators, intelligent perception of model residual fluctuations, and dynamic threshold regulation mechanisms. It not only improves the detection sensitivity and discrimination accuracy of abnormal data, but also realizes hierarchical processing and storage control of suspicious data through the confidence management mechanism. It can achieve efficient identification and dynamic defense of disguised data pollution behaviors, effectively solve the problems of lack of adaptive recognition ability in the prior art and the model misguidance and system instability caused by highly mimetic and continuously evolving disguised data, thereby effectively ensuring the purity of model training data and the reliability of the intelligent decision-making process, and significantly enhancing the robustness, adaptability and long-term operation stability of the system against disguised attacks.
[0113] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0114] Certain exemplary embodiments of the present invention have been described above only by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0115] It should be noted that in this text, if there are relative terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0116] It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above processes does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0117] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0118] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated herein.
[0119] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0120] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0121] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0122] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.
Claims
1. A computer intelligent service management method based on network big data, characterized in that, It includes the following steps: Obtain the original data information of network data characteristics through real-time monitoring of network data streams, business logs, data access processes, and model feedback; Preprocess the collected original feature information, and construct a structured feature dataset based on the preprocessed data to provide a standardized and structured input data basis for subsequent feature extraction, model training, and intelligent evaluation; Based on the dataset, extract feature indicators that reflect the trend and persistent deviation of the data through feature engineering techniques, and comprehensively analyze the extracted feature indicators to complete the quantification of data camouflage features; Input the feature indicators after comprehensive analysis as feature vectors into a pre-trained machine learning model, and use the machine learning model to perform intelligent risk assessment on the current network data to determine whether there is a risk of camouflaged data pollution in the current network data; When the machine learning model determines that there is a risk of camouflaged data pollution in the current data, dynamically narrow the tolerance interval for anomaly determination to improve the detection sensitivity of abnormal data. At the same time, reduce the confidence level of the abnormal data determined to be camouflaged. When the confidence level is lower than the preset threshold, reject the data from being stored in the database to prevent camouflaged data from polluting the subsequent model training and inference processes.
2. The computer intelligent service management method based on network big data according to claim 1, characterized in that, Obtain the original data information of network data characteristics. The specific steps are as follows: First, through real-time monitoring of network data streams, capture network communication behavior data including the transmission path, access frequency, request time, and traffic characteristics of data packets; Secondly, monitor the system business logs and extract the business activity records generated by users in the service call, resource request, and operation behavior links; Subsequently, monitor the data access process and collect the structure information, transmission status, and access sources of the data during the access, distribution, aggregation, and storage processes; Finally, real-time monitor the input and output of the machine learning model, record the confidence level, prediction deviation, and model residual feedback information of the model for network data, and form model response characteristics that reflect the data deviation trend.
3. The computer intelligent service management method based on network big data according to claim 1, characterized in that, Based on the dataset, extract feature indicators that reflect the trend and persistent deviation of the data through feature engineering techniques. The extracted indicators include the activation frequency of low-frequency patterns and the volatility of the model output residuals within a time window. The activation frequency of low-frequency patterns and the volatility of the model output residuals within a time window are comprehensively analyzed under the detection window to generate a low-frequency pattern activation reference value and a model residual fluctuation reference value respectively. The data camouflage features are quantified through the low-frequency pattern activation reference value and the model residual fluctuation reference value.
4. The computer intelligent service management method based on network big data according to claim 3, characterized in that, The specific steps for comprehensively analyzing the activation frequency of low-frequency patterns under the detection window to generate a low-frequency pattern activation reference value are as follows: Within the detection window, first define the low-frequency behavior patterns. For each low-frequency behavior pattern, record the activation times within the current detection window, and at the same time find the corresponding rarity factor in the historical data, and calculate the rare activation contribution value of each low-frequency pattern. The calculation formula is as follows: , Where c i is the rare activation contribution, representing the quantification value of the abnormal activity degree of the i-th low-frequency behavior pattern, and f i is the activation frequency of the low-frequency behavior pattern, representing the actual activation times of the i-th low-frequency behavior pattern, and x i is the historical rarity factor, representing the rarity of the i-th low-frequency behavior pattern during the long-term operation; After obtaining the rare activation contribution value c of all low-frequency behavior patterns i After that, non-linear fusion is performed on all contribution values to generate a low-frequency pattern activation reference value, comprehensively measuring the potential pollution risk of all low-frequency pattern activation behaviors. The calculation expression is as follows: , In the formula, LFPA is the low-frequency pattern activation reference value, N is the total number of low-frequency behavior patterns, and θ is the dynamic adjustment parameter.
5. The computer intelligent service management method based on network big data according to claim 3, characterized in that, The specific steps for comprehensively analyzing the volatility of the model output residuals within a time window under a detection window to generate a reference value for model residual volatility are as follows: Under the detection window, first, based on the residual sequence R = {r q} = {r1, r2, r3, ……, r M}, where r q is the residual value of the model at the q-th time point, M is the total number of time points, calculate the dispersion of the residual amplitude interval of the residual sequence, extract the maximum value of the absolute value of the difference between the residuals of any two points in the sequence, and form a ratio with the sum of the absolute values of the first-order differences of the sequence to obtain the dispersion of the residual amplitude interval. The calculation expression is as follows: , where L is the dispersion of the residual amplitude interval, and r j is the residual value of the model at the j-th time point, and r k is the residual value of the model at the k-th time point, and r k+1 is the residual value of the model at the (k + 1)-th time point, and max(|r q - r j |) is the maximum residual difference value; Based on the same residual sequence, the volatility density is extracted, and the calculation expression is as follows: , In the formula, D is the residual volatility density, ω is the residual volatility determination threshold, which is the determination boundary for judging whether the residual volatility is "significant", δ(·) is the indicator function, which returns 1 when the condition in the parentheses is true and 0 when the condition is false; By synthesizing the residual amplitude interval discreteness L and the residual volatility density D, a reference value for model residual volatility is generated, which is used as the final index to characterize the abnormal degree of the residual sequence volatility within the current window. The calculation expression is as follows: MRF = L·(1 + λD), In the formula, MRF is the reference value for model residual volatility, and λ is the density weighting coefficient.
6. The computer intelligent service management method based on network big data according to claim 3, characterized in that The reference value for low-frequency mode activation after comprehensive analysis and the reference value for model residual volatility are used as feature vectors and input into a pre-trained machine learning model. The machine learning model generates a data pollution risk coefficient, and the current network data is intelligently risk-assessed through the data pollution coefficient to determine whether there is a risk of disguised data pollution in the current network data.
7. The computer intelligent service management method based on network big data according to claim 6, characterized in that The data pollution risk coefficient generated when the pre-trained machine learning model conducts an intelligent risk assessment on the current network data is compared and analyzed with a pre-set reference threshold for the data pollution risk coefficient to determine whether there is a risk of disguised data pollution in the current network data. The judgment logic is as follows: If the data pollution risk coefficient is greater than the pre-set reference threshold for the data pollution risk coefficient, it is determined that there is a risk of disguised data pollution in the current network data; If the data pollution risk coefficient is less than or equal to the pre-set reference threshold for the data pollution risk coefficient, it is determined that there is no risk of disguised data pollution in the current network data.
8. The computer intelligent service management method based on network big data according to claim 7, characterized in that When the machine learning model determines that there is a risk of disguised data pollution in the current data, the tolerance interval for anomaly determination is dynamically narrowed to improve the detection sensitivity for abnormal data. At the same time, the confidence level for the abnormal data determined to be disguised is reduced. When the confidence level is lower than the preset threshold, the specific steps for rejecting data entry are as follows: After the machine learning model evaluates the current network data, the tolerance interval for anomaly determination is adaptively adjusted according to the data pollution risk coefficient DPR to improve the detection sensitivity for abnormal data. The adjustment formula is as follows: Ω t = Ω t-1 · exp(-λ·DPR), where, Ω t is the abnormal determination tolerance interval at time point t, and Ω t-1 is the abnormal determination tolerance interval at the previous moment. λ is the sensitivity control parameter, which is used to adjust the influence intensity of the data contamination risk coefficient on the narrowing speed of the tolerance interval; After obtaining the tolerance interval Ω t the credibility of the current network data is dynamically evaluated to generate a confidence score to measure whether the current data can be received. The generation formula is as follows: C = (1 - DPR) β ·Ω t , In the formula, C is the confidence score, and β is the non-linear adjustment coefficient; Compare the confidence score C with a preset security confidence threshold to determine whether the current network data is allowed to be stored in the database. The judgment rule is as follows: if In the formula, C th is the security confidence threshold. If the current data confidence C is less than the threshold C th , then the current data is determined to have a risk of spoofing and is not stored in the database, thereby blocking the interference of spoofed data on the subsequent model training and inference processes, and ensuring the data quality and operation security of the intelligent system.
Citation Information
Patent Citations
Intelligent network security protection method and system based on big data
CN118413368A
Micro-service scheduling method and system based on edge computing
CN119201407A