Supervision data quality monitoring system and method based on artificial intelligence
By building a dual-index load analysis set based on timestamps and business interface identification, combining weighted fusion calculation and time series prediction model, dynamically analyzing the business interface call trend, the problem of insufficient real-time performance in traditional monitoring methods is solved, efficient and accurate abnormal detection and early warning is achieved, and the data quality control capabilities of the supervision system are improved.
Patent Information
- Application Number
- CN202510840263.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional regulatory data quality monitoring methods are difficult to meet the needs of high-frequency and dynamically changing business interface call scenarios, and lack dynamic analysis capabilities for business interface loads, resulting in insufficient real-time performance and lag in abnormal detection, which affects the stability and data reliability of the regulatory system.
By setting a fixed supervision data acquisition cycle, a dual-index business interface load analysis set based on timestamps and unique identification of the business interface is constructed, combining weighted fusion calculations and time series prediction models, dynamically analyze business interface call trends and abnormal situations, build a multi-level abnormality monitoring system, and use the three Sigma principle to set dynamic monitoring thresholds for hierarchical early warning.
It realizes a forward-looking analysis of the trend of business interface calling, significantly improves the real-time nature of abnormal detection and early warning accuracy, avoids the lag and misjudgment risks of traditional methods, and provides efficient and reliable data quality control for the regulatory system.
Smart Images

Figure CN120336129A_ABST
Abstract
Description
Technical Field
[0001] This invention application relates to the field of artificial intelligence technology, and specifically to a supervision data quality monitoring system and method based on artificial intelligence. Background Art
[0002] In the digital supervision scenario, data quality is the core element to ensure the scientific nature of supervision decisions and the stability of system operation. Accurate and reliable data quality detection can timely discover abnormal calls of business interfaces, effectively avoid the risks of supervision loopholes and decision-making mistakes caused by data deviation, which is of great significance for improving supervision efficiency, reducing operation costs, and enhancing the anti-risk ability of the system. It is also a necessary prerequisite for realizing the digital transformation and sustainable development of supervision.
[0003] In the current field of supervision data quality monitoring, traditional methods are difficult to meet the requirements of high-frequency and dynamically changing business interface call scenarios. Especially in large-scale supervision systems, traditional methods lack the ability to dynamically analyze the load of business interfaces, resulting in insufficient real-time performance and lagging anomaly detection. They usually use fixed thresholds or static models for monitoring, and cannot adapt to the dynamic changes of business interface call frequencies and anomaly ratios, thus seriously affecting the stability of the supervision system and data reliability. Therefore, there is an urgent need for a monitoring method that can dynamically analyze the load of business interfaces, accurately predict call trends and anomalies to improve data quality detection. Summary of the Invention
[0004] The purpose of this invention application is to provide a supervision data quality monitoring system and method based on artificial intelligence to solve the problems raised in the prior art.
[0005] To achieve the above purpose, this invention application provides the following technical solution: A supervision data quality monitoring method based on artificial intelligence, and the intelligent management method includes the following steps: Step S1: Set the supervision data collection period, and obtain the call data of all business interfaces according to the supervision data collection period to construct a business interface load analysis set; Step S1-1: Set the supervision data collection period T, and the supervision data collection period T is a fixed time interval; Step S1-2: Based on the set supervision data collection period T, obtain the call data of all business interfaces according to the supervision data collection period T through application programming interface (API) calls or log parsing; Step S1-3: Perform standardization processing on the obtained business interface call data, including data cleaning and data format unification; Step S1-4: Use the timestamp as the primary index and the unique identifier of the business interface as the secondary index for the standardized business interface call data to construct a business interface load analysis set.
[0006] By setting a fixed acquisition period, orderly obtaining and standardizing the data of business interface calls, and constructing a business interface load analysis set in a dual-index manner, it is possible to ensure the timeliness and standardization of data acquisition, effectively integrate scattered call information, provide a high-quality data foundation with clear structure and unified format for subsequent data statistics, trend analysis, and quality monitoring, and improve the analysis efficiency and accuracy.
[0007] Step S2: Analyze the obtained call times and call exception times data corresponding to the business interface timestamps based on the business interface load analysis set, obtain the call data of the business interface for each regulatory data acquisition period in the history of the business interface for analysis, and obtain the peak call times of the business interface; Step S2-1: Based on the business interface load analysis set, group and count the business interface call data according to timestamps and the unique identifier of the business interface; Step S2-2: Analyze and extract the call times and call exception times for each timestamp of the business interface; Step S2-3: Obtain the call data of the business interface for each regulatory data acquisition period in the history of the business interface, and obtain the maximum call times of the business interface through bubble sort, which is recorded as the peak call times.
[0008] Group and count the data in the business interface load analysis set, extract the call times and exception times of the business interface at each timestamp, obtain the peak call times through bubble sort in combination with historical data, effectively analyze the fluctuation characteristics of the business interface, quantify the call intensity and exception situation of the business interface, identify the historical call extreme values, provide key data support for subsequent call trend prediction, system load assessment, and exception warning, and enhance the pertinence and reliability of data monitoring.
[0009] Step S3: Analyze and predict the call times trend of the business interface through the peak call times of the business interface, and analyze and obtain the call exception ratio trend based on the call exception times data corresponding to the timestamps of the business interface; construct a call exception change trend chart based on the call exception times data of the business interface, and analyze and predict the call exception times of the business interface in combination with the call exception times trend; Step S3-1: Select the timestamp corresponding to the peak call times as the starting point, and obtain the call times data that is only less than the peak call times within the starting point and the current time period of the business interface, which is recorded as the second peak call times; Step S3-2: According to the peak call times and the second peak call times of the business interface, calculate and obtain the maximum call times within the next regulatory data acquisition period T of the business interface through weighted fusion, which is recorded as the call times prediction threshold; The calculation formula for calculating the maximum call times within the next regulatory data acquisition period T of the business interface through weighted fusion is as follows: ; In the formula, F max,new represents the maximum number of calls within the next regulatory data collection period T of the service interface; w represents the weight coefficient of the peak call number; F max,1 represents the peak call number of the service interface; F max,2 represents the call number data that is only less than the peak call number between the starting point of the service interface and the current time period.
[0010] Step S3-3: Obtain the call number data of each historical regulatory data collection period T of the service interface, select the call number prediction threshold as the upper limit of the predicted value of the service interface call number, and perform prediction through the exponential smoothing method in the time series prediction model to obtain the call number of the next regulatory data collection period T of the service interface, denoted as the call number predicted value; The prediction calculation of the exponential smoothing method uses the following formula: ; In the formula, F T+1 represents the predicted value of the call number of the next regulatory data collection period T of the service interface; a represents the average coefficient, and its value range is in the interval of 1 and 0; Y T represents the actual call number of the current regulatory data collection period T; F T represents the predicted value of the call number of the current regulatory data collection period T; By applying the exponential smoothing method in the time series prediction model, combining the call number data of each historical regulatory data collection period of the service interface, and using the call number prediction threshold as the upper limit of the predicted value, the dynamic prediction of the call number of the next regulatory data collection period of the service interface is realized. And the average coefficient in the time series prediction model is adaptively adjusted according to the data fluctuation characteristics, enabling the model to learn the call pattern and trend law from historical data, and having the automatic recognition ability of data characteristics like artificial intelligence.
[0011] When new data is accessed, the time series prediction model will perform weighted processing on the current data based on the historical learning results, and then deduce and predict the future call number. This dynamic parameter adjustment and pattern recognition logic realized through data driving fully reflects the core characteristics of self-optimization and continuous evolution of artificial intelligence technology in the field of data prediction, forming a significant difference from the traditional fixed parameter model.
[0012] Step S3-4: Conduct primary monitoring on the service interface according to the call number prediction threshold, and the primary monitoring warning judgment is as follows: Step S3-4-1: When the call number predicted value exceeds the call number prediction threshold, send out a warning signal for the service interface; Step S3-4-2: When the predicted call count does not exceed the predicted call count threshold, continue to monitor the business interface; Step S3-5: Divide the number of call anomalies in each historical regulatory data collection period T of the business interface by the number of calls in the corresponding regulatory data collection period T to obtain the call anomaly ratio; Step S3-6: Analyze and calculate the average call anomaly ratio based on the call anomaly ratios of each historical regulatory data collection period T of the business interface; Step S3-7: Multiply the predicted call count value of the business interface by the average call anomaly ratio to obtain the call anomaly count of the business interface, denoted as the predicted call anomaly count value.
[0013] Through weighted fusion calculation and combined with the time series prediction model, it can accurately capture the dynamic change characteristics of business interface calls, scientifically set the prediction threshold and generate the predicted call count value, realizing the forward-looking analysis of the business interface call trend; at the same time, calculate the average level based on the historical anomaly ratio, combine the predicted call count value to generate the predicted anomaly count value, construct the anomaly change trend graph, form a multi-level anomaly monitoring system from a single business interface to the overall system, effectively improve the real-time performance of anomaly detection and the accuracy of early warning, and provide a data-driven decision-making basis for system load assessment and quality control. This method uses historical data, applies the time series prediction model to automatically identify and learn the change pattern of business interface call counts, and then deduces and predicts future data, deeply conforming to the application logic and technical characteristics of artificial intelligence in the field of data prediction.
[0014] Step S4: Obtain the total call count data and total call anomaly count data of the system based on the call count trend and call anomaly ratio trend of each business interface; Obtain the historical total call count data and total call anomaly count data of the system to analyze the load trend and anomaly trend of the business interface; Step S4-1: Read the predicted call count value and predicted call anomaly count value of each business interface in the system using the unique identifier of the business interface as the index; Step S4-2: Accumulate the predicted call count values of each business interface in the system to obtain the total call count data; Accumulate the predicted call anomaly count values of each business interface in the system to obtain the total predicted call anomaly count data; Step S4-3: Sort the total call count data and total call anomaly count data in chronological order according to the regulatory data collection period T; Step S4-3-1: Based on the total call count data of each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total call count data as the y-ordinate to construct the load trend graph of the business interface; Step S4-3-2: Based on the total number of call exception times for each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total number of call exception times as the y-ordinate to construct the y-axis, and then construct the exception trend chart of the service interface.
[0015] Taking the unique identifier of the service interface as the index, summarize the calls and exception prediction values of each service interface and accumulate them to form the system total data. After sorting by time sequence, construct the load trend chart and the exception trend chart respectively, which can integrate the scattered service interface data from the macroscopic level of the system, visually present the overall call load change and exception development trend of the regulatory system, effectively make up for the limitations of single service interface analysis, and provide visual and systematic data support for comprehensively grasping the system operation status, identifying potential risks, and formulating dynamic monitoring strategies.
[0016] Step S5: Set the monitoring threshold according to the load trend and exception trend of the system, and use the monitoring threshold to perform data quality detection on the system; According to the analysis of the load trend chart and the exception trend chart, calculate the average value and variance of the service interface call times and call exception times, and set the call times monitoring threshold and the exception times monitoring threshold according to the upper limit value through the three-sigma principle; the calculation formula for setting the call times monitoring threshold is as follows: ; In the formula, E A represents the call times monitoring threshold; A avg represents the average value of the actual call times of the service interface; n represents the number of regulatory data collection periods T participating in the average value calculation; k represents the standard deviation multiple coefficient, and its value is determined by the three-sigma principle; a i represents the call times at the i-th regulatory data collection period T; The value calculation of the standard deviation multiple coefficient k uses the following formula: ; In the formula, k represents the standard deviation multiple coefficient in the three-sigma principle; A max represents the upper limit value of the service interface call times; A min represents the lower limit value of the service interface call times; The calculation formula for setting the exception times monitoring threshold is as follows: ; In the formula, E B represents the exception times monitoring threshold; B avg represents the average value of the service interface call exception times; b i represents the call exception times at the i-th regulatory data collection period T; Perform data quality detection on the system according to the call count monitoring threshold and the exception count monitoring threshold. The specific process of the data quality detection is as follows: When the predicted call count exceeds the call count monitoring threshold, issue a first quality warning signal; When the predicted call count does not exceed the call count monitoring threshold, continuously monitor the business interface calls of the system; When the predicted exception count exceeds the exception count monitoring threshold, issue a second quality warning signal; When the predicted exception count does not exceed the exception count monitoring threshold, continuously monitor the business interface calls of the system.
[0017] Based on the system load and exception trend graph, by calculating the average and variance of the business interface call count and exception count, scientifically set dynamic monitoring thresholds using the three-sigma principle, compare the predicted call count and predicted exception count with the corresponding thresholds respectively, achieve hierarchical warning and continuous monitoring, effectively avoid the lag and misjudgment risks of traditional fixed-threshold monitoring, can accurately capture abnormal fluctuations in system data quality, trigger the warning mechanism in a timely manner, and provide a quantitative and reliable judgment basis for the dynamic control of the regulatory system data quality.
[0018] Furthermore, a regulatory data quality monitoring system based on artificial intelligence, the regulatory data quality monitoring system includes a data acquisition and processing module, a data statistics and analysis module, a trend prediction and analysis module, a system data integration module, and a quality detection and warning module; The data acquisition and processing module is used to set the regulatory data acquisition period and obtain and standardize the business interface call data; the data statistics and analysis module is used to analyze the business interface load data and obtain the call count, exception count, and peak call count of the business interface; the trend prediction and analysis module is used to predict the call count trend and exception count trend of the business interface; the system data integration module is used to summarize the prediction data of each business interface and construct a load and exception trend graph; the quality detection and warning module is used to set the monitoring threshold and detect and warn the system data quality based on the threshold; The output end of the data acquisition and processing module is electrically connected to the input end of the data statistics and analysis module; the output end of the data statistics and analysis module is electrically connected to the input end of the trend prediction and analysis module; the output end of the trend prediction and analysis module is electrically connected to the input end of the system data integration module; the output end of the system data integration module is electrically connected to the input end of the quality detection and warning module; The data acquisition and processing module includes an acquisition period setting unit and a service interface data processing unit; the acquisition period setting unit is used to set the supervision data acquisition period at a fixed time interval; the service interface data processing unit is used to perform a standardized processing process of cleaning and format unification on the obtained service interface call data; The data statistics and analysis module includes a call data extraction unit and a peak data acquisition unit; the call data extraction unit is used to group and statistically extract the call times and exception times of the service interface by timestamp and service interface identifier; the peak data acquisition unit is used to obtain the peak call times of the service interface through historical data sorting; The trend prediction and analysis module includes a call times prediction unit and an exception times prediction unit; the call times prediction unit is used to predict the call times of the service interface based on the peak data and the time series model; the exception times prediction unit is used to estimate the exception times of the service interface according to the historical exception ratio and the call times prediction value; The system data integration module includes a prediction data summary unit and a trend chart construction unit; the prediction data summary unit is used to accumulate the call times and exception times prediction values of each service interface to obtain the total data; the trend chart construction unit is used to construct a trend chart with the acquisition period as the horizontal axis and the total data as the vertical axis; The quality detection and early warning module includes a monitoring threshold setting unit and a quality early warning management unit; the monitoring threshold setting unit is used to calculate and set the monitoring threshold through the three-sigma principle according to the statistical characteristics of the trend chart; the quality early warning management unit is used to compare the prediction value with the threshold and trigger an early warning signal when the threshold is exceeded.
[0019] Compared with the prior art, the beneficial effects of the present invention application are as follows: 1. By setting a fixed supervision data acquisition period, the present invention application orderly acquires and standardizes the service interface call data, constructs a dual-index service interface load analysis set with the timestamp and the unique identifier of the service interface, ensures the timeliness and standardization of data acquisition, and effectively integrates the scattered information. Compared with the traditional method, this method provides a high-quality data basis with clear structure and unified format for subsequent data statistics, trend analysis and quality monitoring, and significantly improves the analysis efficiency and accuracy.
[0020] 2. By combining weighted fusion calculation with the time series prediction model, the present invention application accurately captures the dynamic change characteristics of the service interface calls. Calculate the prediction threshold based on the peak call times and the second peak call times, and predict the call times and exception times through historical data to construct a multi-level exception monitoring system. Compared with the traditional static monitoring, this method can analyze the call trend prospectively and greatly improve the real-time performance of exception detection and the accuracy of early warning.
[0021] 3. This invention application scientifically sets dynamic monitoring thresholds based on the system load and anomaly trend chart, using the three-sigma principle to change the limitations of traditional fixed thresholds. Compared with the monitoring methods of traditional fixed thresholds or static models, it has significant intelligent advantages. By comparing the predicted values with the dynamic thresholds to achieve hierarchical early warning and continuous monitoring, it can accurately capture abnormal fluctuations in data quality, trigger early warnings in a timely manner, effectively avoid risks of lag and misjudgment, and provide quantitative and reliable judgment basis for the dynamic control of the data quality of the supervision system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic flow chart of a method for monitoring the quality of supervision data based on artificial intelligence according to this invention application; Figure 2 is a schematic structural diagram of a system for monitoring the quality of supervision data based on artificial intelligence according to this invention application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of this invention application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this invention application. Obviously, the described embodiments are only a part of the embodiments of this invention application, rather than all the embodiments. Based on the embodiments in this invention application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this invention application.
[0024] Embodiment 1: As Figure 1 shown, this invention application provides a technical solution, a method for monitoring the quality of supervision data based on artificial intelligence. The intelligent management method includes the following steps: Step S1: Set the collection period of supervision data, and obtain the call data of all service interfaces according to the collection period of supervision data to construct a service interface load analysis set; Step S1-1: Set the collection period T of supervision data, and the collection period T of supervision data is a fixed time interval; Step S1-2: Based on the set collection period T of supervision data, obtain the call data of all service interfaces according to the collection period T of supervision data through application programming interface (API) calls or log parsing; Step S1-3: Perform standardization processing on the obtained call data of service interfaces, including data cleaning and unified data format; Step S1-4: Use the timestamp as the primary index and the unique identifier of the service interface as the secondary index for the standardized call data of the service interface to construct a service interface load analysis set.
[0025] In specific implementation, for example, the supervision data collection period T is set to 1 hour, and the data of each business interface call is obtained by pulling or parsing the system log in real time through the API interface, such as the call time and return status of the user authentication interface and the data query interface. The data collection needs to ensure stability to avoid data loss caused by network fluctuations. At the same time, duplicate records need to be removed and missing fields need to be completed during the standardization process to ensure the uniqueness of the double-index structure timestamp + interface ID.
[0026] Step S2: Analyze the data of the call times and call exception times corresponding to the business interface timestamps based on the business interface load analysis set, obtain the call data of the business interface in each historical supervision data collection period of the business interface for analysis, and obtain the peak call times of the business interface. Step S2-1: Group and count the business interface call data based on the timestamp and the unique identifier of the business interface in the business interface load analysis set. Step S2-2: Analyze and extract the call times and call exception times of each timestamp of the business interface. Step S2-3: Obtain the call data of the business interface in each historical supervision data collection period of the business interface, and obtain the maximum call times of the business interface through bubble sort, which is recorded as the peak call times.
[0027] In specific implementation, through the database grouping query statement, such as the GROUP BY in SQL, count the call records according to the timestamp and the business interface ID, and extract the successful call times and the abnormal return times, such as the records with an HTTP status code other than 200. When processing historical data by bubble sort, attention should be paid to the efficiency problem when the data volume is large. It is preferably quick sort or directly obtain the maximum value using the database index to ensure the accuracy of the peak call times.
[0028] Step S3: Analyze and predict the call times trend of the business interface through the peak call times of the business interface, and analyze and obtain the call exception ratio trend according to the call exception times data corresponding to the business interface timestamps; construct a call exception change trend chart based on the call exception times data of the business interface, and analyze and predict the call exception times of the business interface in combination with the call exception times trend analysis. Step S3-1: Select the timestamp corresponding to the peak call times as the starting point, and obtain the call times data that is only less than the peak call times within the starting point and the current time period of the business interface, which is recorded as the second peak call times. Step S3-2: According to the peak call times and the second peak call times of the business interface, calculate the maximum call times within the next supervision data collection period T of the business interface through weighted fusion, which is recorded as the call times prediction threshold. Step S3-3: Obtain the call count data for each regulatory data collection period T in the history of the service interface. Select the call count prediction threshold as the upper limit of the predicted value of the service interface call count. Use the exponential smoothing method in the time series prediction model for prediction to obtain the call count for the next regulatory data collection period T of the service interface, denoted as the call count predicted value. Step S3-4: Conduct primary monitoring on the service interface based on the call count prediction threshold. The primary monitoring early warning judgment is as follows: Step S3-4-1: When the call count predicted value exceeds the call count prediction threshold, send a service interface early warning signal. Step S3-4-2: When the call count predicted value does not exceed the call count prediction threshold, continue to monitor the service interface. Step S3-5: Obtain the call anomaly ratio by dividing the number of call anomalies in each regulatory data collection period T of the service interface history by the call count in the corresponding regulatory data collection period T. Step S3-6: Analyze and calculate the average call anomaly ratio based on the call anomaly ratios in each historical regulatory data collection period T of the service interface. Step S3-7: Multiply the call count predicted value of the service interface by the average call anomaly ratio to obtain the call anomaly count of the service interface, denoted as the call anomaly count predicted value.
[0029] In specific implementation, the weighted fusion calculation can set the weight coefficient w to 0.6, with a higher peak ratio. For example, if the peak call count is 1000 times and the second peak is 800 times, then the prediction threshold is 1000×0.6 + 800×0.4 = 920 times. The time series model can use the simple exponential smoothing method with α taken as 0.3, and combine the call data of the historical 10 periods to predict the next period value. The model parameters need to be dynamically adjusted according to the business fluctuation characteristics. For example, in the high-frequency trading scenario, the α value can be reduced to enhance the weight of recent data and avoid prediction lag.
[0030] Step S4: Obtain the total call count data and total call anomaly count data of the system based on the call count trend and call anomaly ratio trend of each service interface; Obtain the total call count data and total call anomaly count data of the system history to analyze the load trend and anomaly trend of the service interface. Step S4-1: Read the call count predicted value and call anomaly count predicted value of each service interface in the system with the unique identifier of the service interface as the index. Step S4-2: Accumulate the call count predicted values of each service interface in the system to obtain the total call count data; Accumulate the call anomaly count predicted values of each service interface in the system to obtain the total call anomaly count predicted value data. Step S4-3: Sort the total call count data and total call exception count data in chronological order according to the regulatory data collection period T. Step S4-3-1: Based on the total call count data for each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total call count data as the y-ordinate to construct the load trend graph of the business interface. Step S4-3-2: Based on the total call exception count data for each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total call exception count data as the y-ordinate to construct the exception trend graph of the business interface.
[0031] In specific implementation, by traversing the interface list, the predicted values of each interface are accumulated. For example, if the system contains 100 business interfaces, the predicted call counts of each interface are added to obtain the total data. Select an appropriate chart type to construct the trend graph, with the time period on the horizontal axis and the call volume / exception volume on the vertical axis to ensure clear visualization of the trend. When constructing the trend graph, it is necessary to ensure the alignment of cross-period data to avoid deviation in trend analysis caused by collection period errors, which can be achieved by synchronizing the timestamp to the millisecond level.
[0032] Step S5: Set the monitoring threshold according to the load trend and exception trend of the system, and use the monitoring threshold to perform data quality detection on the system. According to the analysis of the load trend graph and exception trend graph, calculate the average value and variance of the business interface call count and call exception count, and set the call count monitoring threshold and exception count monitoring threshold according to the upper limit value by the three-sigma principle; the calculation formula for setting the call count monitoring threshold is as follows: ; In the formula, E A represents the call count monitoring threshold; A avg represents the average value of the actual call count of the business interface; n represents the number of regulatory data collection periods T participating in the average value calculation; k represents the standard deviation multiple coefficient, and its value is determined by the three-sigma principle; a i represents the call count at the i-th regulatory data collection period T. The value calculation of the standard deviation multiple coefficient k uses the following formula: ; In the formula, k represents the standard deviation multiple coefficient in the three-sigma principle; A max represents the upper limit value of the business interface call count; A min represents the lower limit value of the business interface call count. The calculation formula for setting the exception count monitoring threshold is as follows: ; In the formula, E B represents the monitoring threshold for the number of exceptions; B avg represents the average value of the number of exceptions in the business interface calls; b i represents the number of call exceptions at the i-th regulatory data collection period T; Perform data quality detection on the system according to the call count monitoring threshold and the exception count monitoring threshold. The specific process of the data quality detection is as follows: When the predicted call count exceeds the call count monitoring threshold, issue a first quality warning signal; When the predicted call count does not exceed the call count monitoring threshold, continuously monitor the business interface calls of the system; When the predicted exception count of calls exceeds the exception count monitoring threshold, issue a second quality warning signal; When the predicted exception count of calls does not exceed the exception count monitoring threshold, continuously monitor the business interface calls of the system.
[0033] In a specific implementation, taking the total call count data of a regulatory system for 30 consecutive regulatory data collection periods T = 1 hour as an example, calculate its average value μ as 8000 times and the standard deviation σ as 300 times. According to the three-sigma principle, set the upper call count monitoring threshold as μ + 3σ = 8900 times. If the predicted call count for the next period is 9200 times, exceeding the threshold will trigger the first quality warning signal, indicating abnormal system load; if the predicted value is 8500 times, continue to monitor.
[0034] For the number of call exceptions, assume that the historical average number of exceptions is 200 times and the standard deviation is 50 times. Set the upper exception count monitoring threshold as 200 + 3×50 = 350 times. If the predicted exception count for a certain period is 380 times, exceeding the threshold will issue a second quality warning signal, indicating data quality risk; if it is 300 times, maintain the continuous monitoring state.
[0035] Embodiment 2, as Figure 2 shown, the present invention provides a regulatory data quality monitoring system based on artificial intelligence. The intelligent management system includes a data collection and processing module, a data statistics and analysis module, a trend prediction and analysis module, a system data integration module, and a quality detection and warning module; The data acquisition and processing module is used to set the supervision data acquisition period and obtain and standardize the business interface call data; the data statistics and analysis module is used to analyze the business interface load data and obtain the call times, exception times and peak call times of the business interface; the trend prediction and analysis module is used to predict the call times trend and exception times trend of the business interface; the system data integration module is used to summarize the prediction data of each business interface and construct a load and exception trend graph; the quality detection and early warning module is used to set the monitoring threshold and detect and give early warning to the system data quality based on the threshold. The output end of the data acquisition and processing module is electrically connected to the input end of the data statistics and analysis module; the output end of the data statistics and analysis module is electrically connected to the input end of the trend prediction and analysis module; the output end of the trend prediction and analysis module is electrically connected to the input end of the system data integration module; the output end of the system data integration module is electrically connected to the input end of the quality detection and early warning module. The data acquisition and processing module includes an acquisition period setting unit and a business interface data processing unit; the acquisition period setting unit is used to set the supervision data acquisition period at a fixed time interval; the business interface data processing unit is used to perform a standardization processing process of cleaning and format unification on the obtained business interface call data. The data statistics and analysis module includes a call data extraction unit and a peak data acquisition unit; the call data extraction unit is used to group and count and extract the call times and exception times of the business interface according to the time stamp and business interface identifier; the peak data acquisition unit is used to obtain the peak call times of the business interface by sorting historical data. The trend prediction and analysis module includes a call times prediction unit and an exception times prediction unit; the call times prediction unit is used to predict the call times of the business interface based on the peak data and the time series model; the exception times prediction unit is used to estimate the exception times of the business interface according to the historical exception ratio and the call times prediction value. The system data integration module includes a prediction data summary unit and a trend graph construction unit; the prediction data summary unit is used to accumulate the call times and exception times prediction values of each business interface to obtain the total data; the trend graph construction unit is used to construct a trend graph with the acquisition period as the horizontal axis and the total data as the vertical axis. The quality detection and early warning module includes a monitoring threshold setting unit and a quality early warning management unit; the monitoring threshold setting unit is used to calculate and set the monitoring threshold according to the statistical characteristics of the trend graph through the three-sigma principle; the quality early warning management unit is used to compare the prediction value with the threshold and trigger an early warning signal when the threshold is exceeded.
[0036] For those skilled in the art, it is obvious that the present invention application is not limited to the details of the above exemplary embodiments, and the present invention application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention application. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A method for monitoring the quality of regulatory data based on artificial intelligence, characterized in that: The regulatory data quality monitoring method includes the following steps: Step S1: Set the regulatory data collection period, and obtain the call data of all service interfaces according to the regulatory data collection period to construct a service interface load analysis set; Step S2: Analyze the obtained service interface load analysis set to obtain the call times and call exception times corresponding to the service interface timestamps. Obtain the call data of the service interfaces in each historical regulatory data collection period of the service interfaces for analysis to obtain the peak call times of the service interfaces; Step S3: Analyze and predict the call times trend of the service interface through the peak call times of the service interface. Analyze the call exception ratio trend according to the call exception times data corresponding to the service interface timestamps; Construct a call exception change trend chart based on the call exception times data of the service interface, and analyze and predict the call exception times of the service interface in combination with the call exception times trend; Step S4: Obtain the total call times data and total call exception times data of the system according to the call times trend and call exception ratio trend of each service interface; Obtain the total historical call times data and total call exception times data of the system to analyze the load trend and exception trend of the service interfaces; Step S5: Set monitoring thresholds according to the load trend and exception trend of the system, and use the monitoring thresholds to detect the data quality of the system.
2. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 1, wherein: The specific steps of step S1 are as follows: Step S1-1: Set the regulatory data collection period T, and the regulatory data collection period T is a fixed time interval; Step S1-2: Based on the set regulatory data collection period T, obtain the call data of all service interfaces by calling the application program service interface API or parsing the logs according to the regulatory data collection period T; Step S1-3: Perform standardization processing on the obtained service interface call data, including data cleaning and unified data format; Step S1-4: Use the timestamps as the primary index and the unique identifiers of the service interfaces as the secondary index for the standardized service interface call data to construct a service interface load analysis set.
3. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 2, wherein: The specific steps of step S2 are as follows: Step S2-1: Based on the service interface load analysis set, group and count the service interface call data according to the timestamps and the unique identifiers of the service interfaces; Step S2-2: Analyze and extract the call times and call exception times of each timestamp of the service interface; Step S2-3: Obtain the call data of the service interfaces in each historical regulatory data collection period of the service interfaces, and obtain the maximum call times of the service interfaces through bubble sorting, which is recorded as the peak call times.
4. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 3, wherein: The specific steps of step S3 are as follows: Step S3-1: Select the timestamp corresponding to the peak call times as the starting point, and obtain the call times data that is only less than the peak call times within the starting point and the current time period of the service interface, which is recorded as the second peak call times; Step S3-2: According to the peak call times and the second peak call times of the service interface, calculate the maximum call times within the next regulatory data collection period T of the service interface through weighted fusion, which is recorded as the call times prediction threshold.
5. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 4, wherein: In step S3, it also includes: Step S3-3: Obtain the call count data for each regulatory data collection period T in the history of the business interface. Select the call count prediction threshold as the upper limit of the predicted value of the business interface call count. Perform prediction using the exponential smoothing method in the time series prediction model to obtain the call count for the next regulatory data collection period T of the business interface, denoted as the call count predicted value. The prediction calculation using the exponential smoothing method is as follows: ; In the formula, F T+1 represents the predicted value of the number of calls in the next regulatory data collection cycle T of the service interface; a represents the average coefficient, and the value range is in the interval of 1 and 0; Y T represents the actual number of calls in the current regulatory data collection cycle T; F T represents the predicted value of the number of calls in the current regulatory data collection cycle T; Step S3-4: Conduct primary monitoring on the business interface based on the call count prediction threshold. The primary monitoring warning judgment is as follows: Step S3-4-1: When the call count predicted value exceeds the call count prediction threshold, issue a warning signal for the business interface. Step S3-4-2: When the call count predicted value does not exceed the call count prediction threshold, continue to monitor the business interface. Step S3-5: Obtain the call anomaly ratio by dividing the number of call anomalies in each regulatory data collection period T in the history of the business interface by the call count in the corresponding regulatory data collection period T. Step S3-6: Analyze and calculate the average call anomaly ratio based on the call anomaly ratios in each historical regulatory data collection period T of the business interface. Step S3-7: Multiply the call count predicted value of the business interface by the average call anomaly ratio to obtain the number of call anomalies of the business interface, denoted as the call anomaly count predicted value.
6. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 5, characterized in that: The specific steps of step S4 are as follows: Step S4-1: Using the unique identifier of the business interface as an index, read the call count predicted value and the call anomaly count predicted value of each business interface in the system. Step S4-2: Accumulate the call count predicted values of each business interface in the system to obtain the total call count data; accumulate the call anomaly count predicted values of each business interface in the system to obtain the total call anomaly count predicted value data. Step S4-3: Sort the total call count data and the total call anomaly count predicted value data in chronological order according to the regulatory data collection period T. Step S4-3-1: Based on the total call count data for each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total call count data as the y-ordinate, and then construct the load trend graph of the business interface. Step S4-3-2: Based on the total call anomaly count data for each regulatory data collection period T, construct the x-axis with the regulatory data collection period T as the abscissa and the total call anomaly count data as the y-ordinate, and then construct the anomaly trend graph of the business interface.
7. The method for monitoring the quality of regulatory data based on artificial intelligence according to claim 6, characterized in that: In step S5, based on the analysis of the load trend graph and the anomaly trend graph, calculate the average value and variance of the business interface call count and the call anomaly count. Set the call count monitoring threshold and the anomaly count monitoring threshold according to the three-sigma principle based on the upper limit value. The calculation formula for setting the call count monitoring threshold is as follows: ; Where, E A represents the monitoring threshold for the number of invocations; A avg represents the average value of the actual number of invocations of the service interface; n represents the number of regulatory data collection cycles T participating in the average value calculation; k represents the standard deviation multiple coefficient, and its value is determined by the three-sigma principle; a i represents the number of invocations at the i-th regulatory data collection cycle T; The calculation of the value of the standard deviation multiple coefficient k is as follows: ; Where k represents the standard deviation multiple coefficient in the three-sigma principle; A max represents the upper limit value of the number of business interface calls; A min represents the lower limit value of the number of business interface calls; The calculation formula for setting the anomaly count monitoring threshold is as follows: ; Where E B represents the monitoring threshold for the number of exceptions; B avg represents the average number of exceptions in the business interface calls; b i represents the number of call exceptions at the i-th regulatory data collection period T; Conduct data quality detection on the system based on the call count monitoring threshold and the anomaly count monitoring threshold. The specific process of the data quality detection is as follows: When the predicted call count exceeds the call count monitoring threshold, a first quality warning signal is issued; When the predicted call count does not exceed the call count monitoring threshold, the monitoring of the business interface calls of the system continues; When the predicted abnormal call count exceeds the abnormal call count monitoring threshold, a second quality warning signal is issued; When the predicted abnormal call count does not exceed the abnormal call count monitoring threshold, the monitoring of the business interface calls of the system continues.
8. A regulatory data quality monitoring system based on artificial intelligence, which is applied to a regulatory data quality monitoring method based on artificial intelligence described in any one of claims 1-7, characterized in that: The regulatory data quality monitoring system includes a data collection and processing module, a data statistics and analysis module, a trend prediction and analysis module, a system data integration module, and a quality detection and warning module; The data collection and processing module is used to set the regulatory data collection period and obtain and standardize the business interface call data; the data statistics and analysis module is used to analyze the business interface load data and obtain the call count, abnormal count, and peak call count of the business interface; the trend prediction and analysis module is used to predict the call count trend and abnormal count trend of the business interface; The system data integration module is used to summarize the prediction data of each business interface and construct a load and abnormal trend graph; The quality detection and warning module is used to set the monitoring threshold and detect and warn the system data quality based on the threshold; The output end of the data collection and processing module is electrically connected to the input end of the data statistics and analysis module; the output end of the data statistics and analysis module is electrically connected to the input end of the trend prediction and analysis module; the output end of the trend prediction and analysis module is electrically connected to the input end of the system data integration module; the output end of the system data integration module is electrically connected to the input end of the quality detection and warning module.
9. A regulatory data quality monitoring system based on artificial intelligence according to claim 8, wherein: The data collection and processing module includes a collection period setting unit and a business interface data processing unit; the collection period setting unit is used to set the regulatory data collection period at a fixed time interval; the business interface data processing unit is used to perform a standardization processing process of cleaning and format unification on the obtained business interface call data; The data statistics and analysis module includes a call data extraction unit and a peak data acquisition unit; the call data extraction unit is used to group and statistically extract the call count and abnormal count of the business interface by timestamp and business interface identifier; the peak data acquisition unit is used to obtain the peak call count of the business interface by sorting historical data; The trend prediction and analysis module includes a call count prediction unit and an abnormal count prediction unit; the call count prediction unit is used to predict the call count of the business interface based on the peak data and the time series model; the abnormal count prediction unit is used to estimate the abnormal count of the business interface according to the historical abnormal ratio and the predicted call count value.
10. A regulatory data quality monitoring system based on artificial intelligence according to claim 8, wherein: The system data integration module includes a prediction data summary unit and a trend graph construction unit; the prediction data summary unit is used to accumulate the predicted call count and abnormal count values of each business interface to obtain the total data; The trend chart construction unit is used to construct a trend chart with the acquisition period as the horizontal axis and the total data as the vertical axis; The quality inspection and early warning module includes a monitoring threshold setting unit and a quality early warning management unit; The monitoring threshold setting unit is used to calculate and set the monitoring threshold according to the statistical characteristics of the trend chart through the three-sigma principle; the quality early warning management unit is used to compare the predicted value with the threshold and trigger an early warning signal when the threshold is exceeded.
Citation Information
Patent Citations
Network security supervision system and method based on machine learning and abnormal behavior analysis
CN116614277A
Performance monitoring and exception repairing method and device, equipment and medium
CN119473680A
Software application performance optimization method based on big data analysis
CN119917390A
Enterprise operation supervision system based on data analysis
CN120181667A
Method for monitoring time-series data, System for monitoring time-series data and Computer program for the same
KR102011689B1
Cited By
Intelligent configuration management system for vacuum process parameters
CN120560211A