Data access quality intelligent analysis method based on periodic analysis
By identifying the natural cycle and life cycle of data, analyzing the trend of data access quality and building a prediction model, the problem of low efficiency in data access quality management in the prior art is solved, and dynamically adjusting the data access strategy is realized to improve data access quality and system adaptability.
Patent Information
- Application Number
- CN202510511817.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to effectively integrate data from different sources, resulting in information silos, and the static threshold cannot adapt to a dynamically changing environment, resulting in low efficiency in data access quality management.
By collecting and unifying data from the data sources to be collected, statistical methods are used to identify the natural cycle and life cycle of the data, analyzing the trend of data access quality, extracting key feature sets, building prediction models and setting dynamic quality thresholds, and dynamically adjusting the data access strategy.
It improves data access quality and system adaptability, overcomes the limitations of static thresholds, and enhances the system's adaptability to changing environments.
Smart Images

Figure CN120030286A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data quality management, and in particular to a data access quality intelligent analysis method based on periodic analysis. Background Art
[0002] With the advent of the big data era, enterprises and organizations are increasingly relying on data-driven decision-making. High-quality data is the basis for effective decision-making. Data from different sources and types makes data access quality complicated, and unified management and analysis become necessary. However, current technologies are often unable to effectively integrate data from different sources, resulting in information islands, and many existing systems use static thresholds to judge data quality, which cannot adapt to dynamically changing environments.
[0003] Therefore, the present invention provides a data access quality intelligent analysis method based on periodic analysis. Summary of the invention
[0004] The present invention provides an intelligent analysis method for data access quality based on periodic analysis. The method collects and unifies data from the data source to be collected, applies statistical methods to identify the natural cycle and life cycle of the data, and on this basis, analyzes the trend of data access quality, extracts key feature sets, builds a prediction model and sets dynamic quality thresholds. Through real-time monitoring and abnormal judgment, the data access strategy is dynamically adjusted to improve data access quality and system adaptability, overcome the limitations of static thresholds, and enhance the system's adaptability to changing environments.
[0005] The present invention provides a data access quality intelligent analysis method based on periodic analysis, comprising: Step 1: Collect data from the data source to be collected, unify the data format, obtain unified data, apply statistical methods to identify the natural cycle and life cycle of the unified data, and then obtain the periodic pattern of the unified data; Step 2: Analyze the trend of data access quality of unified data in different periods based on the periodic pattern, extract key feature sets, and then identify the peak period and trough period of data access quality; Step 3: Based on the key feature set, a machine learning algorithm is selected to construct a prediction model for data access quality, and access quality thresholds are set according to the peak period and the trough period, thereby obtaining a quality discrimination model; Step 4: Input the actual data into the quality judgment model, make quality anomaly judgments, and dynamically adjust the data access strategy based on the anomaly judgment results.
[0006] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which collects data from a data source to be collected, unifies the data format, obtains unified data, and applies statistical methods to identify the natural cycle and life cycle of the unified data, thereby obtaining a periodic pattern of the unified data, including: Obtaining user needs, determining multiple data sources to be collected according to the user needs, collecting data from the data sources to be collected, and unifying the data formats to obtain unified data; Applying statistical methods to analyze the unified data to obtain statistical analysis results, extracting trend components from the statistical analysis results to obtain the life cycle of the unified data, and extracting seasonal components from the statistical analysis results to obtain the natural cycle of the unified data; Combining the life cycle with the natural cycle results in a cyclical pattern of unified data.
[0007] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which analyzes the trend of data access quality of unified data in different periods based on the periodic pattern, extracts key feature sets, and further identifies the peak period and trough period of data access quality, including: Determining a decomposition method of the unified data in different periods from a data decomposition library according to the periodic pattern, and decomposing the unified data using the decomposition method to obtain a decomposition result; Draw a trend graph, a seasonal graph, and a residual graph of the unified data according to the decomposition results, and extract a key feature set of the unified data based on the trend graph, the seasonal graph, and the residual graph; All peak periods and valley periods of data access quality are identified based on the key feature set.
[0008] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which draws a trend graph, a seasonal graph and a residual graph of unified data according to the decomposition result, and extracts a key feature set of unified data based on the trend graph, the seasonal graph and the residual graph, including: Observe the trend graph, determine the long-term change direction of the data access quality, calculate the slope of the trend line, obtain the trend strength, identify the trend change point from the trend graph, and obtain the first feature by combining the long-term change direction, trend strength and trend change point; Observe the seasonal graph to determine the change pattern of the unified data, calculate the autocorrelation coefficient of the unified data based on the seasonal graph, determine the lag period, calculate the seasonal index of the unified data in combination with the lag period and the change pattern, and then determine the second feature; Observe the residual graph, evaluate the amplitude and frequency of random fluctuations, identify outliers in the residuals, perform cycle and trend analysis on the outliers, and determine the third feature based on the cycle and trend analysis results; The first feature, the second feature and the third feature are combined to obtain a key feature set of unified data.
[0009] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which selects a machine learning algorithm based on the key feature set to construct a prediction model for data access quality, and sets access quality thresholds according to the peak period and the trough period, thereby obtaining a quality discrimination model, including: The Fire Cicada algorithm is used to screen the key feature set to obtain the optimal feature, and a machine learning algorithm is selected from the feature-algorithm library based on the optimal feature to build a prediction model for data access quality; The access quality peak threshold is set according to the peak period, and the access quality trough threshold is set according to the trough period. The quality discrimination model is obtained by combining the prediction model with the access quality peak threshold and the access quality trough threshold.
[0010] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which uses the Fire Cicada algorithm to screen the key feature set and obtain the optimal features, including: Based on the key feature set combined with the Fire Cicada algorithm, a fitness function is defined to evaluate the quality of each key feature subset, and the fitness value of each key feature subset is calculated, and the attractiveness formula of the key feature subset is set according to the fitness value and the Fire Cicada algorithm; , where X represents the mutual attraction between key feature subsets; represents the initial attraction between subsets of key features; Indicates the speed at which the mutual attraction decays with distance; represents the distance between key feature subset i and key feature subset j; T represents the total running time of the Fire Cicada algorithm; t represents the process of the Fire Cicada algorithm; represents the adjustment parameter of the change in fitness on the degree of mutual attraction; Indicates the difference between the current fitness and the previous fitness; Indicates the maximum fitness value in the current group; Indicates the degree to which the influence distance weakens the mutual attraction; D indicates the maximum possible distance between key feature subsets; represents the random noise term; The key feature set is screened according to the attraction formula to obtain the optimal feature.
[0011] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which sets an access quality peak threshold according to a peak period and sets an access quality valley threshold according to a valley period, including: Classify the collected data sources into types, and determine the standard threshold range corresponding to each type classification result from the type-data quality table according to the type classification results; Based on the standard threshold range, a peak access quality threshold is set in combination with a peak period, and a valley access quality threshold is set in combination with a valley period.
[0012] The present invention provides a data access quality intelligent analysis method based on periodic analysis, which inputs actual data into a quality discrimination model, performs quality abnormality judgment, and dynamically adjusts the data access strategy according to the abnormality judgment result, including: Input the actual data into the quality judgment model, and judge the abnormality of data access quality based on the output of the model; Record the abnormal judgment results, evaluate the data access strategy, and analyze the abnormal causes; Based on the abnormal cause and the abnormal judgment result, an adjustment plan is formulated to adjust the data access strategy.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: by collecting and unifying data from the data sources to be collected, applying statistical methods to identify the natural cycles and life cycles of the data, and on this basis, analyzing the trend of data access quality, extracting key feature sets, building a prediction model and setting dynamic quality thresholds, and dynamically adjusting data access strategies through real-time monitoring and anomaly judgment to improve data access quality and system adaptability, overcome the limitations of static thresholds, and enhance the system's adaptability to a changing environment.
[0014] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0015] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 It is a flow chart of a method for intelligent analysis of data access quality based on periodic analysis provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0018] Embodiment 1: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, such as Figure 1 As shown, including: Step 1: Collect data from the data source to be collected, unify the data format, obtain unified data, apply statistical methods to identify the natural cycle and life cycle of the unified data, and then obtain the periodic pattern of the unified data; Step 2: Analyze the trend of data access quality of unified data in different periods based on the periodic pattern, extract key feature sets, and then identify the peak period and trough period of data access quality; Step 3: Based on the key feature set, a machine learning algorithm is selected to construct a prediction model for data access quality, and access quality thresholds are set according to the peak period and the trough period, thereby obtaining a quality discrimination model; Step 4: Input the actual data into the quality judgment model, make quality anomaly judgments, and dynamically adjust the data access strategy based on the anomaly judgment results.
[0019] In this embodiment, the data source to be collected refers to a source that can provide relevant data, which may include: sales records: historical sales data, customer purchase behavior, market data: competitor sales data, market trend reports, social media: user feedback, comments and interaction data.
[0020] In this embodiment, the life cycle refers to different stages of data or products within a period of time. For example, the life cycle of a product can be divided into an introduction period, a growth period, a maturity period, and a decline period.
[0021] In this embodiment, natural cycles refer to the inherent periodic fluctuations in the data, not limited to seasonality. For example, some consumer products may have increased sales in certain months of each year (such as the beginning of the school year).
[0022] In this embodiment, the periodic pattern refers to identifying the overall fluctuation pattern of the data through comprehensive life cycle and natural cycle analysis. For example, the sales of a certain product peak in spring and autumn every year, and are relatively low in summer and winter.
[0023] In this embodiment, by obtaining user needs, multiple data sources to be collected are determined, and the format is unified to generate unified data. Statistical methods are applied to analyze the unified data, extract trends and seasonal components, and then identify the life cycle and natural cycle of the data. Finally, this information is combined to derive the periodic pattern of the unified data.
[0024] In this embodiment, the key feature set is important features extracted from the trend graph, the seasonal graph, and the residual graph, for example: the annual growth rate of sales, the average sales value of each season, and the standard deviation of the residual (used to evaluate the volatility of data access quality).
[0025] In this embodiment, peak period and trough period refer to time periods showing significant growth or decline in the data. For example, peak period: during summer and holidays (such as Christmas), beverage sales increase significantly, and trough period: during winter or non-holidays, beverage sales decrease significantly.
[0026] In this embodiment, trend analysis is performed by drawing a trend graph, a seasonal graph, and a residual graph of unified data.
[0027] In this embodiment, a decomposition method of unified data in different periods is determined based on a periodic pattern, the unified data is decomposed to obtain a decomposition result, and a key feature set is extracted by drawing a trend graph, a seasonal graph, and a residual graph, thereby identifying the peak and trough periods of data access quality.
[0028] In this embodiment, the key feature set is obtained by integrating the first feature, the second feature and the third feature.
[0029] In this embodiment, the input of the prediction model is the optimal feature set (such as "user behavior", "timestamp", "access method"), and the output is: predicted access quality (such as "high quality", "low quality" or a specific quality score). For example, the input features are user behavior, timestamp and access method, and the output is an access quality score (such as 1-100 points).
[0030] In this embodiment, the access quality peak threshold is the critical value of the access quality during the peak period. Exceeding this value indicates good access quality. For example, assuming that the peak threshold of transaction data is set to 98%, that is, during the peak period, the quality score of the transaction data must exceed 98%.
[0031] In this embodiment, the access quality valley threshold is a critical value of the access quality during the valley period. A value lower than this value indicates poor access quality. For example, assuming that the valley threshold of the sensor data is set to 85%, the quality score of the sensor data must be lower than 85% during the valley period.
[0032] In this embodiment, the input of the quality discrimination model is (access quality score) and the set peak threshold and trough threshold, and the output is: access quality discrimination result (such as "excellent", "good", "poor"). For example, the input is: the access quality score is 75 points, the peak threshold is 80, and the trough threshold is 50. The output is: the judgment is "good" because 75 points is higher than the trough threshold but lower than the peak threshold.
[0033] In this embodiment, the Fire Cicada algorithm is used to perform feature screening on the key feature set, extract the optimal features, select a suitable machine learning algorithm from the feature-algorithm library, build a prediction model for data access quality, set access quality thresholds for peak and trough periods, and combine the prediction model with these thresholds to form a quality discrimination model.
[0034] In this embodiment, the quality abnormality judgment is to evaluate the actual data through the quality discrimination model to determine whether there is an abnormality, which usually means that the data quality is lower than the set threshold. For example, assuming that the access quality score of a data source is 70 points, and the set valley threshold is 75 points. Since the score is lower than the threshold, it is judged that the data source is abnormal.
[0035] In this embodiment, the abnormal judgment result is to determine whether the data access quality is abnormal based on the result output by the model, which is usually divided into "normal" or "abnormal". For example, in the above example, the judgment result is "abnormal" because 70 points is lower than the set low threshold.
[0036] In this embodiment, the evaluation of data access strategy is to analyze and evaluate the existing data access strategy to determine its effectiveness and adaptability, and to identify factors that may cause abnormal data quality. For example, suppose the data access strategy stipulates that data should be collected once an hour, but the strategy is not effectively implemented during peak hours, resulting in data missing or delays, thereby affecting data quality.
[0037] In this embodiment, actual data is input into the quality judgment model, and an abnormal judgment of data access quality is performed based on the model output results. After the abnormal judgment results are recorded, the data access strategy is evaluated and the abnormal causes are analyzed. Based on the abnormal causes and the judgment results, corresponding adjustment plans are formulated to optimize the data access strategy.
[0038] The working principle and beneficial effects of the above technical solution are: by collecting and unifying data from the data source to be collected, applying statistical methods to identify the natural cycle and life cycle of the data, and on this basis, analyzing the trend of data access quality, extracting key feature sets, building a prediction model and setting dynamic quality thresholds, and dynamically adjusting data access strategies through real-time monitoring and anomaly judgment to improve data access quality and system adaptability, overcome the limitations of static thresholds, and enhance the system's adaptability to a changing environment.
[0039] Embodiment 2: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which collects data from a data source to be collected, unifies the format of the data, obtains unified data, and applies statistical methods to identify the natural cycle and life cycle of the unified data, thereby obtaining a periodic pattern of the unified data, including: Obtaining user needs, determining multiple data sources to be collected according to the user needs, collecting data from the data sources to be collected, and unifying the data formats to obtain unified data; Applying statistical methods to analyze the unified data to obtain statistical analysis results, extracting trend components from the statistical analysis results to obtain the life cycle of the unified data, and extracting seasonal components from the statistical analysis results to obtain the natural cycle of the unified data; Combining the life cycle with the natural cycle results in a cyclical pattern of unified data.
[0040] In this embodiment, user demand refers to the information or problem that the user hopes to obtain through data analysis. For example, a retail company hopes to understand the sales trend of its products in order to optimize inventory management and marketing strategies.
[0041] In this embodiment, the unified data is a consistent data set formed by formatting and standardizing the data from the data source to be collected. For example, the sales data from different stores are cleaned and formatted to form a unified data set containing sales records of all stores.
[0042] In this embodiment, the statistical analysis results are conclusions and information obtained after statistical methods are applied to the unified data, for example, calculating average sales, standard deviation, trend line, etc.
[0043] In this embodiment, trend component extraction refers to identifying the long-term change trend of data from the statistical analysis results. For example, the upward trend of sales in the past few years is determined by linear regression analysis.
[0044] In this embodiment, seasonal component extraction refers to extracting the periodic fluctuation part from the data, which is usually related to the season, month or specific time period. For example, retail data will have sales peaks around the holidays every year.
[0045] The working principle and beneficial effects of the above technical solution are: by obtaining user needs, determining multiple data sources to be collected, and unifying the format to generate unified data, applying statistical methods to analyze the unified data, extracting trends and seasonal components, and then identifying the life cycle and natural cycle of the data, and finally integrating this information to derive the periodic pattern of the unified data to support subsequent analysis and decision-making, ensure the relevance and effectiveness of the collected data, and improve data integration efficiency.
[0046] Embodiment 3: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which analyzes the trend of data access quality of unified data in different periods based on the periodic pattern, extracts key feature sets, and further identifies the peak period and trough period of data access quality, including: Determining a decomposition method of the unified data in different periods from a data decomposition library according to the periodic pattern, and decomposing the unified data using the decomposition method to obtain a decomposition result; Draw a trend graph, a seasonal graph, and a residual graph of the unified data according to the decomposition results, and extract a key feature set of the unified data based on the trend graph, the seasonal graph, and the residual graph; All peak periods and valley periods of data access quality are identified based on the key feature set.
[0047] In this embodiment, the decomposition method is a technique used to decompose time series data into different components such as trend, seasonality, and residual.
[0048] In this embodiment, the data decomposition library is a resource library containing different decomposition methods and algorithms, which can be used to select a suitable decomposition method. For example, the library can contain a package based on seasonal decomposition.
[0049] In this embodiment, the decomposition result is the various components obtained after the decomposition method is used. For example, after decomposition, you may obtain trend components (such as sales increasing year by year), seasonal components (such as peak sales in summer), and residual components (random fluctuations).
[0050] In this embodiment, the trend graph is a graph showing the long-term trend of data. For example, a graph showing the gradual increase in sales over time; the seasonal graph is a graph showing the seasonal fluctuations of data. For example, sales plotted by month show the peak sales in summer every year; the residual graph is a graph showing the part of the data that is not explained by the trend and seasonal components, and is usually used to analyze random fluctuations.
[0051] The working principle and beneficial effects of the above technical solution are: determining the decomposition method of unified data in different periods based on the periodic pattern, decomposing the unified data to obtain the decomposition results, and extracting the key feature set by drawing trend charts, seasonal charts and residual charts, so as to identify the peak and trough periods of data access quality, support subsequent analysis and decision-making, improve the accuracy of data access quality analysis, optimize decision support, and enhance the flexibility of data processing.
[0052] Embodiment 4: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which draws a trend graph, a seasonal graph, and a residual graph of unified data according to the decomposition result, and extracts a key feature set of the unified data based on the trend graph, the seasonal graph, and the residual graph, including: Observe the trend graph, determine the long-term change direction of the data access quality, calculate the slope of the trend line, obtain the trend strength, identify the trend change point from the trend graph, and obtain the first feature by combining the long-term change direction, trend strength and trend change point; Observe the seasonal graph to determine the change pattern of the unified data, calculate the autocorrelation coefficient of the unified data based on the seasonal graph, determine the lag period, calculate the seasonal index of the unified data in combination with the lag period and the change pattern, and then determine the second feature; Observe the residual graph, evaluate the amplitude and frequency of random fluctuations, identify outliers in the residuals, perform cycle and trend analysis on the outliers, and determine the third feature based on the cycle and trend analysis results; The first feature, the second feature and the third feature are combined to obtain a key feature set of unified data.
[0053] In this embodiment, the long-term change direction is to observe a trend graph (such as a time series graph) to determine whether the data is rising, falling, or remaining stable. For example, sales data has increased year by year in the past five years, indicating that the long-term change direction is rising.
[0054] In this embodiment, the trend strength is quantified by calculating the slope of the trend line to quantify the strength of the change. The larger the slope, the stronger the change. For example, if the sales volume of a product increases from 1,000 to 5,000, the slope is 4, indicating strong sales growth.
[0055] In this embodiment, the trend change point is to identify the turning point in the trend graph, such as a sudden drop in sales at a certain point in time, which may indicate a market change or increased competition.
[0056] In this embodiment, the second characteristics are the variation pattern, the autocorrelation coefficient, the lag period and the seasonal index.
[0057] In this embodiment, the variation pattern is determined by a seasonal graph (such as a periodic graph) to determine the variation pattern of the data. For example, electricity consumption may fluctuate significantly in summer and winter.
[0058] In this embodiment, the autocorrelation coefficient is calculated to evaluate the correlation of the data and determine the lag period of the data. For example, if the autocorrelation coefficient is significant in both lag period 1 and lag period 12, it indicates that the data has seasonality.
[0059] In this embodiment, the seasonal index is calculated by combining the lag period and the change pattern to indicate the performance of a particular season. For example, if the sales index in summer is 1.5, it means that the sales in summer are 50% higher than those in the base period.
[0060] In this embodiment, the cycle and trend analysis is to evaluate the amplitude and frequency of random fluctuations by observing the residual graph and identify outliers. For example, the sales data of a certain month is abnormally high, which may be caused by promotional activities.
[0061] In this embodiment, the outlier analysis is to perform period and trend analysis on the identified outliers to determine whether they are part of a long-term trend or are just short-term fluctuations.
[0062] The working principle and beneficial effects of the above technical solution are: by observing the trend graph, the long-term change direction of data access quality is determined, the slope of the trend line is calculated to obtain the trend strength, and the trend change point is identified, thereby forming the first feature. The change pattern and autocorrelation coefficient are analyzed through the seasonal graph, the lag period is determined and the seasonal index is calculated to obtain the second feature. The random fluctuations and outliers are evaluated through the residual graph, and the cycle and trend analysis is performed to form the third feature. Finally, these features are combined to obtain the key feature set of unified data, improve the accuracy of data access quality analysis, and support more effective monitoring and decision optimization.
[0063] Embodiment 5: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which selects a machine learning algorithm based on the key feature set to build a prediction model for data access quality, and sets access quality thresholds according to the peak period and the trough period, thereby obtaining a quality discrimination model, including: The Fire Cicada algorithm is used to screen the key feature set to obtain the optimal feature, and a machine learning algorithm is selected from the feature-algorithm library based on the optimal feature to build a prediction model for data access quality; The access quality peak threshold is set according to the peak period, and the access quality trough threshold is set according to the trough period. The quality discrimination model is obtained by combining the prediction model with the access quality peak threshold and the access quality trough threshold.
[0064] In this embodiment, feature screening is to select the features that have the greatest influence on the prediction target from a large number of features through an algorithm (such as the Fire Cicada algorithm), generate multiple feature subsets, calculate the fitness value of each feature subset, and select the optimal feature subset based on the fitness value. For example, assuming there are 10 features (such as user behavior, timestamp, access method, etc.), after screening through the Fire Cicada algorithm, 3 optimal features (such as user behavior, timestamp, access method) are selected.
[0065] In this embodiment, the optimal feature is a feature set that can provide the best prediction performance after feature screening. For example, the optimal features may be "user behavior", "timestamp" and "access method".
[0066] In this embodiment, the feature-algorithm library is a collection of multiple machine learning algorithms. Users can select appropriate algorithms for model building based on the features. For example, common machine learning algorithms include: linear regression, decision tree, random forest, support vector machine, and neural network.
[0067] The working principle and beneficial effects of the above technical solution are: use the Fire Cicada algorithm to screen the key feature set, extract the optimal features, select the appropriate machine learning algorithm from the feature-algorithm library, build a prediction model for data access quality, set access quality thresholds for peak and trough periods, and combine the prediction model with these thresholds to form a quality discrimination model to evaluate data access quality, improve the accuracy of data access quality prediction, and enhance anomaly detection capabilities.
[0068] Embodiment 6: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which uses the Fire Cicada algorithm to screen the key feature set and obtain the optimal feature, including: Based on the key feature set combined with the Fire Cicada algorithm, a fitness function is defined to evaluate the quality of each key feature subset, and the fitness value of each key feature subset is calculated, and the attractiveness formula of the key feature subset is set according to the fitness value and the Fire Cicada algorithm; , where X represents the mutual attraction between key feature subsets; represents the initial attraction between subsets of key features; Indicates the speed at which the mutual attraction decays with distance; represents the distance between key feature subset i and key feature subset j; T represents the total running time of the Fire Cicada algorithm; t represents the process of the Fire Cicada algorithm; represents the adjustment parameter of the change in fitness on the degree of mutual attraction; Indicates the difference between the current fitness and the previous fitness; Indicates the maximum fitness value in the current group; Indicates the degree to which the influence distance weakens the mutual attraction; D indicates the maximum possible distance between key feature subsets; represents the random noise term; The key feature set is screened according to the attraction formula to obtain the optimal feature.
[0069] In this embodiment, the degree of excellence refers to the effectiveness and importance of each key feature subset in the prediction model, which is usually evaluated based on the performance indicators of the model (such as accuracy, recall rate, F1-score, etc.). Feature subset generation: generating multiple feature subsets from the key feature set; model training and evaluation: using each feature subset to train the machine learning model and calculate the performance indicators of the model; excellence scoring: assigning a excellence score to each feature subset based on the model performance indicators.
[0070] In this embodiment, the fitness value is a numerical value used to quantify the quality of the feature subset, which is usually directly related to the model performance. The higher the fitness value, the more effective the feature subset is in the model.
[0071] In this embodiment, the fitness function is defined as follows: a fitness function is defined, for example, using the accuracy or other performance indicators of the model, a fitness value is calculated: a fitness value is calculated for each feature subset, typically by mapping the model performance indicator to a range of fitness values (such as 0 to 1), and a fitness value is recorded: the fitness value of each feature subset is recorded for subsequent comparison and selection.
[0072] In this embodiment, an attraction formula is defined based on factors such as the degree of mutual attraction between feature subsets, initial attraction, and distance attenuation rate. The attraction is calculated by using the formula: the attraction between each feature subset is calculated. Feature screening: feature screening is performed based on attraction and fitness value, and feature subsets with high attraction and good fitness value are selected as the optimal features.
[0073] The working principle and beneficial effects of the above technical solution are as follows: Based on the key feature set and the fire cicada algorithm, a fitness function is defined to evaluate the pros and cons of each key feature subset and calculate the fitness value. By setting the attraction formula, the mutual attraction between feature subsets, fitness changes and distance attenuation are considered, the attraction between feature subsets is dynamically adjusted, the feature set is screened according to the attraction, the optimal features are identified, the efficiency and accuracy of feature selection are effectively improved, and the performance of the data access quality prediction model is optimized.
[0074] Embodiment 7: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which sets an access quality peak threshold according to a peak period and sets an access quality valley threshold according to a valley period, including: Classify the collected data sources into types, and determine the standard threshold range corresponding to each type classification result from the type-data quality table according to the type classification results; Based on the standard threshold range, a peak access quality threshold is set in combination with a peak period, and a valley access quality threshold is set in combination with a valley period.
[0075] In this embodiment, type classification is to classify the collected data sources into different types according to specific standards (such as data source, data format, data purpose, etc.). Assuming that there are the following data sources: user behavior data, transaction data, sensor data, and log data, according to the characteristics of the data, they can be divided into: Type A: behavior data (user behavior data), Type B: transaction data (transaction data), Type C: sensor data (sensor data), Type D: log data (log data).
[0076] In this embodiment, the type-data quality table is a table that lists different types of data sources and their corresponding data quality standards and threshold ranges. The table is used to evaluate and manage the quality of different types of data.
[0077] In this embodiment, the type classification result is based on the type classification result to determine the type of each data source. For example, in the actual data collection process, a data source is identified as user behavior data, so its type classification result is "Type A".
[0078] In this embodiment, the standard threshold range is the quality standard and threshold range set for each data type, which is used to evaluate the quality of the data. For example, assuming that the standard threshold range for behavioral data is: completeness threshold: 95%, accuracy threshold: 90%, consistency threshold: 98%, timeliness threshold: 24 hours.
[0079] The working principle and beneficial effects of the above technical solution are: by classifying the collected data sources into types, determining the standard threshold range for each type from the type-data quality table based on the classification results, and combining the characteristics of peak and trough periods to set the peak threshold and trough threshold of access quality respectively, so as to facilitate subsequent quality evaluation and management, clarify the quality standards of different data types, improve the monitoring and evaluation capabilities of data access quality, and ensure the validity and stability of data in different time periods.
[0080] Embodiment 8: The embodiment of the present invention provides a data access quality intelligent analysis method based on periodic analysis, which inputs actual data into a quality discrimination model, performs quality abnormality judgment, and dynamically adjusts the data access strategy according to the abnormality judgment result, including: Input the actual data into the quality judgment model, and judge the abnormality of data access quality based on the output of the model; Record the abnormal judgment results, evaluate the data access strategy, and analyze the abnormal causes; Based on the abnormal cause and the abnormal judgment result, an adjustment plan is formulated to adjust the data access strategy.
[0081] In this embodiment, the abnormal cause is the specific reason that causes the abnormal data access quality, which may include technical problems, improper policy execution, data source problems, etc. For example, in the above example, the abnormal cause may be: insufficient data collection frequency (collected once every hour, but the data volume is large during peak hours and cannot be processed in time), unstable data source (sensor failure causes data loss).
[0082] In this embodiment, the adjustment plan is an adjustment measure formulated based on the abnormal cause and judgment result to optimize the data access strategy and improve data quality. For example, the frequency of data collection is increased (for example, data is collected every 30 minutes), a data quality monitoring tool is introduced, data access is monitored in real time, and the data source is upgraded or replaced to ensure data stability.
[0083] The working principle and beneficial effects of the above technical solution are: input the actual data into the quality judgment model, make abnormal judgments on the data access quality based on the model output results, record the abnormal judgment results, evaluate the data access strategy and analyze the abnormal causes, and formulate corresponding adjustment plans based on the abnormal causes and judgment results to optimize the data access strategy, ensure data quality, quickly identify anomalies and make targeted adjustments, improve the reliability and stability of data access, and enhance decision-making support capabilities.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data access quality intelligent analysis method based on periodic analysis, characterized in that: include: Step 1: Collect data from the data source to be collected, unify the data format, obtain unified data, apply statistical methods to identify the natural cycle and life cycle of the unified data, and then obtain the periodic pattern of the unified data; Step 2: Analyze the trend of data access quality of unified data in different periods based on the periodic pattern, extract key feature sets, and then identify the peak period and trough period of data access quality; Step 3: Based on the key feature set, a machine learning algorithm is selected to construct a prediction model for data access quality, and access quality thresholds are set according to the peak period and the trough period, thereby obtaining a quality discrimination model; Step 4: Input the actual data into the quality judgment model, make quality anomaly judgments, and dynamically adjust the data access strategy based on the anomaly judgment results.
2. According to claim 1, a data access quality intelligent analysis method based on periodic analysis is characterized in that: Collect data from the data sources to be collected, unify the data format, obtain unified data, apply statistical methods to identify the natural cycle and life cycle of the unified data, and then obtain the periodic pattern of the unified data, including: Obtaining user needs, determining multiple data sources to be collected according to the user needs, collecting data from the data sources to be collected, and unifying the data formats to obtain unified data; Applying statistical methods to analyze the unified data to obtain statistical analysis results, extracting trend components from the statistical analysis results to obtain the life cycle of the unified data, and extracting seasonal components from the statistical analysis results to obtain the natural cycle of the unified data; Combining the life cycle with the natural cycle results in a cyclical pattern of unified data.
3. The method for intelligent analysis of data access quality based on periodic analysis according to claim 1, characterized in that: Based on the periodic pattern, the trend of data access quality of unified data in different periods is analyzed, and key feature sets are extracted to identify the peak and trough periods of data access quality, including: Determining a decomposition method of the unified data in different periods from a data decomposition library according to the periodic pattern, and decomposing the unified data using the decomposition method to obtain a decomposition result; Draw a trend graph, a seasonal graph, and a residual graph of the unified data according to the decomposition results, and extract a key feature set of the unified data based on the trend graph, the seasonal graph, and the residual graph; All peak periods and valley periods of data access quality are identified based on the key feature set.
4. The method for intelligent analysis of data access quality based on periodic analysis according to claim 3 is characterized in that: According to the decomposition results, a trend graph, a seasonal graph, and a residual graph of the unified data are drawn, and a key feature set of the unified data is extracted based on the trend graph, the seasonal graph, and the residual graph, including: Observe the trend graph, determine the long-term change direction of the data access quality, calculate the slope of the trend line, obtain the trend strength, identify the trend change point from the trend graph, and obtain the first feature by combining the long-term change direction, trend strength and trend change point; Observe the seasonal graph to determine the change pattern of the unified data, calculate the autocorrelation coefficient of the unified data based on the seasonal graph, determine the lag period, calculate the seasonal index of the unified data in combination with the lag period and the change pattern, and then determine the second feature; Observe the residual graph, evaluate the amplitude and frequency of random fluctuations, identify outliers in the residuals, perform cycle and trend analysis on the outliers, and determine the third feature based on the cycle and trend analysis results; The first feature, the second feature and the third feature are combined to obtain a key feature set of unified data.
5. The method for intelligent analysis of data access quality based on periodic analysis according to claim 1, characterized in that: A machine learning algorithm is selected based on the key feature set to build a prediction model for data access quality, and access quality thresholds are set according to the peak period and the trough period, thereby obtaining a quality discrimination model, including: The Fire Cicada algorithm is used to screen the key feature set to obtain the optimal feature, and a machine learning algorithm is selected from the feature-algorithm library based on the optimal feature to build a prediction model for data access quality; The access quality peak threshold is set according to the peak period, and the access quality trough threshold is set according to the trough period. The quality discrimination model is obtained by combining the prediction model with the access quality peak threshold and the access quality trough threshold.
6. The method for intelligent analysis of data access quality based on periodic analysis according to claim 5, characterized in that: The Fire Cicada algorithm is used to screen the key feature set and obtain the optimal features, including: Based on the key feature set combined with the Fire Cicada algorithm, a fitness function is defined to evaluate the quality of each key feature subset, and the fitness value of each key feature subset is calculated, and the attractiveness formula of the key feature subset is set according to the fitness value and the Fire Cicada algorithm; , where X represents the mutual attraction between key feature subsets; represents the initial attraction between subsets of key features; Indicates the speed at which the mutual attraction decays with distance; represents the distance between key feature subset i and key feature subset j; T represents the total running time of the Fire Cicada algorithm; t represents the process of the Fire Cicada algorithm; represents the adjustment parameter of the change in fitness on the degree of mutual attraction; Indicates the difference between the current fitness and the previous fitness; Indicates the maximum fitness value in the current population; Indicates the degree to which the influence distance weakens the mutual attraction; D indicates the maximum possible distance between key feature subsets; represents the random noise term; The key feature set is screened according to the attraction formula to obtain the optimal feature.
7. The method for intelligent analysis of data access quality based on periodic analysis according to claim 2, characterized in that: Set the access quality peak threshold according to the peak period and the access quality valley threshold according to the valley period, including: Classify the collected data sources into types, and determine the standard threshold range corresponding to each type classification result from the type-data quality table according to the type classification results; Based on the standard threshold range, a peak access quality threshold is set in combination with a peak period, and a valley access quality threshold is set in combination with a valley period.
8. The method for intelligent analysis of data access quality based on periodic analysis according to claim 1, characterized in that: Input the actual data into the quality judgment model to judge the quality anomaly, and dynamically adjust the data access strategy according to the anomaly judgment results, including: Input the actual data into the quality judgment model, and judge the abnormality of data access quality based on the output of the model; Record the abnormal judgment results, evaluate the data access strategy, and analyze the abnormal causes; Based on the abnormal cause and the abnormal judgment result, an adjustment plan is formulated to adjust the data access strategy.
Citation Information
Patent Citations
Bearing fault diagnosis model construction method, diagnosis method and electronic equipment
CN110766100A
Default probability prediction method for optimizing recurrent neural network based on gravitational search algorithm
CN113379536A
Using method for integrating various networks and video conference equipment
CN118158089A
Cable channel structure settlement monitoring system
CN118260723A
Internet traffic consumption real-time reminding system
CN118264592A
Cited By
Analysis and display method and system for network quality detection, medium and product
CN121309384A