Data identification method and system based on artificial intelligence
Time series data is processed through adaptive denoising and coupling correction technology, and combined with Lyapunov index and information entropy analysis, SOM network and K-means clustering algorithm are constructed, which solves the problems of low accuracy of time series data recognition and neglected data correlation in the existing technology, and achieves more efficient and accurate data recognition.
Patent Information
- Application Number
- CN202510194288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing data identification technologies often face noise problems when processing timing data, resulting in reduced accuracy of identification results, and fail to fully consider the potential coupling relationship and interdependence between data, ignoring the inherent correlation and dynamic variability of data.
By collecting the data to be identified, the time series is constructed, the feedback coefficient is calculated to perform adaptive denoising on the time series, and coupling correction is performed based on the time series after adaptive denoising. The coupled corrected time series is converted into phase spatial data, the projection trajectory is constructed and the Lyapunov index is calculated, the critical network is constructed using weighted undirected graphs, the information entropy and dependency metrics are calculated, and the best matching unit for the feature vector is constructed to calculate the eigenvectors of the SOM network, and finally the event type is identified using the K-means clustering algorithm.
It improves the accuracy and stability of data identification, enhances the accuracy and robustness of classification, and reduces data redundancy and computational complexity.
Smart Images

Figure CN120123804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a data recognition method and system based on artificial intelligence. Background Art
[0002] With the rapid development of information technology and big data, data plays an increasingly important role in various fields. Especially in fields such as industry, finance, and healthcare, a large amount of data is continuously generated through sensors, the Internet, and various monitoring devices. How to efficiently and accurately process and recognize this data has become a hot issue in current research and applications. Data recognition technology plays a crucial role in these fields. Especially with the support of artificial intelligence (AI) and machine learning technologies, the data processing ability and accuracy are continuously improved. The data recognition method based on artificial intelligence has gradually become an important development trend, which can automatically learn from data, identify patterns, and has strong generalization ability.
[0003] The existing data recognition technologies still have some deficiencies. When dealing with time series data, the existing methods often face the problem of noise, resulting in a reduction in the accuracy of the recognition results. The existing methods often fail to fully consider the potential coupling relationships and interdependencies between data, resulting in the neglect of the internal relevance and dynamic variability of data during modeling. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a data recognition method and system based on artificial intelligence, which solves the problems that the existing methods often face the problem of noise when dealing with time series data, resulting in a reduction in the accuracy of the recognition results, and the existing methods often fail to fully consider the potential coupling relationships and interdependencies between data, resulting in the neglect of the internal relevance and dynamic variability of data during modeling.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a data recognition method based on artificial intelligence, which includes,
[0008] Collect the data to be recognized to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the adaptively denoised time series;
[0009] Convert the time series after coupling correction into phase space data, construct the projection trajectory, calculate the Lyapunov exponent of the projection trajectory, set the Lyapunov exponent method as the critical point, construct the critical network using the weighted undirected graph, calculate the information entropy and dependence measure respectively and merge them into the feature vector, and construct the SOM network to calculate the best matching unit of the feature vector;
[0010] Use the K-means clustering algorithm to cluster the best matching units of the feature vectors, identify the event types, and store the data to be identified generated by collection and analysis.
[0011] As a preferred solution of the data recognition method based on artificial intelligence described in the present invention, wherein: collecting the data to be identified to construct a time series, calculating the feedback coefficient to perform adaptive denoising on the time series, and performing coupling correction based on the time series after adaptive denoising, including:
[0012] Collect the data to be identified from the time series database through the APL interface;
[0013] Sort the collected data to be identified in chronological order to generate a time series;
[0014] The data to be identified includes temperature and humidity, air quality index, GDP, number of likes, heart rate, and trading volume data;
[0015] Perform standardization processing on the time series, use the interquartile range method to identify and delete abnormal data, and use the mode imputation method to fill in the missing data;
[0016] Use the mean filtering method to perform preliminary denoising on the standardized time series, and calculate the feedback coefficient K(x i );
[0017] Use the k-nearest neighbor method to calculate the distance neighborhood data point set N(i) between the time series after preliminary denoising;
[0018] Use the feedback coefficient K(x i ) to perform adaptive denoising on the time series after preliminary denoising, and obtain the i-th data point x i ′ ;
[0019] Calculate the coupling degree C(x i ,x j ) between the i-th and j-th data points;
[0020] Based on the coupling degree, calculate the strong coupling correction coefficient α and the weak coupling correction coefficient γ respectively;
[0021] The percentile method is used to set the classification threshold. The data points after adaptive denoising with a coupling degree greater than or equal to the classification threshold are subjected to strong coupling correction, and the data points after adaptive denoising with a coupling degree less than the classification threshold are subjected to weak coupling correction.
[0022] As a preferred embodiment of the data recognition method based on artificial intelligence according to the present invention, wherein: the converting the time series after coupling correction into phase space data and constructing a projection trajectory includes:
[0023] The asymptotic dominance criterion is used to set the delay time and embedding dimension, and the delay coordinate method is used to convert the time series after coupling correction into phase space data;
[0024] The interval probability distribution and joint probability distribution of the phase space data are calculated using statistical analysis methods, the single-dimensional entropy is calculated using the Shannon entropy formula, and the joint entropy is calculated using the joint entropy formula;
[0025] Based on the difference between the single-dimensional entropy and the joint entropy, the information gain is calculated;
[0026] The information gain selection method is used to select the maximum information gain as the section plane of the phase space;
[0027] The threshold method is used to set the intersection points of the Poincaré section, and the continuity check method is used to construct the projection trajectory.
[0028] As a preferred embodiment of the data recognition method based on artificial intelligence according to the present invention, wherein: calculating the Lyapunov exponent of the projection trajectory, using the Lyapunov exponent method to set it as the critical point, and constructing a critical network using a weighted undirected graph, includes:
[0029] Randomly select a pair of adjacent phase space data in the projection trajectory as the initial points;
[0030] The Euclidean distance formula is used to calculate the Euclidean distance between the initial points;
[0031] The second trajectory is generated by using the small perturbation method for the initial points, and the Euclidean distance between the projection trajectory and the second trajectory is calculated using the Euclidean distance formula;
[0032] The Lyapunov exponent calculation method is used to calculate the Lyapunov exponent at time t;
[0033] The Lyapunov exponent method is used to set it as the critical point, the mutation detection method is used to set the window size, and the phase space data within the time window before and after the critical point is extracted;
[0034] The phase space data within the time window is set as the nodes of the critical network, the dynamic time warping algorithm is used to calculate the similarity between two nodes, and the reciprocal of the similarity is used to calculate the connection weight between the nodes;
[0035] Construct a critical network using a weighted undirected graph.
[0036] As a preferred solution of the data recognition method based on artificial intelligence according to the present invention, wherein: the steps of respectively calculating the information entropy and the dependence metric and combining them into a feature vector, and constructing a SOM network to calculate the best matching unit of the feature vector include:
[0037] Summarize all the nodes of the critical network to obtain a node set, calculate the information flow transmission probability using the weight normalization method, and calculate the information entropy of the nodes using the information entropy formula;
[0038] Based on the information entropy of the nodes, calculate the dependence metric E between the u-th and v-th nodes u,v ;
[0039] Normalize the information entropy and the dependence metric of the nodes respectively, and combine them into a feature vector T;
[0040] Use principal component analysis to reduce the dimension of the feature vector T;
[0041] Collect the data to be trained from the UCI machine learning library, construct a historical critical network for the data to be trained, calculate the historical information entropy and the historical dependence metric respectively, and perform normalization processing to generate a historical feature vector. Use principal component analysis to reduce the dimension of the historical feature vector to generate a training set;
[0042] Construct a SOM network, including an input layer and an output layer;
[0043] Set the input layer as the feature vector after dimension reduction;
[0044] Randomly select the feature vector after dimension reduction from the training set as a sample, initialize the weight vector of the neuron using a random value, calculate the Euclidean distance between the sample and the weight vector of the neuron using the Euclidean distance formula, and select the neuron closest to the sample as the best matching unit b BMU ;
[0045] Set the initial neighborhood radius and learning rate using the empirical rule, and update the neighborhood radius and learning rate respectively using the linear attenuation formula;
[0046] Set the neighborhood function using the Gaussian function;
[0047] Update the weight vector of the neuron using the SOM weight update formula;
[0048] Set the maximum number of iterations based on the convergence criterion method, update the SOM network in each round of iteration, and stop the iteration when the maximum number of iterations is reached;
[0049] Input the feature vectors after dimensionality reduction into the SOM network to obtain the best matching unit of the feature vectors after dimensionality reduction.
[0050] As a preferred embodiment of the data recognition method based on artificial intelligence according to the present invention, wherein: use the K-means clustering algorithm to cluster the best matching units of the feature vectors to identify event types, including:
[0051] Use the ROC curve method to set the classification threshold, use the elbow method to set the number of clusters K, randomly select K load data initial cluster centers from the best matching units of the feature vectors after dimensionality reduction, use the Euclidean distance formula to calculate the Euclidean distance from the best matching units of the feature vectors after dimensionality reduction to the K distance centers, assign the best matching units of the feature vectors after dimensionality reduction to the nearest cluster center, and recalculate the K cluster centers after each assignment. Stop the iteration when the calculated cluster centers are less than the judgment threshold to obtain K classified event types.
[0052] As a preferred embodiment of the data recognition method based on artificial intelligence according to the present invention, wherein: store the data to be recognized generated by collection and analysis, including:
[0053] Store the collected data to be recognized and the event types generated by analysis in the central database, and set security access measures. The central database backs up the stored data to the cloud, and regularly performs integrity detection on the stored data and the backup data. After the detection is completed, generate an integrity detection record and synchronously store it in the central database.
[0054] In a second aspect, the present invention provides a data recognition system based on artificial intelligence, including,
[0055] A collection and correction module, configured to collect the data to be recognized to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the time series after adaptive denoising;
[0056] A calculation and construction module, configured to convert the time series after coupling correction into phase space data and construct a projection trajectory, calculate the Lyapunov exponent of the projection trajectory, use the Lyapunov exponent method to set it as a critical point, construct a critical network using a weighted undirected graph, calculate the information entropy and the dependence metric respectively and merge them into a feature vector, and construct a SOM network to calculate the best matching unit of the feature vector;
[0057] A clustering and storage module, configured to use the K-means clustering algorithm to cluster the best matching units of the feature vectors to identify event types, and store the data to be recognized generated by collection and analysis.
[0058] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the data recognition method based on artificial intelligence as described in the first aspect of the present invention is implemented.
[0059] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the data recognition method based on artificial intelligence as described in the first aspect of the present invention is implemented.
[0060] The beneficial effects of the present invention are as follows: The present invention constructs a time series by collecting data to be recognized, calculates a feedback coefficient to perform adaptive denoising on the time series, and performs coupling correction based on the adaptively denoised time series; converts the coupled and corrected time series into phase space data and constructs a projection trajectory, calculates the Lyapunov exponent of the projection trajectory, sets it as a critical point using the Lyapunov exponent method, constructs a critical network using a weighted undirected graph, calculates the information entropy and dependence measure respectively and combines them into a feature vector, constructs a SOM network to calculate the best matching unit of the feature vector; uses the K-means clustering algorithm to cluster the best matching units of the feature vectors, identifies the event type, improves the accuracy and stability of data recognition, enhances the accuracy and robustness of classification, and reduces data redundancy and computational complexity. Description of the Drawings
[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is a flowchart of the data recognition method based on artificial intelligence in Embodiment 1.
[0063] Figure 2 It is a structural diagram of the data recognition system based on artificial intelligence in Embodiment 1. Detailed Embodiments
[0064] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.
[0065] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0066] Secondly, as used herein, an "embodiment" or "embodiments" refers to specific features, structures, or characteristics that may be included in at least one implementation of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive of other embodiments.
[0067] Example 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a data recognition method based on artificial intelligence, including the following steps:
[0068] S1. Collect the data to be recognized to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the adaptively denoised time series;
[0069] Specifically, collecting the data to be recognized to construct a time series, calculating the feedback coefficient to perform adaptive denoising on the time series, and performing coupling correction based on the adaptively denoised time series includes:
[0070] Collect the data to be recognized through the APL interface from the time series database;
[0071] Sort the collected data to be recognized in chronological order to generate a time series;
[0072] The data to be recognized includes temperature and humidity, air quality index, GDP, number of likes, heart rate, and trading volume data;
[0073] Perform standardization processing on the time series, use the interquartile range method to identify and delete abnormal data, and use the mode imputation method to fill in missing data;
[0074] Use the mean filter method to perform preliminary denoising processing on the standardized time series, and calculate the feedback coefficient K(x i ) of the preliminarily denoised time series. The formula is:
[0075]
[0076] where x i is the data point of the i-th preliminarily denoised time series, μ x is the mean of the preliminarily denoised time series, and σ x is the standard deviation of the preliminarily denoised time series;
[0077] Calculate the set of distance neighborhood data points N(i) between the time series after preliminary denoising using the k-nearest neighbor method;
[0078] Use the feedback coefficient K(x i ) to perform adaptive denoising on the time series after preliminary denoising to obtain the i-th data point x i ′, and the formula is:
[0079] x i ′ = x i + K(x i )·(x i - μ x ),
[0080] Calculate the coupling degree C(x i , x j ) between the i-th and j-th data points, and the formula is:
[0081]
[0082] Calculate the strong coupling correction coefficient α and the weak coupling correction coefficient γ respectively based on the coupling degree, and the formula is:
[0083]
[0084] where C max is the maximum coupling degree in the time series after preliminary denoising;
[0085] Use the percentile method to set the classification threshold, perform strong coupling correction on the data points after adaptive denoising with a coupling degree greater than or equal to the classification threshold, and perform weak coupling correction on the data points after adaptive denoising with a coupling degree less than the classification threshold. The formula is:
[0086]
[0087] where x i strong is the i-th strongly coupled corrected data point, x i weak is the i-th weakly coupled corrected data point, and x j ′ is the j-th data point after adaptive denoising.
[0088] By using the data collected from data sources (such as temperature and humidity, air quality index, GDP, etc.), the state of the system can be comprehensively described, multi-dimensional and fine-grained information can be obtained, which helps to perform more accurate data analysis and identification. The data values are converted to a unified scale using a standardized method to eliminate the dimensional differences between different data sources. The interquartile range method is used to identify and remove abnormal data, which can effectively remove extreme values that do not conform to statistical laws and ensure the accuracy and stability of the data. The mode imputation method is used to fill in the missing data, effectively solving the problem of missing data and making the data more complete, providing a guarantee for further processing and analysis. The calculated feedback coefficient helps to reflect the degree of change in the denoised data and provides a feedback signal for subsequent adaptive denoising. By calculating the distance neighborhood data point set between data points using the k-nearest neighbor method, the correlation relationships in the data can be better identified, thereby improving the denoising effect. Based on the strong and weak coupling correction of the coupling degree, it is ensured that different data points are appropriately corrected according to their similarity, which can not only optimize the data processing effect but also make the correction process more refined and flexible, and can adapt to complex data in different scenarios.
[0089] S2. Convert the coupled-corrected time series into phase space data and construct a projection trajectory, calculate the Lyapunov exponent of the projection trajectory, set it as the critical point using the Lyapunov exponent method, construct a critical network using a weighted undirected graph, calculate the information entropy and dependence measure respectively and merge them into a feature vector, and construct a SOM network to calculate the best matching unit of the feature vector;
[0090] Specifically, converting the coupled-corrected time series into phase space data and constructing a projection trajectory includes:
[0091] Set the delay time and embedding dimension using the asymptotic dominance criterion, and convert the coupled-corrected time series into phase space data using the delay coordinate method;
[0092] Use statistical analysis methods to calculate the interval probability distribution and joint probability distribution of the phase space data respectively, calculate the single-dimensional entropy using the Shannon entropy formula, and calculate the joint entropy using the joint entropy formula;
[0093] Based on the difference between the single-dimensional entropy and the joint entropy, calculate the information gain, and the formula is:
[0094] IG(X i ,H(X j ))=H(X i )+H(X j )-H(X i ,H(X j )),
[0095] Use the information gain selection method to select the maximum information gain as the section plane of the phase space;
[0096] Set the intersection points of the Poincaré section using the threshold method, construct the projection trajectory using the continuity check method, record the phase space data of the intersection points of each crossing of the Poincaré section, and connect the intersection points in chronological order to form the projection trajectory.
[0097] The asymptotic dominance criterion compares different combinations of delay times and dimensions, selects the parameters that can best reflect the internal dynamic law of the data, thus providing support for the accurate modeling of the data. Transforming time series data into phase space data through the delay coordinate method helps to reveal the nonlinearity and time series dependence of the data. Shannon entropy provides the degree of chaos of a single data dimension, while joint entropy reveals the correlation between multiple data dimensions. The information gain selection method improves the efficiency and accuracy of data analysis. By setting the intersection points of the Poincaré section, complex dynamic trajectories can be transformed into simplified two-dimensional intersection points, which is convenient for further analyzing the periodicity and dynamic behavior of the data. Combining with the continuity check method ensures that the constructed phase space data trajectory has physical rationality and does not show incoherent changes.
[0098] Furthermore, calculate the Lyapunov exponent of the projection trajectory, set it as the critical point using the Lyapunov exponent method, and construct a critical network using a weighted undirected graph, including:
[0099] Randomly select a pair of adjacent phase space data in the projection trajectory as the initial points;
[0100] Calculate the Euclidean distance between the initial points using the Euclidean distance formula;
[0101] Generate a second trajectory for the initial points using the small perturbation method, and calculate the Euclidean distance between the projection trajectory and the second trajectory using the Euclidean distance formula;
[0102] Calculate the Lyapunov exponent at time t using the Lyapunov exponent calculation method, and the formula is:
[0103]
[0104] where is, d 0 is the Euclidean distance between the initial points, and d(t) is the Euclidean distance between the projection trajectory at time t and the second trajectory;
[0105] Set it as the critical point using the Lyapunov exponent method, set the window size using the mutation detection method, and extract the phase space data within the time window before and after the critical point;
[0106] Set the phase space data within the time window as the nodes of the critical network, calculate the similarity between two nodes using the dynamic time warping algorithm, and calculate the connection weight between nodes using the reciprocal of the similarity. The formula is as follows:
[0107]
[0108] where ω uv is the connection weight between the u-th and v-th nodes, X u ′ is the phase space data within the time window before and after the u-th critical point, and X v ′ is the phase space data within the time window before and after the v-th critical point;
[0109] Construct a critical network using a weighted undirected graph.
[0110] By comparing the distances between adjacent data points, the distribution of data in the phase space can be effectively evaluated. By making small perturbations to the initial point, new trajectories are generated to explore the sensitivity of the system to initial conditions. Through the calculation of Lyapunov exponents, the dynamic behavior of the system can be accurately revealed, and key critical points are selected for in-depth analysis based on the calculation results. By analyzing the data before and after the critical points, the critical moments when the system undergoes mutations can be identified. The use of the DTW algorithm can not only solve the out-of-sync problem in time series but also reveal the potential correlations in the data. The construction of a weighted undirected graph can clearly reflect the dependence relationships between different states in the system, helping to identify the dynamic structure and interrelationships of the system.
[0111] Furthermore, calculate the information entropy and dependence measure respectively and merge them into a feature vector, and construct a SOM network to calculate the best matching unit of the feature vector, including:
[0112] Summarize all the nodes of the critical network to obtain a node set, calculate the information flow transmission probability using the weight normalization method, and calculate the information entropy of the nodes using the information entropy formula;
[0113] Based on the information entropy of the nodes, calculate the dependence measure E u,v between the u-th and v-th nodes, and the formula is:
[0114]
[0115] where A(u) is the information entropy of the u-th node, and A(v) is the information entropy of the v-th node;
[0116] Normalize the information entropy and dependence measure of the nodes respectively, and merge them into a feature vector T, T = [A(u) ′ , E u,v ′ , where A(u) ′is the information entropy of the u-th node after normalization, E u,v ′ is the dependency measure between the u-th and v-th nodes after normalization;
[0117] Use principal component analysis to reduce the dimensionality of the feature vector T;
[0118] Collect the data to be trained from the UCI Machine Learning Repository, construct a historical critical network for the data to be trained, calculate the historical information entropy and historical dependency measure respectively, and perform normalization processing to generate a historical feature vector. Use principal component analysis to reduce the dimensionality of the historical feature vector to generate a training set;
[0119] Construct a SOM network, including an input layer and an output layer;
[0120] Set the input layer as the feature vector after dimensionality reduction;
[0121] Randomly select the feature vector after dimensionality reduction from the training set as a sample, initialize the weight vector of the neuron with a random value, calculate the Euclidean distance between the sample and the weight vector of the neuron using the Euclidean distance formula, and select the neuron closest to the sample as the best matching unit b BMU ;
[0122] Set the initial neighborhood radius and learning rate using the empirical rule, and update the neighborhood radius and learning rate respectively using the linear decay formula;
[0123] Set the neighborhood function using the Gaussian function, and the formula is:
[0124]
[0125] where is the output of the neighborhood function between the b-th neuron and the best matching unit b at the p-th iteration, and σ(t) is the neighborhood radius at the p-th iteration, which gradually decreases with time; BMU between;
[0126] Update the weight vector of the neuron using the SOM weight update formula, and the formula is:
[0127]
[0128] where W b (p + 1) is the weight vector of the b-th neuron at the (p + 1)-th iteration, and W b (p) is the weight vector of the b-th neuron at the p-th iteration, η(p) is the learning rate at the p-th iteration, and Y is the sample;
[0129] Set the maximum number of iterations based on the convergence criterion method, update the SOM network in each round of iteration, and stop the iteration when the maximum number of iterations is reached;
[0130] Input the feature vector after dimensionality reduction into the SOM network to obtain the best matching unit of the feature vector after dimensionality reduction.
[0131] By using the weight normalization method to calculate the information flow transmission probability and calculating the information entropy of each node according to the information entropy formula, the complexity and uncertainty of each node can be quantified. The level of the dependence metric reflects the strength of the information transmission or interaction between two nodes, and thus the nodes or regions that play a key role in the system can be identified. After normalizing the information entropy and dependence metric of the nodes respectively and merging them into a feature vector, different data features can be integrated into a unified expression, which is convenient for subsequent processing and analysis. When dealing with high-dimensional data, PCA not only improves the efficiency of the algorithm but also helps reduce the interference of noise and improve the robustness of the data model. Introducing historical data not only provides sufficient training samples for the model but also captures the past behavior patterns of the system, providing strong support for future predictions. Under the framework of unsupervised learning, the SOM network can adaptively identify the structure and patterns of input data, and is particularly suitable for the classification and clustering analysis of complex data. By comparing the distance between the sample and the neuron, the neuron that best represents the characteristics of the input data is selected to ensure that the network can effectively classify and map the input data. The Gaussian function is used to set the neighborhood function, and the weight vector of the neuron is adjusted through the SOM weight update formula, ensuring that the SOM network gradually improves during the training process, can better adapt to the input data, and optimizes the weight distribution of the network through iteration, thereby improving the accuracy of classification and mapping.
[0132] S3. Use the K-means clustering algorithm to cluster the best matching units of the feature vector, identify the event types, and store the data to be recognized generated by collection and analysis;
[0133] Specifically, use the K-means clustering algorithm to cluster the best matching units of the feature vector and identify the event types, including:
[0134] Use the ROC curve method to set the classification threshold, use the elbow method to set the number of clusters K, randomly select K load data initial cluster centers from the best matching units of the feature vector after dimensionality reduction, use the Euclidean distance formula to calculate the Euclidean distance from the best matching units of the feature vector after dimensionality reduction to the K distance centers, and assign the best matching units of the feature vector after dimensionality reduction to the nearest cluster center. After each assignment, recalculate the K cluster centers, and stop the iteration when the calculated cluster center is less than the judgment threshold to obtain the K event types after classification.
[0135] By comparing the relationship between the false positive rate and the true positive rate, the ROC curve helps to select the best decision boundary and balance sensitivity and specificity. Selecting the appropriate K value can ensure that each cluster is sufficiently representative and improve the stability and interpretability of the clustering results. By calculating the Euclidean distance between each sample and the center, the event type can be quickly and effectively divided. The K-means algorithm can adaptively adjust the cluster center to ensure that each cluster can better represent the data points it contains. Setting a reasonable stopping criterion can effectively reduce the number of iterations, thereby improving the efficiency of the algorithm. Automatically assigning data points to different clusters can effectively distinguish different types of events and reveal potential patterns in the data. The K-means clustering algorithm is efficient and scalable, and can quickly generate clustering results for further analysis.
[0136] Furthermore, the data to be identified generated by storage, collection and analysis includes:
[0137] The collected data to be identified and the event types generated by the analysis are stored in the central database, and security access measures are set. The central database will back up the stored data to the cloud and regularly perform integrity checks on the stored data and backup data. After the test is completed, the integrity test record is generated and stored synchronously in the central database.
[0138] The central database can efficiently organize and retrieve large amounts of data, providing strong support for subsequent data analysis, auditing and decision-making. Security access measures include data encryption, enhanced identity authentication, role permission management, etc., which can ensure that only users who comply with security policies can access and operate sensitive data. Cloud backup can not only ensure high availability and disaster recovery of data, but also improve system reliability through redundant storage and distributed computing, generate integrity detection records and store them synchronously in the central database, which not only provides a basis for data recovery and auditing, but also provides managers with real-time feedback on the health status of the data, helping to promptly discover potential storage problems, and restore data through cloud backup to reduce system downtime caused by data loss or damage. By combining security access measures and cloud backup, the system can achieve intelligent management and efficient access to data.
[0139] This embodiment also provides a data recognition system based on artificial intelligence, including:
[0140] The collection and correction module is used to collect the data to be identified to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the time series after adaptive denoising;
[0141] A computational construction module is used to convert the coupled corrected time series into phase space data and construct a projection trajectory, calculate the Lyapunov index of the projection trajectory, set it as a critical point using the Lyapunov index method, construct a critical network using a weighted undirected graph, calculate the information entropy and dependency measure separately and merge them into a feature vector, and construct a SOM network to calculate the best matching unit of the feature vector;
[0142] The clustering storage module is used to cluster the best matching units of the feature vector using the K-means clustering algorithm, identify the event type, and store the data to be identified generated by collection and analysis.
[0143] This embodiment also provides a computer device, which is suitable for the case of an artificial intelligence-based data recognition method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the artificial intelligence-based data recognition method proposed in the above embodiment.
[0144] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0145] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the data recognition method based on artificial intelligence proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0146] In summary, the present invention constructs a time series by collecting data to be identified, calculates a feedback coefficient to adaptively denoise the time series, and performs coupling correction based on the time series after adaptive denoising; converts the time series after coupling correction into phase space data and constructs a projection trajectory, calculates the Lyapunov exponent of the projection trajectory, sets it as a critical point using the Lyapunov exponent method, constructs a critical network using a weighted undirected graph, calculates information entropy and dependency metrics respectively and merges them into a feature vector, constructs a SOM network to calculate the best matching unit of the feature vector; uses a K-means clustering algorithm to cluster the best matching units of the feature vector, identifies the event type, improves the accuracy and stability of data recognition, enhances the accuracy and robustness of classification, and reduces data redundancy and computational complexity.
[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A data recognition method based on artificial intelligence, characterized in that: include, Collect the data to be identified to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the time series after adaptive denoising; The time series after coupling correction is converted into phase space data and the projection trajectory is constructed. The Lyapunov index of the projection trajectory is calculated and set as the critical point using the Lyapunov index method. The critical network is constructed using a weighted undirected graph. The information entropy and dependency metric are calculated and merged into a feature vector. The SOM network is constructed to calculate the best matching unit of the feature vector. The K-means clustering algorithm is used to cluster the best matching units of the feature vector, identify the event type, and store the data to be identified generated by collection and analysis.
2. The artificial intelligence-based data recognition method according to claim 1, characterized in that: The collecting of the data to be identified to construct a time series, calculating the feedback coefficient to adaptively denoise the time series, and performing coupling correction based on the adaptively denoised time series include: Collect the data to be identified from the time series database through the APL interface; Sort the collected data to be identified in chronological order to generate a time series; The data to be identified include temperature and humidity, air quality index, GDP, number of likes, heart rate, and transaction volume data; The time series were standardized, the interquartile range method was used to identify and delete abnormal data, and the mode interpolation method was used to fill in missing data; The mean filter method is used to perform preliminary denoising on the standardized time series, and the feedback coefficient K(x i ); Use the k-nearest neighbor method to calculate the distance neighborhood data point set N(i) between the time series after preliminary denoising; Using the feedback factor K(x i ) Adaptively denoise the time series after preliminary denoising to obtain the i-th adaptive denoised data point x i ′; Calculate the coupling degree C(x) between the i-th and j-th data points i , x j ); Based on the coupling degree, the strong coupling correction coefficient α and the weak coupling correction coefficient γ are calculated respectively; The classification threshold is set using the percentile method. The adaptive denoising data points with a coupling degree greater than or equal to the classification threshold are subjected to strong coupling correction, and the adaptive denoising data points with a coupling degree less than the classification threshold are subjected to weak coupling correction.
3. The artificial intelligence-based data recognition method according to claim 2, characterized in that: The step of converting the coupled corrected time series into phase space data and constructing a projection trajectory includes: The delay time and embedding dimension are set using the asymptotic dominance criterion, and the coupled corrected time series are converted into phase space data using the delayed coordinate method. The interval probability distribution and joint probability distribution of phase space data are calculated using statistical analysis methods, the single-dimensional entropy is calculated using the Shannon entropy formula, and the joint entropy is calculated using the joint entropy formula; Calculate information gain based on the difference between single-dimensional entropy and joint entropy; Use the information gain selection method to select the maximum information gain as the slice of the phase space; The intersection points of the Poincaré sections were set using a threshold method, and the projection trajectory was constructed using a continuity check method.
4. The artificial intelligence-based data recognition method according to claim 3, characterized in that: The Lyapunov exponent of the calculated projection trajectory is set as a critical point using the Lyapunov exponent method, and a critical network is constructed using a weighted undirected graph, including: Randomly select a pair of adjacent phase space data in the projection trajectory as the initial point; Use the Euclidean distance formula to calculate the Euclidean distance between the initial points; The second trajectory is generated by using a small perturbation method for the initial point, and the Euclidean distance formula is used to calculate the Euclidean distance between the projected trajectory and the second trajectory; Calculate the Lyapunov exponent at time t using the Lyapunov exponent calculation method; The critical point is set using the Lyapunov exponent method, the window size is set using the mutation detection method, and the phase space data in the time window before and after the critical point is extracted; The phase space data in the time window is set as the nodes of the critical network, the dynamic time warping algorithm is used to calculate the similarity between two nodes, and the connection weight between the nodes is calculated using the inverse of the similarity; Constructing critical networks using weighted undirected graphs.
5. The artificial intelligence-based data recognition method according to claim 4, characterized in that: The information entropy and dependency measure are calculated separately and combined into a feature vector, and a SOM network is constructed to calculate the best matching unit of the feature vector, including: All the nodes of the critical network are aggregated to obtain a node set, the information flow transmission probability is calculated using the weight normalization method, and the information entropy of the node is calculated using the information entropy formula; Based on the information entropy of the nodes, calculate the dependency measure E between the uth and vth nodes u,v ; Normalize the information entropy and dependency measure of the node separately and merge them into the feature vector T; Use principal component analysis to reduce the dimension of the feature vector T; Collect the training data from the UCI machine learning library, build a historical critical network for the training data, calculate the historical information entropy and historical dependency measure, and perform normalization to generate historical feature vectors. Use principal component analysis to reduce the dimension of the historical feature vectors and generate a training set. Construct a SOM network, including input layer and output layer; Set the input layer to the feature vector after dimensionality reduction; Randomly select the reduced eigenvector from the training set as the sample, use the random value to initialize the neuron weight vector, use the Euclidean distance formula to calculate the Euclidean distance between the sample and the neuron weight vector, and select the neuron closest to the sample as the best matching unit b. BMU ; Use empirical rules to set the initial neighborhood radius and learning rate, and use linear decay formulas to update the neighborhood radius and learning rate respectively; Use Gaussian function to set the neighborhood function; Update the neuron's weight vector using the SOM weight update formula; The maximum number of iterations is set based on the convergence criterion method, the SOM network is updated in each round of iteration, and the iteration is stopped when the maximum number of iterations is reached; The feature vector after dimension reduction is input into the SOM network to obtain the best matching unit of the feature vector after dimension reduction.
6. The artificial intelligence-based data recognition method according to claim 5, characterized in that: The K-means clustering algorithm is used to cluster the best matching units of the feature vector to identify the event type, including: The ROC curve method is used to set the classification threshold, the elbow rule is used to set the number of clusters K, K initial cluster centers of load data are randomly selected from the best matching unit of the feature vector after dimensionality reduction, and the Euclidean distance formula is used to calculate the Euclidean distance from the best matching unit of the feature vector after dimensionality reduction to the K distance centers. The best matching unit of the feature vector after dimensionality reduction is assigned to the cluster center with the nearest distance. After each assignment, the K cluster centers are recalculated. When the calculated cluster center is less than the judgment threshold, the iteration is stopped to obtain the classified K event types.
7. The artificial intelligence-based data recognition method according to claim 6, characterized in that: The data to be identified generated by the storage, collection and analysis includes: The collected data to be identified and the event types generated by the analysis are stored in the central database, and security access measures are set. The central database will back up the stored data to the cloud and regularly perform integrity checks on the stored data and backup data. After the test is completed, the integrity test record is generated and stored synchronously in the central database.
8. An artificial intelligence-based data recognition system, based on the artificial intelligence-based data recognition method according to any one of claims 1 to 7, characterized in that: include, The collection and correction module is used to collect the data to be identified to construct a time series, calculate the feedback coefficient to perform adaptive denoising on the time series, and perform coupling correction based on the time series after adaptive denoising; A computational construction module is used to convert the coupled corrected time series into phase space data and construct a projection trajectory, calculate the Lyapunov index of the projection trajectory, set it as a critical point using the Lyapunov index method, construct a critical network using a weighted undirected graph, calculate the information entropy and dependency measure separately and merge them into a feature vector, and construct a SOM network to calculate the best matching unit of the feature vector; The clustering storage module is used to cluster the best matching units of the feature vector using the K-means clustering algorithm, identify the event type, and store the data to be identified generated by collection and analysis.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the artificial intelligence-based data recognition method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the artificial intelligence-based data recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Data processing method and device for human-computer interaction and electronic equipment
CN121680707A