Personalized customization data mining and analysis system based on artificial intelligence

Through a personalized customized data mining and analysis system based on artificial intelligence, the management and analysis problems of multi-source heterogeneous and dynamically changing edge data are solved, efficient anomaly detection and personalized response are achieved, and the data availability and adaptability of the system are improved.

CN120632708AInactive Publication Date: 2025-09-12SHAANXI SHOUYI NETWORK TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510617968.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data mining technologies are difficult to effectively manage and analyze multi-source heterogeneous and dynamically changing edge data, lack personalized response capabilities, and the anomaly detection system lacks generalization capabilities in multiple business scenarios.

Method used

It adopts a personalized customized data mining and analysis system based on artificial intelligence, including edge data acquisition, heterogeneous tensor decomposition, dynamic clustering analysis and abnormal data labeling modules. Through multi-dimensional adaptive attention-enhanced heterogeneous tensor decomposition and cluster evolution graph model, it realizes dynamic analysis and anomaly detection of high-order data.

Benefits of technology

It improves data availability and maintainability, enhances the recognition and response efficiency of abnormal information, and enhances the system's adaptability and personalized analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632708A_ABST
    Figure CN120632708A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized customization data mining and analysis system based on artificial intelligence, and the system comprises the following modules: an edge data collection module which is used for generating a structured multi-source heterogeneous data set; the high-order data tensor construction module is used for constructing a high-order data tensor based on the structured multi-source heterogeneous data set; the heterogeneous tensor decomposition module is used for extracting a potential feature matrix and a core tensor; the feature fusion and unified expression module is used for organizing all the fused feature vectors according to a time sequence to form a unified feature expression sequence; the dynamic clustering analysis module is used for generating a clustering evolution diagram; the clustering stability evaluation and feedback module is used for obtaining a fusion feature vector in a cluster structure mutation state; and the abnormal data labeling module outputs an abnormal data labeling result set. According to the method, the overall data availability and maintainability of the system are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a personalized customized data mining and analysis system based on artificial intelligence. Background Art

[0002] With the continuous evolution of edge computing and cloud computing collaboration, massive amounts of real-time data are rapidly generated by various terminals and edge nodes. Data types include sensor signals, user behavior logs, network communication traffic, and other forms, creating a highly heterogeneous, multi-dimensional and dynamic data environment. In complex systems such as intelligent transportation, smart manufacturing, and industrial monitoring, how to effectively manage, accurately mine, and intelligently analyze edge data from different sources and types has become a key research direction for big data intelligent applications.

[0003] Currently, mainstream data mining technologies are primarily based on unified modeling paradigms, focusing on global optimization analysis in static data environments. These techniques, such as those based on traditional cluster analysis, statistical anomaly detection, or deep learning-based pattern recognition models, offer some versatility in extracting preliminary information, but they suffer from significant shortcomings in scenarios involving dynamically changing, multi-source, and heterogeneous data.

[0004] First, existing methods struggle to support the fusion modeling of high-dimensional, heterogeneous data. Faced with multi-type, cross-domain data from edge nodes, traditional low-level feature representation and linear processing methods are unable to preserve the interactive structure of multidimensional information, making it difficult to mine deep patterns and causing significant information loss. Furthermore, the lack of flexible abstraction mechanisms tailored to user customization requirements results in insufficient generalization capabilities in practical deployments.

[0005] Secondly, current cluster analysis generally uses static modeling, ignoring data evolution trends and temporal context. This makes it difficult to adapt to the time-varying and behaviorally volatile nature of data streams in edge environments, leading to delayed or even distorted clustering results. In personalized application scenarios, users expect models to track and respond to their business changes in real time, a requirement that existing technologies clearly struggle to meet.

[0006] Thirdly, most existing anomaly detection systems still remain at the level of unified threshold judgment or coarse-grained scoring, failing to combine cluster evolution history and behavioral trends to intelligently reason and respond to abnormal events. In a multi-business, multi-role parallel data platform, the system cannot generate personalized anomaly analysis reports based on different user needs, limiting its usability in tasks such as adaptation, security monitoring, and intelligent recommendation.

[0007] In summary, there is an urgent need to design a personalized data mining method to make up for the above technical deficiencies. Summary of the Invention

[0008] One purpose of the present invention is to propose a personalized customized data mining and analysis system based on artificial intelligence, which greatly improves the overall data availability and maintainability of the system.

[0009] According to an embodiment of the present invention, a personalized customized data mining and analysis system based on artificial intelligence includes the following modules:

[0010] The edge data acquisition module is used to collect and preprocess multi-source heterogeneous raw data from multiple cloud computing edge nodes to generate structured multi-source heterogeneous data sets;

[0011] High-order data tensor construction module, used to construct high-order data tensors based on structured multi-source heterogeneous data sets;

[0012] The heterogeneous tensor decomposition module introduces the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decomposes the high-order data tensor using heterogeneous tensor decomposition technology, and extracts the potential feature matrix and core tensor;

[0013] The feature fusion and unified expression module is used to perform structured concatenation and joint encoding of the extracted latent feature matrix and the core tensor to construct a fused feature vector, and organize all fused feature vectors in chronological order to form a unified feature representation sequence;

[0014] Dynamic clustering analysis module, used to perform dynamic clustering analysis on the unified feature representation sequence in a continuous time window, construct a cluster set, and generate a cluster evolution graph based on the similarity relationship between clusters in adjacent time windows;

[0015] The cluster stability evaluation and feedback module is used to calculate the cluster stability score and stability change rate of each time window based on the cluster evolution graph, mark the cluster structure mutation state, and obtain the fusion feature vector under the cluster structure mutation state;

[0016] The abnormal data annotation module is used to classify and annotate the abnormal behavior of the fused feature vector under the state of cluster structure mutation, and output the abnormal data annotation result set.

[0017] A personalized customization data mining and analysis method based on artificial intelligence, used to execute the personalized customization data mining and analysis system based on artificial intelligence, characterized by comprising the following steps:

[0018] S1. Collect multi-source heterogeneous datasets from multiple cloud computing edge nodes, perform data preprocessing on the multi-source heterogeneous datasets, and generate structured multi-source heterogeneous datasets;

[0019] S2. Construct high-order data tensors based on structured multi-source heterogeneous datasets;

[0020] S3. Introduce the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decompose the high-order data tensor using heterogeneous tensor decomposition technology, and extract the latent feature matrix and core tensor;

[0021] S4. Perform joint feature fusion processing on the extracted latent feature matrix and core tensor to construct a fused feature vector. The fused feature vector is then organized in chronological order to form a unified feature representation sequence. The unified feature representation sequence contains both local structural information and global interaction information of abnormal data at the cloud computing edge node.

[0022] S5. Input the unified feature representation sequence into the dynamic clustering module based on the cluster evolution graph, set a set of continuous time windows, perform dynamic clustering to generate a set of cluster clusters in each time window, and construct a cluster evolution graph. Calculate the cluster stability score and stability change rate. When the cluster stability score is lower than the abnormal warning threshold and the stability change rate is lower than the stability decline threshold, the system triggers the abnormal clustering feedback mechanism and marks the time window as a cluster structure mutation state. The fused feature vector in the cluster structure mutation state is classified as abnormal behavior and the abnormal data annotation result set is output.

[0023] Optionally, the S1 includes the following steps:

[0024] S11. Set the edge node set N in the cloud computing environment, n in the edge node set N k represents the kth edge node, from each edge node n k Collect multi-source heterogeneous raw data, including sensor data, log data, business traffic data and user behavior data, and build an initial multi-source heterogeneous dataset D raw ;

[0025] S12. Set the sampling time window T, the sampling time window T t m Indicates the mth sampling time. At each sampling time t m , edge node n k Collect raw data of corresponding types The original data points in the form of triples in, Indicates that the s-th data source is at edge node n k , sampling time t m The original data points at is the corresponding structured eigenvalue;

[0026] S13. Initial multi-source heterogeneous dataset D raw Perform unified preprocessing operations on all types of data in;

[0027] S14. All kinds of pre-processed data are sorted by edge node n k , sampling time t m , the triple dimensions of data source type s are recombined to generate a structured multi-source heterogeneous dataset:

[0028]

[0029] Among them, D struct is a structured multi-source heterogeneous dataset, K is the total number of edge nodes, M is the total number of sampling times, and S represents the number of data source categories.

[0030] Optionally, the preprocessing operation includes:

[0031] Time alignment processing interpolates or truncates different types of data based on a unified sampling time window T, so that all data are at a unified sampling time t m The above has a consistent alignment structure;

[0032] Data cleaning and processing to remove missing, duplicate or abnormally formatted raw data points and fill in missing values;

[0033] Format unification processing, converting unstructured data into structured form, so that all data can be represented as standard numerical feature vectors;

[0034] Noise removal processing, structural eigenvalues ​​of time series The filtering algorithm is used for smoothing to remove high-frequency noise interference.

[0035] Optionally, the S2 includes the following steps:

[0036] S21. In structured multi-source heterogeneous dataset D struct On this basis, a unified three-dimensional index system is established, which includes the edge node dimension N, the sampling time dimension T and the number of data source categories S;

[0037] S22. Based on the triple-dimensional index system, edge node n k , sampling time t m And the data source type s is a joint index, which extracts the edge node n from the structured multi-source heterogeneous data set k , sampling time t m The structured feature value corresponding to the s-th data source type Used to represent observation data of a single tensor unit;

[0038] S23. Arrange all the structured eigenvalues ​​extracted based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S in three dimensions to construct a third-order tensor structure and generate a high-order data tensor X. The third-order structure of the high-order data tensor X corresponds to the edge node dimension, the sampling time dimension, and the number of data source categories, respectively. The value of each tensor unit in the high-order data tensor is the structured eigenvalue.

[0039] Optionally, S3 includes the following steps:

[0040] S31. A high-order data tensor X is introduced into the heterogeneous tensor decomposition module of the multi-dimensional adaptive attention enhancement mechanism. By calculating the adaptive attention weight matrix based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S, the module automatically identifies the data dimensions and features that contribute most to abnormal events when the cloud computing edge node is in an abnormal state.

[0041] S32. The probability of abnormal events at edge nodes, the magnitude of abnormal data fluctuations, and the historical abnormal correlation strength are integrated into a unified multi-dimensional abnormal attention index. The attention weight matrix of the edge node dimension is set as The attention weight matrix of the sampling time dimension is The attention weight matrix of the number of data source categories is Among them, the edge node attention weight element Represents edge node n k The comprehensive abnormal attention degree is a combination of the cumulative frequency of historical abnormal events and the fluctuation amplitude of the current node data, and the sampling time attention weight element represents the sampling time t m The intensity of data fluctuations and the combined abnormal attention of abnormal trend prediction values, data source type attention weight element The abnormal attention degree is represented by the contribution of data source type s to the detection of historical abnormal events and the difference between the current node data. The Hadamard element-by-element product operation of the three-dimensional tensor is used to calculate the attention weight matrix of the edge node dimension, the attention weight matrix of the time dimension, and the attention weight matrix of the number of data source categories to obtain the adaptive attention weight matrix;

[0042] S33. Use the adaptive attention weight matrix to perform element-by-element multi-dimensional adaptive weighted adjustment on the high-order data tensor X to form an abnormal sensitivity enhanced weighted high-order data tensor X * :

[0043]

[0044] in, It is a weighted high-order data tensor element after abnormal sensitivity enhancement, reflecting the weight amplification effect of data points with higher abnormal contribution in the abnormal event mining task of cloud computing edge nodes;

[0045] S34. Using Tucker decomposition structure with non-negative constraints to weight high-order data tensor X * Implement a heterogeneous tensor decomposition process based on the joint optimization of non-negativity constraints and anomaly-sensitive loss functions to obtain the latent feature matrix and core tensor:

[0046] X * ≈G×1U N ×2U T ×3U S ;

[0047] The decomposition process is achieved by defining an abnormality-sensitive loss function L anom Optimization solution:

[0048]

[0049] Among them, G is the core tensor that captures the interaction relationship of multi-dimensional abnormal data, and the potential feature matrix U N 、U T 、U S They respectively reflect the abnormal characteristics in three dimensions: edge nodes, sampling time, and data source type. λ is the sparse regularization coefficient used to control the sensitivity to abnormalities.

[0050] Optionally, the S4 includes the following steps:

[0051] S41. Potential feature matrix U N 、U T 、U S Perform joint feature fusion processing with the core tensor G, structurally splice the tensor-level global interaction features with the low-rank features of each dimension to form a ternary joint feature representation set:

[0052] F k,m,s =Concat([U N (k,:)],[U T (m,:)],[U S (s,:)],[G :: ]);

[0053] Among them, F k,m,s Indicates that at edge node n k , sampling time t m , fusion feature vector under the condition of data source type s, Concat represents vector-level splicing operation, G ::Represents the vector representation of the expanded local sub-tensor in the core tensor corresponding to the (k, m, s) triple combination, and the concatenation result dimension is unified to the fusion feature dimension Where D = R N +R T +R S +R G ;

[0054] S42. All fused feature vectors F k,m,s Perform serialization operations in time order t1, t2, ..., t M Reorganize and construct a unified feature representation sequence F seq ,The unified feature representation sequence simultaneously contains local features and multi-dimensional interaction information in abnormal data of cloud computing edge nodes.

[0055] Optionally, the S5 includes the following steps:

[0056] S51. Unify the feature representation sequence F seq In the dynamic clustering module based on the evolution graph, the continuous time window set is set to W = {W1, W2, ..., W L}, where W l represents the lth continuous time window, L is the total number of continuous time windows;

[0057] S52. In each continuous time window W l In the sequence F, the uniform feature is used to represent seq The fused feature vector F in k,m,s Input data for clustering, based on a unified fusion feature dimension Perform initial clustering to generate an initial cluster set in, Indicates that in the lth time window W l The qth cluster in the cluster, Q l is the time window W l The total number of cohesive clusters;

[0058] S53. Based on adjacent continuous time windows W l-1 and W l The cluster set C l-1 and C l , construct cluster evolution graph G evo =(V evo ,E evo ), where the graph node set V evo The nodes in the graph represent clusters, and the graph edge set E evo The edges in represent the evolutionary relationship between clusters in adjacent time windows;

[0059] S54. Representing sequence F according to unified features seq The temporal changes of the fusion feature vector are calculated to calculate the cluster evolution graph G evo Similarity between adjacent window clusters in :

[0060]

[0061] in, Represents the uth cluster in the l-1th time window and the vth cluster in the lth time window The similarity between x 、F y Respectively represent the fusion feature vectors within the cluster;

[0062] S55. Cluster evolution graph G evo Bind with the corresponding relationship of the unified feature representation sequence to establish a mapping relationship between the continuous time clustering change trajectory and the data source behavior evolution;

[0063] S56. Introduce the cluster stability measurement function, calculate the cluster stability score of the lth time window, and construct the cluster stability score sequence:

[0064]

[0065] Cluster stability score Stab(W l ) represents the continuity degree of the cluster structure in the lth time window. The higher the stability score, the more stable the cluster structure and the smaller the change.

[0066] S57. Construct the cluster stability change rate function ΔStab(W) for each time window based on the cluster stability score sequence l )=Stab(W l )-Stab(W l-1 ), where the cluster stability change rate ΔStab(W l ) represents the current time window W l The stability score of the previous time window W l-1 The numerical difference between the stability scores of

[0067] When the cluster stability change rate ΔStab(W l ) and ΔStab(W l+1 ) are all less than the preset stability drop threshold ε1, the system determines that the clustering structure is in a continuously unstable state, and at this time the time window adaptive adjustment mechanism is triggered to shorten the next time window W l+2 The duration of the adjusted time window is the current time window length minus the window reduction step δ t, and at the same time ensure that the adjusted window length is not less than the minimum allowed window length W set by the system min ;

[0068] On the contrary, when the cluster stability change rate of two consecutive time windows is greater than the preset stability rising threshold ε2, it means that the cluster structure tends to be stable continuously. At this time, the window expansion mechanism is triggered and the next time window W is extended. l+2 The length of the time window is set to the current time window length plus the window growth step δ t and ensure that the maximum allowable window length W set by the system is not exceeded max ;

[0069] When the clustering stability score Stab(W l ) is less than the abnormal warning threshold θ alert , and the cluster stability change rate of this time window ΔStab(W l ) is less than the stability decline threshold ε1, the system determines that there is a mutation behavior in the clustering structure of the time window, triggers the abnormal clustering feedback mechanism, and marks the time window as a clustering structure mutation state, entering the abnormal data labeling process.

[0070] Optionally, the abnormal data labeling process includes the following steps:

[0071] Combined with the time window W marked as the cluster structure mutation state in the abnormal cluster feedback mechanism l , extract all fusion feature vectors F in the window k,m,s , and identify n according to the edge node k , sampling time t m and data source type s to establish a triple index correspondence to form a mutation window feature subset;

[0072] Extract the definition rules of common anomaly types of edge nodes from historical anomaly analysis records and build an abnormal behavior classification rule set. The abnormal behavior classification rule set is defined based on the abnormal fluctuation amplitude, cross-time feature offset degree, and low-density clustering isolation index in the feature dimension.

[0073] According to the following abnormal behavior classification rules, the fusion feature vectors in the cluster structure mutation window are labeled, where:

[0074] Anomaly type A: The current feature has the lowest cluster density in the history of the node dimension, and similar features appear isolated across two consecutive time windows;

[0075] Anomaly type B: The time change direction of the current feature is opposite to the historical cluster center trend, and the feature offset is greater than the average offset threshold δ drift ;

[0076] Anomaly type C: The sampling time corresponding to the current feature undergoes concentrated feature jumps between nodes, and obvious multi-source fusion conflict indicators appear, including contradictory feature vectors from multiple data sources;

[0077] S514. The abnormal fusion feature vector F that has been completed classification labeling k,m,s According to the corresponding edge node number n k and sampling time t m Perform structured output and build an abnormal data annotation result set, which records the abnormal label, abnormal cause and associated context information of each fused feature vector.

[0078] The beneficial effects of the present invention are:

[0079] (1) The present invention introduces a heterogeneous tensor decomposition mechanism with multi-dimensional adaptive attention enhancement. By constructing independent attention weight matrices in the edge node dimension, sampling time dimension and data source type dimension, and integrating comprehensive indicators such as historical abnormal frequency, node fluctuation amplitude and feature offset strength, the active enhancement processing of abnormal sensitive features in high-order tensors is achieved. The data tensor is reconstructed through three-dimensional weighted Hadamard operation and decomposed using the Tucker structure with non-negative constraints, which effectively improves the decomposition accuracy and discrimination of high-dimensional heterogeneous data in the abnormal expression layer, and improves the identifiability and aggregation of abnormal information.

[0080] (2) The present invention constructs an evolution graph clustering model based on a continuous time window, and designs a clustering stability measurement function and a change rate feedback mechanism. By calculating the similarity evolution trend between the cluster structures of each time window in real time, the time window length is dynamically adjusted to match the data mutation frequency, thereby realizing the rhythm control of clustering updates and the adaptive evolution of clustering structures, enhancing the time sensitivity and trend tracking ability of the clustering process, and improving the response efficiency and recognition accuracy of abnormal structure identification.

[0081] (3) The present invention proposes an abnormal feedback mechanism and a structured annotation process for cluster mutation status. When the cluster stability score is lower than the warning threshold and continues to decline, the abnormal feedback mark is automatically triggered, and the abnormal status is semantically classified based on the multi-dimensional behavioral characteristics of the fusion features, and finally an abnormal data annotation result set with contextual information is formed. Compared with the traditional method that relies on single-point threshold judgment, it has more structural stability, contextual logic traceability and abnormal classification interpretability, providing a reliable basis for subsequent edge collaborative warning, model incremental training and personalized response strategy, greatly improving the overall data availability and maintainability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0083] Figure 1 This is a flow chart of an artificial intelligence-based personalized customized data mining and analysis system proposed by the present invention;

[0084] Figure 2 This is a schematic diagram of a high-order data modeling structure based on three-dimensional tensors in an artificial intelligence-based personalized customized data mining and analysis system proposed by the present invention;

[0085] Figure 3 This is a process diagram of dynamic clustering and cluster evolution graph construction based on continuous time windows in an artificial intelligence-based personalized customized data mining and analysis system proposed by the present invention. DETAILED DESCRIPTION

[0086] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0087] refer to Figure 1-Figure 3 , a personalized customized data mining and analysis system based on artificial intelligence, including the following modules:

[0088] The edge data acquisition module is used to collect and preprocess multi-source heterogeneous raw data from multiple cloud computing edge nodes to generate structured multi-source heterogeneous data sets;

[0089] High-order data tensor construction module, used to construct high-order data tensors based on structured multi-source heterogeneous data sets;

[0090] The heterogeneous tensor decomposition module introduces the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decomposes the high-order data tensor using heterogeneous tensor decomposition technology, and extracts the potential feature matrix and core tensor;

[0091] The feature fusion and unified expression module is used to perform structured concatenation and joint encoding of the extracted latent feature matrix and the core tensor to construct a fused feature vector, and organize all fused feature vectors in chronological order to form a unified feature representation sequence;

[0092] Dynamic clustering analysis module, used to perform dynamic clustering analysis on the unified feature representation sequence in a continuous time window, construct a cluster set, and generate a cluster evolution graph based on the similarity relationship between clusters in adjacent time windows;

[0093] The cluster stability evaluation and feedback module is used to calculate the cluster stability score and stability change rate of each time window based on the cluster evolution graph, mark the cluster structure mutation state, and obtain the fusion feature vector under the cluster structure mutation state;

[0094] The abnormal data annotation module is used to classify and annotate the abnormal behavior of the fused feature vector under the state of cluster structure mutation, and output the abnormal data annotation result set.

[0095] A personalized data mining and analysis method based on artificial intelligence, used to execute a personalized data mining and analysis system based on artificial intelligence, characterized by comprising the following steps:

[0096] S1. Collect multi-source heterogeneous datasets from multiple cloud computing edge nodes, perform data preprocessing on the multi-source heterogeneous datasets, and generate structured multi-source heterogeneous datasets;

[0097] S2. Construct high-order data tensors based on structured multi-source heterogeneous datasets;

[0098] S3. Introduce the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decompose the high-order data tensor using heterogeneous tensor decomposition technology, and extract the latent feature matrix and core tensor;

[0099] S4. Perform joint feature fusion processing on the extracted latent feature matrix and core tensor to construct a fused feature vector. The fused feature vector is then organized in chronological order to form a unified feature representation sequence. The unified feature representation sequence contains both local structural information and global interaction information of abnormal data at the cloud computing edge node.

[0100] S5. Input the unified feature representation sequence into the dynamic clustering module based on the cluster evolution graph, set a set of continuous time windows, perform dynamic clustering to generate a set of cluster clusters in each time window, and construct a cluster evolution graph. Calculate the cluster stability score and stability change rate. When the cluster stability score is lower than the abnormal warning threshold and the stability change rate is lower than the stability decline threshold, the system triggers the abnormal clustering feedback mechanism and marks the time window as a cluster structure mutation state. The fused feature vector in the cluster structure mutation state is classified as abnormal behavior and the abnormal data annotation result set is output.

[0101] In this embodiment, S1 includes the following steps:

[0102] S11. Set the edge node set N in the cloud computing environment, n in the edge node set N k represents the kth edge node, from each edge node n kCollect multi-source heterogeneous raw data, including sensor data, log data, business traffic data and user behavior data, and build an initial multi-source heterogeneous dataset D raw ;

[0103] S12. Set the sampling time window T, the sampling time window T t m Indicates the mth sampling time. At each sampling time t m , edge node n k Collect raw data of corresponding types The original data points in the form of triples in, Indicates that the s-th data source is at edge node n k , sampling time t m The original data points at is the corresponding structured eigenvalue;

[0104] S13. Initial multi-source heterogeneous dataset D raw Perform unified preprocessing operations on all types of data in;

[0105] S14. All kinds of pre-processed data are sorted by edge node n k , sampling time t m , the triple dimensions of data source type s are recombined to generate a structured multi-source heterogeneous dataset:

[0106]

[0107] Among them, D struct is a structured multi-source heterogeneous dataset, K is the total number of edge nodes, M is the total number of sampling times, and S represents the number of data source categories.

[0108] In this embodiment, the pre-processing operation includes:

[0109] Time alignment processing interpolates or truncates different types of data based on a unified sampling time window T, so that all data are at a unified sampling time t m The above has a consistent alignment structure;

[0110] Data cleaning and processing to remove missing, duplicate or abnormally formatted raw data points and fill in missing values;

[0111] Format unification processing, converting unstructured data into structured form, so that all data can be represented as standard numerical feature vectors;

[0112] Noise removal processing, structural eigenvalues ​​of time series The filtering algorithm is used for smoothing to remove high-frequency noise interference.

[0113] In this embodiment, S2 includes the following steps:

[0114] S21. In structured multi-source heterogeneous dataset D struct On this basis, a unified three-dimensional index system is established, which includes the edge node dimension N, the sampling time dimension T and the number of data source categories S;

[0115] S22. Based on the triple-dimensional index system, edge node n k , sampling time t m And the data source type s is a joint index, which extracts the edge node n from the structured multi-source heterogeneous data set k , sampling time t m The structured feature value corresponding to the s-th data source type Used to represent observation data of a single tensor unit;

[0116] S23. Arrange all the structured eigenvalues ​​extracted based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S in three dimensions to construct a third-order tensor structure and generate a high-order data tensor X. The third-order structure of the high-order data tensor X corresponds to the edge node dimension, the sampling time dimension, and the number of data source categories, respectively. The value of each tensor unit in the high-order data tensor is the structured eigenvalue.

[0117] In this embodiment, S3 includes the following steps:

[0118] S31. A high-order data tensor X is introduced into the heterogeneous tensor decomposition module of the multi-dimensional adaptive attention enhancement mechanism. By calculating the adaptive attention weight matrix based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S, the module automatically identifies the data dimensions and features that contribute most to abnormal events when the cloud computing edge node is in an abnormal state.

[0119] S32. The probability of abnormal events at edge nodes, the magnitude of abnormal data fluctuations, and the historical abnormal correlation strength are integrated into a unified multi-dimensional abnormal attention index. The attention weight matrix of the edge node dimension is set as The attention weight matrix of the sampling time dimension is The attention weight matrix of the number of data source categories is Among them, the edge node attention weight element Represents edge node n k The comprehensive abnormal attention degree is a combination of the cumulative frequency of historical abnormal events and the fluctuation amplitude of the current node data, and the sampling time attention weight element represents the sampling time t mThe intensity of data fluctuations and the combined abnormal attention of abnormal trend prediction values, data source type attention weight element The abnormal attention degree is represented by the contribution of data source type s to the detection of historical abnormal events and the difference between the current node data. The Hadamard element-by-element product operation of the three-dimensional tensor is used to calculate the attention weight matrix of the edge node dimension, the attention weight matrix of the time dimension, and the attention weight matrix of the number of data source categories to obtain the adaptive attention weight matrix;

[0120] S33. Use the adaptive attention weight matrix to perform element-by-element multi-dimensional adaptive weighted adjustment on the high-order data tensor X to form an abnormal sensitivity enhanced weighted high-order data tensor X * :

[0121]

[0122] in, It is a weighted high-order data tensor element after abnormal sensitivity enhancement, reflecting the weight amplification effect of data points with higher abnormal contribution in the abnormal event mining task of cloud computing edge nodes;

[0123] S34. Using Tucker decomposition structure with non-negative constraints to weight high-order data tensor X * Implement a heterogeneous tensor decomposition process based on the joint optimization of non-negativity constraints and anomaly-sensitive loss functions to obtain the latent feature matrix and core tensor:

[0124] X * ≈G×1U N ×2U T ×3U S ;

[0125] The decomposition process is achieved by defining an abnormality-sensitive loss function L anom Optimization solution:

[0126]

[0127] Among them, G is the core tensor that captures the interaction relationship of multi-dimensional abnormal data, and the potential feature matrix U N 、U T 、U S They respectively reflect the abnormal characteristics in three dimensions: edge nodes, sampling time, and data source type. λ is the sparse regularization coefficient used to control the sensitivity to abnormalities.

[0128] This paper introduces a heterogeneous tensor decomposition method with a multi-dimensional adaptive attention mechanism, which significantly improves the flexibility, pertinence, and high-dimensional feature expression capabilities of edge node abnormal data modeling:

[0129] Unlike traditional tensor decomposition methods that assign equal weights to each dimension, this method combines the distribution characteristics of abnormal behavior at the edge of cloud computing nodes and introduces adaptive attention weights in the edge node dimension, sampling time dimension, and data source type dimension. This method uses statistical indicators such as the historical frequency of node anomalies, data fluctuation amplitude, time evolution trend, and data source differences to construct a multidimensional anomaly attention weight matrix. While maintaining the integrity of the tensor structure, data units with higher anomaly contributions are given greater modeling weights, making the latent tensor structure more sensitive to key behavioral characteristics, thereby enhancing the expressiveness and discriminative power of anomaly structure detection.

[0130] The joint optimization model of anomaly-sensitive weighted tensors and tensor decomposition with non-negative constraints is constructed to effectively alleviate the sparsity and redundancy problems of high-dimensional heterogeneous data. The present invention applies the three-dimensional attention matrix to the original high-order data tensor through element-by-element Hadamard product operations to generate an anomaly-sensitive enhanced weighted tensor. On this basis, the Tucker decomposition structure is introduced. At the same time, a joint optimization objective function is constructed by combining non-negative constraints and anomaly-sensitive regularization terms, encouraging the latent feature matrix to pay attention to local nonlinear offset phenomena in the data while maintaining a sparse structure. This method not only retains the advantages of tensor decomposition for modeling high-order structures, but also introduces a directional modeling mechanism in anomaly detection tasks, achieving a dynamic balance between structural compression and information amplification, and showing stronger convergence and stability in large-scale heterogeneous data processing environments.

[0131] Combining the multi-dimensional attention tensor weighting mechanism with the cloud-edge abnormal event modeling task has achieved a paradigm shift in the anomaly detection task from "structural learning" to "behavioral attention". Existing tensor modeling methods are mostly used for compression and reconstruction tasks, and lack an identification mechanism for high-risk behavior fragments. The present invention converts the node behavior evolution, time disturbance trend and data source disturbance expression under abnormal conditions into a structural weight matrix, forming a tensor control strategy driven by "abnormal behavior contribution" as the core, with strong contextual adaptability and dynamic expression capabilities, so that the model can not only compress information redundancy at the data level, but also enhance the recognition ability of complex edge abnormal patterns at the structural level.

[0132] In summary, the heterogeneous tensor decomposition method in this implementation breaks through the limitations of equal-weight modeling and static optimization of traditional tensor decomposition in anomaly detection. It has significant technical innovation and practical value both in theoretical design and application. It can effectively solve the problem of anomaly mining of high-dimensional, heterogeneous, and dynamic data in the cloud computing edge node environment, and improve the system's intelligent analysis capabilities and response efficiency.

[0133] In this embodiment, S4 includes the following steps:

[0134] S41. Potential feature matrix U N 、UT 、U S Perform joint feature fusion processing with the core tensor G, structurally splice the tensor-level global interaction features with the low-rank features of each dimension to form a ternary joint feature representation set:

[0135] F k,m,s =Concat([U N (k,:)],[U T (m,:)],[U S (s,:)],[G :: ]);

[0136] Among them, F k,m,s Indicates that at edge node n k , sampling time t m , fusion feature vector under the condition of data source type s, Concat represents vector-level splicing operation, G :: Represents the vector representation of the expanded local sub-tensor in the core tensor corresponding to the (k, m, s) triple combination, and the concatenation result dimension is unified to the fusion feature dimension Where D = R N +R T +R S +R G ;

[0137] S42. All fused feature vectors F k,m,s Perform serialization operations in time order t1, t2, ..., t M Reorganize and construct a unified feature representation sequence F seq ,The unified feature representation sequence simultaneously contains local features and multi-dimensional interaction information in abnormal data of cloud computing edge nodes.

[0138] In this embodiment, S5 includes the following steps:

[0139] S51. Unify the feature representation sequence F seq In the dynamic clustering module based on the evolution graph, the continuous time window set is set to W = {W1, W2, ..., W L}, where W l represents the lth continuous time window, L is the total number of continuous time windows;

[0140] S52. In each continuous time window W l In the sequence F, the uniform feature is used to represent seq The fused feature vector F in k,m,s Input data for clustering, based on a unified fusion feature dimension Perform initial clustering to generate an initial cluster set in, Indicates that in the lth time window W l The qth cluster in the cluster, Q l is the time window W l The total number of cohesive clusters;

[0141] S53. Based on adjacent continuous time windows W l-1 and W l The cluster set C l-1 and C l , construct cluster evolution graph G evo =(V evo ,E evo ), where the graph node set V evo The nodes in the graph represent clusters, and the graph edge set E evo The edges in represent the evolutionary relationship between clusters in adjacent time windows;

[0142] S54. Representing sequence F according to unified features seq The temporal changes of the fusion feature vector are calculated to calculate the cluster evolution graph G evo Similarity between adjacent window clusters in :

[0143]

[0144] in, Represents the uth cluster in the l-1th time window and the vth cluster in the lth time window The similarity between x 、F y Respectively represent the fusion feature vectors within the cluster;

[0145] S55. Cluster evolution graph G evo Bind with the corresponding relationship of the unified feature representation sequence to establish a mapping relationship between the continuous time clustering change trajectory and the data source behavior evolution;

[0146] S56. Introduce the cluster stability measurement function, calculate the cluster stability score of the lth time window, and construct the cluster stability score sequence:

[0147]

[0148] Cluster stability score Stab(W l ) represents the continuity degree of the cluster structure in the lth time window. The higher the stability score, the more stable the cluster structure and the smaller the change.

[0149] S57. Construct the cluster stability change rate function ΔStab(W) for each time window based on the cluster stability score sequence l)=Stab(W l )-Stab(W l-1 ), where the cluster stability change rate ΔStab(W l ) represents the current time window W l The stability score of the previous time window W l-1 The numerical difference between the stability scores of

[0150] When the cluster stability change rate ΔStab(W l ) and ΔStab(W l+1 ) are all less than the preset stability drop threshold ε1, the system determines that the clustering structure is in a continuously unstable state, and at this time the time window adaptive adjustment mechanism is triggered to shorten the next time window W l+2 The duration of the adjusted time window is the current time window length minus the window reduction step δ t , and at the same time ensure that the adjusted window length is not less than the minimum allowed window length W set by the system min ;

[0151] On the contrary, when the cluster stability change rate of two consecutive time windows is greater than the preset stability rising threshold ε2, it means that the cluster structure tends to be stable continuously. At this time, the window expansion mechanism is triggered and the next time window W is extended. l+2 The length of the time window is set to the current time window length plus the window growth step δ t and ensure that the maximum allowable window length W set by the system is not exceeded max ;

[0152] When the clustering stability score Stab(W l ) is less than the abnormal warning threshold θ alert , and the cluster stability change rate of this time window ΔStab(W l ) is less than the stability decline threshold ε1, the system determines that there is a mutation behavior in the clustering structure of the time window, triggers the abnormal clustering feedback mechanism, and marks the time window as a clustering structure mutation state, entering the abnormal data labeling process.

[0153] This implementation introduces a dynamic clustering mechanism based on a cluster evolution graph, supplemented by a cluster stability metric and a time window adaptive adjustment algorithm. This effectively addresses key issues in existing anomaly detection algorithms for dynamic data streams at cloud computing edge nodes, such as static modeling lag, difficulty identifying mutations, and lack of flexibility in cluster scheduling. Specifically, it achieves the following three innovative and beneficial effects:

[0154] A cluster evolution graph modeling mechanism is proposed to achieve structured expression and tracking of temporal continuous clustering behavior. The present invention establishes node representations of cluster clusters and edge connections of cross-window similarity in continuous time windows to form a cluster evolution graph. This structure models behaviors such as cluster centers, cluster merging, and evolutionary trends at the graph level, enabling the system to process cluster change trajectories under complex time series using graph computing. Compared with existing methods that can only identify anomalies through static analysis of independent windows, the present invention significantly enhances the structural interpretability of abnormal behaviors and the traceability of behavioral evolution by starting from evolutionary relationships.

[0155] A clustering stability metric function and a rate-of-change function are introduced to implement an adaptive time window control mechanism based on behavioral trends. A stability metric is formed by calculating the maximum matching similarity between the clustering results of each time window and the clustering results of the previous window. The window size is adjusted in real time based on continuous downward or upward trends, allowing the clustering process to keep pace with the evolution of the data stream. Compared with traditional fixed-window clustering methods, this method can effectively adapt to changes in data anomaly density and fluctuations in evolution rate, improving detection accuracy while controlling computational load. It has high engineering feasibility and theoretical novelty.

[0156] An abnormal clustering structure mutation triggering mechanism and behavioral feedback process have been constructed to enhance the system's self-learning and dynamic adjustment capabilities. The present invention automatically triggers the "clustering structure mutation" judgment when the cluster stability fluctuation exceeds the warning threshold and the cluster score is lower than the set value, and sends the corresponding window to the abnormal analysis module to perform abnormal type classification, context feature annotation and structural data output. Unlike the existing algorithms that simply rely on sample outlier degree for judgment, the present invention triggers abnormal judgment based on the rate of change of structural evolution, which is closer to the real system evolution mechanism, significantly reduces the false alarm rate and missed detection rate, and improves the system's autonomous response capability.

[0157] In this embodiment, the abnormal data annotation process includes the following steps:

[0158] Combined with the time window W marked as the cluster structure mutation state in the abnormal cluster feedback mechanism l , extract all fusion feature vectors F in the window k,m,s , and identify n according to the edge node k , sampling time t m and data source type s to establish a triple index correspondence to form a mutation window feature subset;

[0159] Extract the definition rules of common anomaly types of edge nodes from historical anomaly analysis records and build an abnormal behavior classification rule set. The abnormal behavior classification rule set is defined based on the abnormal fluctuation amplitude, cross-time feature offset degree, and low-density clustering isolation index in the feature dimension.

[0160] According to the following abnormal behavior classification rules, the fusion feature vectors in the cluster structure mutation window are labeled, where:

[0161] Anomaly type A: The current feature has the lowest cluster density in the history of the node dimension, and similar features appear isolated across two consecutive time windows;

[0162] Anomaly type B: The time change direction of the current feature is opposite to the historical cluster center trend, and the feature offset is greater than the average offset threshold δ drift ;

[0163] Anomaly type C: The sampling time corresponding to the current feature undergoes concentrated feature jumps between nodes, and obvious multi-source fusion conflict indicators appear, including contradictory feature vectors from multiple data sources;

[0164] S514. The abnormal fusion feature vector F that has been completed classification labeling k,m,s According to the corresponding edge node number n k and sampling time t m Perform structured output and build an abnormal data annotation result set, which records the abnormal label, abnormal cause and associated context information of each fused feature vector.

[0165] Example 1:

[0166] On September 5, 2024, when a provincial power dispatching data center was performing distribution monitoring tasks, multiple edge computing nodes deployed between the substation and the feeder terminal equipment began to continuously report multiple abnormal signals. The data center applied the present invention to perform real-time monitoring and abnormality analysis of operating data. Each edge node has real-time access to five types of data sources, including control command logs of distribution automation devices, low-voltage load curves, current and voltage fluctuation information, external environmental sensor data, and local network access logs.

[0167] After the system was started, it entered a continuous operation phase, sampling data in real time at 10-second frames. At 3:17:30 PM on September 5, 2024, the system noticed significant anomalies in the tensor weighting results of multiple samples in the fused feature sequence uploaded by edge node "Node-B23." One record showed: edge node ID "Node-B23," sampling time "3:17:30 PM," data source type "low-voltage load," and the corresponding fused feature value "F{B23,109,2}=2.714," significantly higher than three standard deviations of the seven-day historical average.

[0168] Further cluster evolution graph analysis found that the stability score of the current time window is 0.28, which is lower than the threshold θ alert=0.5, and the similarity with the previous window drops by more than 0.2, the system identifies it as a "cluster structure mutation state." At this point, the dynamic clustering module automatically triggers the abnormal data feedback mechanism and enters the data annotation process.

[0169] The system first extracts all fusion feature sequences within the current anomaly window and finds that the following features have highly consistent "anomaly contributions":

[0170] From the edge node "Node-B23", the source device is "Feeder Terminal FT-007";

[0171] In the voltage fluctuation data, the three-phase instantaneous voltage showed a non-periodic sharp drop, with the maximum drop of 31.7% and a duration of about 5.5 seconds;

[0172] The load power factor deviated abnormally, dropping from the normal value of 0.92 to 0.35, and recovered slowly;

[0173] During the same time period, network logs recorded that a host with the source IP address "192.168.23.104" sent a large number of short-term, high-frequency UDP requests to "Node-B23" with the destination port number "8883," potentially attempting to interfere with the MQTT data channel of the edge node.

[0174] The system identifies this type of event as a typical "heterogeneous channel conflict interference anomaly."

[0175] Combined with the historical template, the present invention automatically labels the abnormal event as "composite communication-load behavior mutation abnormality" and classifies it as abnormality type C.

[0176] The abnormal event report is as follows:

[0177] Event ID: ANOM_202409051517_B23;

[0178] Edge node: Node-B23;

[0179] Abnormal start time: 2024-09-05 15:17:30;

[0180] Detection source: tensor feature sequence fusion abnormal points + clustering structure mutation;

[0181] Main abnormal characteristics: transient voltage abnormality, sudden drop in power factor, and UDP high-frequency injection;

[0182] Source IP address: 192.168.23.104

[0183] Destination port: 8883

[0184] Abnormal label: Class C - Communication interference type load mutation;

[0185] System response: The 15-second window reduction mechanism is triggered, and an alarm signal is sent to the cloud-based control platform;

[0186] Disposal operation: The administrator received the alarm at 15:18:00, remotely disconnected the "Node-B23" MQTT channel, and switched to the backup node "Node-B18". The system returned to normal at 15:18:32.

[0187] To verify the superiority of the method of the present invention, the project team conducted comparative tests on the method of the present invention and the traditional method in real abnormal sample identification tasks during a 14-day continuous monitoring mission in September 2024. The test scenarios covered four major types of typical events, including communication anomalies, current fluctuations, sensor drift, and control logic conflicts. The total number of sampling windows reached 17,200 frames, with 638 abnormal samples and 16,562 normal samples.

[0188] Table 1 Experimental comparison data results (partial)

[0189]

[0190] In addition, in terms of specific event recognition capabilities, the method of the present invention still maintains an accurate recognition rate of over 90% when multiple abnormal sources exist simultaneously (load mutation superimposed on network injection), while traditional methods often misclassify abnormalities, especially when the fusion data dimension increases to five dimensions or above, the recognition effect decreases significantly.

[0191] Finally, the management platform combined the anomaly annotation dataset generated by the method of the present invention and successfully used it for subsequent automatic model updates and policy adjustments, realizing intelligent autonomy of edge nodes. It was officially deployed and launched in mid-September 2024, and the average daily number of warning triggers dropped from the original average of 8.7 times to 2.3 times, effectively avoiding most false alarms and missed detections.

[0192] The present invention introduces a heterogeneous tensor decomposition mechanism with multi-dimensional adaptive attention enhancement. By constructing independent attention weight matrices in the edge node dimension, sampling time dimension and data source type dimension, and integrating comprehensive indicators such as historical abnormality frequency, node fluctuation amplitude and feature offset strength, it realizes active enhancement processing of abnormal sensitive features in high-order tensors. The data tensor is reconstructed through three-dimensional weighted Hadamard operation and decomposed using the Tucker structure with non-negative constraints, which effectively improves the decomposition accuracy and discrimination of high-dimensional heterogeneous data in the abnormal expression layer, and improves the identifiability and aggregation of abnormal information.

[0193] The present invention constructs an evolution graph clustering model based on continuous time windows, and designs a clustering stability measurement function and a change rate feedback mechanism. By calculating the similarity evolution trend between the cluster structures of each time window in real time, the time window length is dynamically adjusted to match the data mutation frequency, thereby realizing the rhythm control of clustering updates and the adaptive evolution of clustering structures, enhancing the time sensitivity and trend tracking ability of the clustering process, and improving the response efficiency and recognition accuracy of abnormal structure identification.

[0194] The present invention proposes an abnormal feedback mechanism and structured labeling process for cluster mutation status. When the cluster stability score is lower than the warning threshold and continues to decline, the abnormal feedback mark is automatically triggered, and the abnormal status is semantically classified based on the multi-dimensional behavioral characteristics of the fusion features, and finally an abnormal data labeling result set with contextual information is formed. Compared with the traditional method that relies on single-point threshold judgment, it has more structural stability, contextual logic traceability and abnormal classification interpretability, providing a reliable basis for subsequent edge collaborative warning, model incremental training and personalized response strategies, and greatly improving the overall data availability and maintainability of the system.

[0195] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A personalized customized data mining and analysis system based on artificial intelligence, characterized by: Includes the following modules: The edge data acquisition module is used to collect and preprocess multi-source heterogeneous raw data from multiple cloud computing edge nodes to generate structured multi-source heterogeneous data sets; High-order data tensor construction module, used to construct high-order data tensors based on structured multi-source heterogeneous data sets; The heterogeneous tensor decomposition module introduces the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decomposes the high-order data tensor using heterogeneous tensor decomposition technology, and extracts the potential feature matrix and core tensor; The feature fusion and unified expression module is used to perform structured concatenation and joint encoding of the extracted latent feature matrix and the core tensor to construct a fused feature vector, and organize all fused feature vectors in chronological order to form a unified feature representation sequence; Dynamic clustering analysis module, used to perform dynamic clustering analysis on the unified feature representation sequence in a continuous time window, construct a cluster set, and generate a cluster evolution graph based on the similarity relationship between clusters in adjacent time windows; The cluster stability evaluation and feedback module is used to calculate the cluster stability score and stability change rate of each time window based on the cluster evolution graph, mark the cluster structure mutation state, and obtain the fusion feature vector under the cluster structure mutation state; The abnormal data annotation module is used to classify and annotate the abnormal behavior of the fused feature vector under the state of cluster structure mutation, and output the abnormal data annotation result set.

2. A personalized customization data mining and analysis method based on artificial intelligence, used to execute the personalized customization data mining and analysis system based on artificial intelligence according to claim 1, characterized in that: The steps include: S1. Collect multi-source heterogeneous datasets from multiple cloud computing edge nodes, perform data preprocessing on the multi-source heterogeneous datasets, and generate structured multi-source heterogeneous datasets; S2. Construct high-order data tensors based on structured multi-source heterogeneous datasets; S3. Introduce the constructed high-order data tensor input into the heterogeneous tensor decomposition module of the attention weight mechanism, decompose the high-order data tensor using heterogeneous tensor decomposition technology, and extract the latent feature matrix and core tensor; S4. Perform joint feature fusion processing on the extracted latent feature matrix and core tensor to construct a fused feature vector. The fused feature vector is then organized in chronological order to form a unified feature representation sequence. The unified feature representation sequence contains both local structural information and global interaction information of abnormal data at the cloud computing edge node. S5. Input the unified feature representation sequence into the dynamic clustering module based on the cluster evolution graph, set a set of continuous time windows, perform dynamic clustering to generate a set of cluster clusters in each time window, and construct a cluster evolution graph. Calculate the cluster stability score and stability change rate. When the cluster stability score is lower than the abnormal warning threshold and the stability change rate is lower than the stability decline threshold, the system triggers the abnormal clustering feedback mechanism and marks the time window as a cluster structure mutation state. The fused feature vector in the cluster structure mutation state is classified as abnormal behavior and the abnormal data annotation result set is output.

3. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 2, characterized in that: Said S1 comprises the following steps: S11. Set the edge node set N in the cloud computing environment, n in the edge node set N k represents the kth edge node, from each edge node n k Collect multi-source heterogeneous raw data, including sensor data, log data, business traffic data and user behavior data, and build an initial multi-source heterogeneous dataset D raw ; S12. Set the sampling time window T, the sampling time window T t m Indicates the mth sampling time. At each sampling time t m , edge node n k Collect raw data of corresponding types The original data points in the form of triples in, Indicates that the s-th data source is at edge node n k , sampling time t m The original data points at is the corresponding structured eigenvalue; S13. Initial multi-source heterogeneous dataset D raw Perform unified preprocessing operations on all types of data in; S14. All kinds of pre-processed data are sorted by edge node n k , sampling time t m , the triple dimension of data source type s is recombined to generate a structured multi-source heterogeneous dataset D struct .

4. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 3, characterized in that: The pre-processing operation includes: Time alignment processing interpolates or truncates different types of data based on a unified sampling time window T, so that all data are at a unified sampling time t m The above has a consistent alignment structure; Data cleaning and processing to remove missing, duplicate or abnormally formatted raw data points and fill in missing values; Format unification processing, converting unstructured data into structured form, so that all data can be represented as standard numerical feature vectors; Noise removal processing uses a filtering algorithm to smooth the structured eigenvalues ​​of the time series and remove high-frequency noise interference.

5. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 3, characterized in that: The S2 comprises the following steps: S21. In structured multi-source heterogeneous dataset D struct On this basis, a unified three-dimensional index system is established, which includes the edge node dimension N, the sampling time dimension T and the number of data source categories S; S22. Based on the triple-dimensional index system, edge node n k , sampling time t m And the data source type s is a joint index, which extracts the edge node n from the structured multi-source heterogeneous data set k , sampling time t m The structured feature value corresponding to the s-th data source type Used to represent observation data of a single tensor unit; S23. Arrange all the structured eigenvalues ​​extracted based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S in three dimensions to construct a third-order tensor structure and generate a high-order data tensor X. The third-order structure of the high-order data tensor X corresponds to the edge node dimension, the sampling time dimension, and the number of data source categories, respectively. The value of each tensor unit in the high-order data tensor is the structured eigenvalue.

6. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 5, characterized in that: The S3 includes the following steps: S31. A high-order data tensor X is introduced into the heterogeneous tensor decomposition module of the multi-dimensional adaptive attention enhancement mechanism. By calculating the adaptive attention weight matrix based on the edge node dimension N, the sampling time dimension T, and the number of data source categories S, the module automatically identifies the data dimensions and features that contribute most to abnormal events when the cloud computing edge node is in an abnormal state. S32. The probability of abnormal events at edge nodes, the magnitude of abnormal data fluctuations, and the historical abnormal correlation strength are integrated into a unified multi-dimensional abnormal attention index. The attention weight matrix of the edge node dimension is set as The attention weight matrix of the sampling time dimension is The attention weight matrix of the number of data source categories is Among them, the edge node attention weight element Represents edge node n k The comprehensive abnormal attention degree is a combination of the cumulative frequency of historical abnormal events and the fluctuation amplitude of the current node data, and the sampling time attention weight element represents the sampling time t m The intensity of data fluctuations and the combined abnormal attention of abnormal trend prediction values, data source type attention weight element The abnormal attention degree is represented by the contribution of data source type s to the detection of historical abnormal events and the difference between the current node data. The Hadamard element-by-element product operation of the three-dimensional tensor is used to calculate the attention weight matrix of the edge node dimension, the attention weight matrix of the time dimension, and the attention weight matrix of the number of data source categories to obtain the adaptive attention weight matrix; S33. Use the adaptive attention weight matrix to perform element-by-element multi-dimensional adaptive weighted adjustment on the high-order data tensor X to form an abnormal sensitivity enhanced weighted high-order data tensor X * ; S34. Using Tucker decomposition structure with non-negative constraints to weight high-order data tensor X * Implement a heterogeneous tensor decomposition process based on the joint optimization of non-negativity constraints and anomaly-sensitive loss functions to obtain the latent feature matrix and core tensor: X * ≈G×1U N ×2U T ×3U S ; The decomposition process is achieved by defining an abnormality-sensitive loss function L anom Optimization solution: Among them, G is the core tensor that captures the interaction relationship of multi-dimensional abnormal data, and the potential feature matrix U N 、U T 、U S They respectively reflect the abnormal characteristics in three dimensions: edge nodes, sampling time, and data source type. λ is the sparse regularization coefficient used to control the sensitivity to abnormalities.

7. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 6, characterized in that: The S4 comprises the following steps: S41. Potential feature matrix U N 、U T 、U S Perform joint feature fusion processing with the core tensor G, structurally splice the tensor-level global interaction features with the low-rank features of each dimension to form a ternary joint feature representation set: F k,m,s =Concat([U N (k,:)],[U T (m,:)],[U S (s,:)],[G :: ]); Among them, F k,m,s Indicates that at edge node n k , sampling time t m , fusion feature vector under the condition of data source type s, Concat represents vector-level splicing operation, G :: Represents the vector representation of the expanded local sub-tensor in the core tensor corresponding to the (k, m, s) triple combination, and the concatenation result dimension is unified to the fusion feature dimension Where D = R N +R T +R S +R G ; S42. All fused feature vectors F k,m,s Perform serialization operations in time order t1, t2, ..., t M Reorganize and construct a unified feature representation sequence F seq .

8. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 7, characterized in that: The S5 comprises the following steps: S51. Unify the feature representation sequence F seq In the dynamic clustering module based on the evolution graph, the continuous time window set is set to W = {W1, W2, ..., W L }, where W l represents the lth continuous time window, L is the total number of continuous time windows; S52. In each continuous time window W l In the sequence F, the uniform feature is used to represent seq The fused feature vector F in k,m,s Input data for clustering, based on a unified fusion feature dimension Perform initial clustering to generate an initial cluster set in, Indicates that in the lth time window W l The qth cluster in the cluster, Q l is the time window W l The total number of cohesive clusters; S53. Based on adjacent continuous time windows W l-1 and W l The cluster set C l-1 and C l , construct cluster evolution graph G evo =(V evo ,E evo ), where the graph node set V evo The nodes in the graph represent clusters, and the graph edge set E evo The edges in represent the evolutionary relationship between clusters in adjacent time windows; S54. Representing sequence F according to unified features seq The temporal changes of the fusion feature vector are calculated to calculate the cluster evolution graph G evo Similarity between adjacent window clusters in : in, Represents the uth cluster in the l-1th time window and the vth cluster in the lth time window The similarity between x 、F y Respectively represent the fusion feature vectors within the cluster; S55. Cluster evolution graph G evo Bind with the corresponding relationship of the unified feature representation sequence to establish a mapping relationship between the continuous time clustering change trajectory and the data source behavior evolution; S56. Introduce the cluster stability measurement function, calculate the cluster stability score of the lth time window, and construct the cluster stability score sequence: Cluster stability score Stab(W l ) represents the continuity degree of the cluster structure in the lth time window. The higher the stability score, the more stable the cluster structure and the smaller the change. S57. Construct the cluster stability change rate function ΔStab(W) for each time window based on the cluster stability score sequence l )=Stab(W l )-Stab(W l-1 ), where the cluster stability change rate ΔStab(W l ) represents the current time window W l The stability score of the previous time window W l-1 The numerical difference between the stability scores of When the cluster stability change rate ΔStab(W l ) and ΔStab(W l+1 ) are all less than the preset stability drop threshold ε1, the system determines that the clustering structure is in a continuously unstable state, and at this time the time window adaptive adjustment mechanism is triggered to shorten the next time window W l+2 The duration of the adjusted time window is the current time window length minus the window reduction step δ t , and at the same time ensure that the adjusted window length is not less than the minimum allowed window length W set by the system min ; On the contrary, when the cluster stability change rate of two consecutive time windows is greater than the preset stability rising threshold ε2, it means that the cluster structure tends to be stable continuously. At this time, the window expansion mechanism is triggered and the next time window W is extended. l+2 The length of the time window is set to the current time window length plus the window growth step δ t and ensure that the maximum allowable window length W set by the system is not exceeded max ; When the clustering stability score Stab(W l ) is less than the abnormal warning threshold θ alert , and the cluster stability change rate of this time window ΔStab(W l ) is less than the stability decline threshold ε1, the system determines that there is a mutation behavior in the clustering structure of the time window, triggers the abnormal clustering feedback mechanism, and marks the time window as a clustering structure mutation state, entering the abnormal data labeling process.

9. The method for personalized customized data mining and analysis based on artificial intelligence according to claim 8, characterized in that: The abnormal data annotation process includes the following steps: Combined with the time window W marked as the cluster structure mutation state in the abnormal cluster feedback mechanism l , extract all fusion feature vectors F in the window k,m,s , and identify n according to the edge node k , sampling time t m and data source type s to establish a triple index correspondence to form a mutation window feature subset; Extract the definition rules of common anomaly types of edge nodes from historical anomaly analysis records and build an abnormal behavior classification rule set. The abnormal behavior classification rule set is defined based on the abnormal fluctuation amplitude, cross-time feature offset degree, and low-density clustering isolation index in the feature dimension. According to the following abnormal behavior classification rules, the fusion feature vectors in the cluster structure mutation window are labeled, where: Anomaly type A: The current feature has the lowest cluster density in the history of the node dimension, and similar features appear isolated across two consecutive time windows; Anomaly type B: The time change direction of the current feature is opposite to the historical cluster center trend, and the feature offset is greater than the average offset threshold δ drift ; Anomaly type C: The sampling time corresponding to the current feature undergoes concentrated feature jumps between nodes, and obvious multi-source fusion conflict indicators appear, including contradictory feature vectors from multiple data sources; S514. The abnormal fusion feature vector F that has been completed classification labeling k,m,s According to the corresponding edge node number n k and sampling time t m Perform structured output and build an abnormal data annotation result set, which records the abnormal label, abnormal cause and associated context information of each fused feature vector.

Citation Information

Cited By

  • Large model tensor database construction method and system

    CN120910029A

  • Tensor database construction method and system of large model

    CN120910029B

  • Communication equipment production intelligent management system based on machine learning

    CN121350587A