Automatic data labeling and classifying method for artificial intelligence model
By combining clustering algorithms, graph convolution networks and reinforcement learning strategies, the labeling error problem caused by ignoring the complex relationship of data points in the existing technology is solved, and high-precision data node classification and continuous optimization are achieved.
Patent Information
- Application Number
- CN202510584938.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data annotation method based on clustering algorithm ignores the complex relationship between data points, resulting in errors in the annotation results. Especially in the processing of high-dimensional data or nonlinear data, inaccurate classification is prone to occur.
The clustering algorithm is used to form data nodes, and the node local features are extracted through the graph convolution network, combined with mutual information loss function and graph regularization term loss function for training, the node hierarchy structure is represented through the generalized Voronoi graph, the node and edge features are updated, the graph structure is formed for node classification annotation, and the graph convolution network is optimized based on reinforcement learning strategy.
It effectively improves the accuracy of node feature extraction, improves the accuracy of node classification, realizes continuous optimization of real-time data, and significantly improves the efficiency of data processing and classification accuracy.
Smart Images

Figure CN120105166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing classification, and in particular to a method for automatic data labeling and classification for an artificial intelligence model. Background Art
[0002] The rapid development of artificial intelligence (AI) has promoted the progress of data labeling and classification technology. In the training process of AI models, data labeling is a crucial step. Traditional data labeling methods usually rely on manual operations, which are time-consuming and susceptible to human bias. With the development of deep learning and machine learning technologies, more and more automated data labeling methods have emerged. Existing automatic labeling technologies are mainly implemented through rule-based systems, supervised learning, semi-supervised learning and other methods. Among them, automatic labeling methods based on clustering algorithms have become an important research direction. Clustering algorithms automatically group data by analyzing the similarity or distance relationship of data to form a preliminary label set. At the same time, the application of graph neural networks (GNNs) in data labeling and classification has gradually increased. Graph convolutional networks (GCNs), as an important GNN method, can effectively improve the accuracy of data analysis by aggregating local node features, especially when processing complex data structures. In recent years, automatic labeling methods combining clustering algorithms and graph convolutional networks have gradually become a research hotspot, especially in the scenarios of large-scale data sets and real-time data processing, and have high application prospects.
[0003] However, the existing technology still has shortcomings. Data labeling methods based on clustering algorithms often ignore the complex relationships between data points, resulting in errors in the labeling results. Especially in the processing of high-dimensional data or nonlinear data, inaccurate classification is prone to occur, and there is often a lack of effective modeling of the relationship between nodes, resulting in inaccurate node feature updates, which in turn affects the final classification effect. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for automatic data labeling and classification for artificial intelligence models, which solves the problem that the prior art often ignores the complex relationship between data points, resulting in errors in the labeling results and often lacks effective modeling of the relationship between nodes, resulting in inaccurate node feature updates.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for automatic data annotation and classification for an artificial intelligence model, which comprises: Based on the database, we collect structured numerical data, use clustering algorithms to perform clustering analysis on the data to form data nodes, extract local features of the nodes through graph convolutional networks, and train the discriminator through mutual information loss functions to filter local features of the nodes, and perform dimensionality reduction processing simultaneously; The hierarchical structure of nodes is represented by a generalized Voronoi diagram, and a graph regularization loss function is defined. Graph convolution operations are used in each layer to update node features. Edge features are updated through the interaction of graph regularization and node features, and node features are updated again through nonlinear transformation using the updated edge features. A graph structure is formed based on the updated node features and edge features, and node classification and labeling is performed. The node classification annotations are displayed in real time, the classification records are stored in the database, and the graph convolutional network is continuously optimized through reinforcement learning based on the reinforcement learning strategy.
[0007] As a preferred solution of the method for automatic data annotation and classification for artificial intelligence models described in the present invention, wherein: the clustering algorithm is used to perform cluster analysis on the data to form data nodes, the local features of the nodes are extracted through the graph convolutional network, and the discriminator is trained through the mutual information loss function to filter the local features of the nodes, and the dimensionality reduction processing is performed simultaneously, which means that the data is clustered by the K-means clustering algorithm, and each cluster in the clustering result forms a data node, the basic statistics of all data in the data node are extracted as the initial feature vector, all data nodes are connected to form connection edges, the Euclidean distance between data nodes is calculated as the initial edge feature, the Euclidean distance between nodes is less than the set distance threshold, and the nodes are defined as adjacent nodes, and the graph convolutional neural network is used to extract the local features of each data node. ; Get the local features of each data node Then the maximum pooling operation is used to aggregate the local features of all nodes to form the global feature s; Using the depth map information maximization method to build the discriminator ; The discriminator is iteratively trained through the loss function until the loss function value converges; According to the trained discriminator Match the local features of each node, when the discriminator When the output node local feature score is less than the set matching threshold, it is judged as a match failure, the node local feature is screened out from the data node, and the node local feature is regenerated through the graph convolutional network for matching until the match is successful; The principal component analysis method is used to reduce the dimension of the local features of the matched data nodes.
[0008] As a preferred solution of the method for automatic data annotation and classification for artificial intelligence models described in the present invention, wherein: the hierarchical structure of the nodes is represented by a generalized Voronoi diagram, a graph regularization term loss function is defined, a graph convolution operation is used in each layer to update the node feature value to define the number of layers L, the local feature cosine similarity between the first layer data nodes is calculated, the node j with the highest similarity to the node i is selected to merge to generate the next layer of nodes, and the local features of the nodes i and j are synchronously merged to form the local features of the nodes in the next layer , traverse all nodes in the first layer to generate nodes in the next layer and obtain local features of nodes in the next layer; Recursively generate L-layer node hierarchy, connect nodes in each layer to form connection edges, and calculate the data integrity of each layer of nodes by combining the similarities between all nodes in each layer ; The data integrity of each layer of nodes is used as the layer weight, and the local features of each layer of nodes are updated according to the layer weight. .
[0009] As a preferred solution of the data automatic labeling and classification method for artificial intelligence models described in the present invention, wherein: the updating of edge features through the interaction of graph regularization and node features and the updating of node features again through nonlinear transformation through the updated edge features, forming a graph structure based on the updated node features and edge features and performing node classification labeling means that the edge features are updated from the first layer through regularization calculation to obtain a positive update value and update the value in reverse ; The ratio of the forward update value to the reverse update value is used as the edge feature update value , update the node local features again based on the edge feature update value; Taking the updated node local features of the first layer as the benchmark, calculating the similarity between the updated node local features in the remaining layers and the updated node local features in the first layer, selecting the updated node local features with the highest similarity and traversing all layers, weighting all the updated node local features with the highest similarity in the remaining layers and the updated node local features in the first layer to form the final node features, and forming a graph structure based on the final node features and edge feature update values; A convolutional neural network is constructed based on the final node features, and the convolutional neural network is trained based on the training data. The final node features are input into the trained convolutional neural network to obtain node classification, and labels are generated to mark the data in the nodes.
[0010] As a preferred solution of the method for automatic data labeling and classification for artificial intelligence models described in the present invention, the method comprises: collecting structured numerical data based on a database and performing preprocessing operations.
[0011] As a preferred solution of the method for automatic data labeling and classification for artificial intelligence models described in the present invention, the real-time display of node classification labels refers to synchronously displaying the node label tags and node data, and adding classification timestamp data.
[0012] As a preferred solution of the method for automatic data labeling and classification for artificial intelligence models described in the present invention, wherein: the formation of classification records and storage in the database refers to classifying and labeling all nodes to form classification records and storing them in the database, and the database synchronously generates a backup for each classification record and uploads it to the cloud blockchain storage.
[0013] As a preferred solution of the method for automatic data labeling and classification for artificial intelligence models described in the present invention, the continuous reinforcement learning optimization of the graph convolutional network based on the reinforcement learning strategy refers to using the Q-learning algorithm to perform reinforcement learning training on the graph convolutional network and regularly evaluating the accuracy of the graph convolutional network.
[0014] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the method for automatic data labeling and classification for an artificial intelligence model as described in the first aspect of the present invention.
[0015] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for automatic data labeling and classification for an artificial intelligence model as described in the first aspect of the present invention.
[0016] The beneficial effects of the present invention are as follows: the present invention performs cluster analysis on data through a clustering algorithm, and combines the mutual information loss function and the graph regularization term loss function to effectively improve the accuracy of node feature extraction, introduces a generalized Voronoi diagram to represent the hierarchical structure of nodes, and effectively improves the accuracy of node classification through the interactive update features of graph regularization and node features, thereby achieving continuous optimization of real-time data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0018] Figure 1 This is a flow chart of the method for automatically labeling and classifying data for an artificial intelligence model in Example 1. DETAILED DESCRIPTION
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0020] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.
[0022] Example 1, reference Figure 1 , which is the first embodiment of the present invention, and provides a method for automatic data labeling and classification for an artificial intelligence model, comprising the following steps: S1. Collect structured numerical data based on the database, use clustering algorithm to perform cluster analysis on the data to form data nodes, extract local features of nodes through graph convolution network, and train the discriminator through mutual information loss function to filter local features of nodes, and perform dimensionality reduction processing simultaneously; Specifically, a clustering algorithm is used to perform cluster analysis on the data to form data nodes, local features of the nodes are extracted through a graph convolutional network, and the discriminator is trained through a mutual information loss function to filter local features of the nodes. Simultaneous dimensionality reduction processing refers to clustering analysis of the data through a K-means clustering algorithm, and each cluster in the clustering result forms a data node, extracting the basic statistics of all data in the data node as the initial feature vector, including mean, standard deviation, maximum value, minimum value, skewness, kurtosis, etc., connecting all data nodes to form connecting edges, calculating the Euclidean distance between data nodes as the initial edge feature, defining nodes with an Euclidean distance less than a set distance threshold as adjacent nodes, and using a graph convolutional neural network to extract local features of each data node :
[0023] in is the weight matrix, For Node The initial eigenvector of the jth neighboring node, is the set of adjacent nodes of node i, To learn the parameters, is a nonlinear activation function; Get the local features of each data node Then the maximum pooling operation is used to aggregate the local features of all nodes to form the global feature s; Using the depth map information maximization method to build the discriminator :
[0024] in Local features The transpose of is the sigmoid activation function; The discriminator is iteratively trained by the loss function until the loss function value converges. The loss function is:
[0025] Where N and M are the number of positive and negative samples, obtained through training data, is the local feature of the negative sample, generated by the graph erosion function; According to the trained discriminator Match the local features of each node, when the discriminator When the output node local feature score is less than the set matching threshold, it is judged as a match failure, the node local feature is screened out from the data node, and the node local feature is regenerated through the graph convolutional network for matching until the match is successful; The principal component analysis method is used to reduce the dimension of the local features of the matched data nodes.
[0026] The K-means clustering algorithm converts complex and high-dimensional data sets into a more easily processable set of nodes through statistical features (such as mean, standard deviation, etc.), greatly reducing the amount of calculation, while enabling the subsequent graph convolutional network to more effectively capture the intrinsic structure of the data. Through this process, the data nodes can better reflect the similarity of the data, which is convenient for the subsequent model training and optimization. The local features of each data node are extracted through graph convolution operations, which can effectively enhance the diversity and information richness of node representation. This is especially important for processing high-dimensional data, because traditional feature extraction methods often tend to ignore the complex relationship between data. Through the graph convolutional network, the features of the node can reflect its local structure and the influence of surrounding nodes, significantly improving the accuracy of node classification, and further enhancing the representation of the graph by maximizing the mutual information between local features and global features. The beneficial effect of this key step is that it can effectively extract high-quality features from unsupervised data and screen and optimize node features through the discriminator. Compared with traditional supervised learning methods, this method does not require labeled data and can make full use of the structural information of the data, thereby improving the accuracy of automatic labeling and classification. The quality of feature extraction is further optimized by screening the local features of each node and regenerating the node features through the graph convolutional network. The uniqueness of this step is that it screens features through the discriminator and eliminates unmatched features to ensure that the final node features are more accurate. The iterative optimization method enables each round of feature update to improve the representation ability of the node, gradually correct potential errors, and improve the final classification results. Principal component analysis is used to reduce the dimensionality of local node features to reduce feature dimensions and improve model efficiency.
[0027] S2. The hierarchical structure of nodes is represented by a generalized Voronoi diagram, and a graph regularization loss function is defined. Graph convolution operations are used in each layer to update node features. Edge features are updated through the interaction of graph regularization and node features, and node features are updated again through nonlinear transformation through the updated edge features. A graph structure is formed based on the updated node features and edge features, and node classification and labeling is performed. Specifically, the hierarchical structure of nodes is represented by a generalized Voronoi diagram, a graph regularization loss function is defined, and the number of layers L is defined by using graph convolution operations to update node features in each layer. The cosine similarity of local features between the first layer data nodes is calculated, and the node j with the highest similarity to node i is selected to merge to generate the next layer of nodes. The local features of node i and node j are simultaneously merged to form the local features of the next layer of nodes. , traverse all nodes in the first layer to generate nodes in the next layer and obtain local features of nodes in the next layer. The number of nodes in all layers is the same; Recursively generate L-layer node hierarchy, connect nodes in each layer to form connection edges, and calculate the data integrity of each layer of nodes by combining the similarities between all nodes in each layer :
[0028] in is the number of nodes in the lth layer, is the similarity between node i and node j in layer l; The data integrity of each layer of nodes is used as the layer weight, and the local features of each layer of nodes are updated according to the layer weight. :
[0029] in is the set of adjacent nodes of node i in layer l, is the local feature of node j in layer l.
[0030] By using the generalized Voronoi diagram to represent the hierarchical structure of nodes, the difficulty of modeling the relationship between nodes in traditional graph convolutional networks is solved. In traditional methods, node relationships are often defined only based on physical or simple distances, ignoring the complex feature relationships between nodes. The introduction of the generalized Voronoi diagram enables the relationship between nodes to be dynamically adjusted according to the similarity in the feature space, thereby more accurately representing the local relationship between nodes in the multidimensional data space. This innovation effectively improves the expression ability of data nodes in high-dimensional feature space and enhances the information flow in subsequent graph convolution operations, so that the node features of each layer can more accurately capture the structural information of the data. By calculating the cosine similarity of the local features of the nodes in the first layer and selecting the nodes with the highest similarity for merging, the problem of how to deal with the similarity between nodes in a high-dimensional feature space is effectively solved. Node merging based on cosine similarity not only ensures the feature similarity between the merged nodes, but also retains the representative features of each node, which helps to maintain the integrity and consistency of the data in subsequent layers. Compared with simple adjacency relationship construction, node merging based on similarity can better reflect the intrinsic structure of the data and improve the learning efficiency and classification accuracy of the graph convolutional network. Through this recursive structure, the model can extract higher and higher levels of abstract features at each layer and ensure that the nodes are close to each other. The connection can retain its similarity. Unlike the traditional single-level structure, the multi-level recursive structure of the present invention can better capture the complexity and hierarchical relationship of the data, especially when processing large-scale data with multi-level information, it can significantly improve the efficiency of data processing and classification accuracy. In the traditional graph convolutional network, the node feature update of each layer usually does not have a clear weight assignment, which may cause the information of some layers to be ignored or over-learned. By introducing hierarchical weights, the contribution of each layer to the final result can be dynamically adjusted according to the similarity and data integrity of each layer of nodes. This method can ensure that the information of each layer can play an important role in the final node classification, thereby improving The accuracy and robustness of the entire model. In the graph convolutional network, updating edge features through graph regularization terms and updating node features twice in combination with nonlinear transformations are effective means to improve the learning ability of the graph model. The introduction of regularization terms enables the model to automatically balance the learning of local information and global information during training, thereby avoiding the risk of overfitting. At the same time, the updating of edge features and the synchronous optimization of node features help the model to more accurately capture the relationship between nodes and improve the accuracy of node classification. The innovation of this process lies in that through the combination of regularization and nonlinear transformation, the model can flexibly adapt to complex data distribution and improve the expressive power of the graph convolutional network.
[0031] Furthermore, the edge features are updated through the interaction of graph regularization and node features, and the node features are updated again through nonlinear transformation through the updated edge features. The graph structure is formed based on the updated node features and edge features, and the node classification and labeling are performed. The edge features are updated from the first layer through regularization calculation to obtain the positive update value. and update the value in reverse :
[0032]
[0033] in and are the local features of node i and node j in layer l, respectively. is the local feature of node k in layer l, and are the forward update value and reverse update value of the edge features of node i and node j in the l-1th layer, respectively. and are the forward update value and reverse update value of the edge features of node i and node k in the l-1th layer, respectively. is the activation function, l is a positive integer greater than 1. When l-1 is 1, the forward update value and reverse update value of the edge feature of the l-1th layer are equal to the initial edge feature value; The ratio of the forward update value to the reverse update value is used as the edge feature update value , based on the edge feature update value, the node local features are updated again:
[0034] in is the local feature of the node after updating again, For splicing operation; Taking the updated node local features of the first layer as the benchmark, calculating the similarity between the updated node local features in the remaining layers and the updated node local features in the first layer, selecting the updated node local features with the highest similarity and traversing all layers, weighting all the updated node local features with the highest similarity in the remaining layers and the updated node local features in the first layer to form the final node features, and forming a graph structure based on the final node features and edge feature update values; A convolutional neural network is constructed based on the final node features, and the convolutional neural network is trained based on the training data. The final node features are input into the trained convolutional neural network to obtain node classification, and labels are generated to mark the data in the nodes.
[0035] The combination of graph regularization and interactive updating of edge features by node features not only helps the graph convolutional network maintain stability during the node feature update process, but also effectively avoids overfitting. Especially in large-scale data sets, the similarity and structural relationship of nodes are crucial to node classification. By combining graph regularization with interactive updating of edge features, this method can better preserve the global structure of the graph, so that the feature update of each layer can reflect the real relationship between nodes and optimize the process of feature propagation, thereby improving classification accuracy. The operation of updating node features using nonlinear transformation can not only enhance the feature representation ability, but also improve the adaptability of graph convolutional networks in complex data. In traditional graph convolutional networks, linear updates often limit the expressiveness of the model and fail to capture the nonlinear relationship between nodes. By introducing nonlinear transformations, the model can handle more complex data structures and enhance its performance in high-dimensional, nonlinear data. Especially in node classification tasks, it can greatly improve the performance of the model. The ratio of the forward and reverse update values is used as the edge feature update value, which effectively captures the bidirectional information between nodes and enhances the graph convolutional network's understanding of node relationships. The edge feature update in traditional graph convolutional networks often only considers one-way propagation of information, while the introduction of bidirectional updates can better reflect the mutual influence between nodes, improve the stability of the graph structure and the effectiveness of node feature updates, and the node features are updated. The features are updated through the splicing operation, which significantly enhances the representation ability of node features. In the multi-layer graph convolutional network, the nodes must not only retain the local information of the current layer, but also need to fuse information from different levels. The splicing operation can effectively integrate the node features of multiple layers, avoiding the problem of information loss, making the node features more comprehensive and adapting to complex and changeable data structures. The steps of similarity calculation and weighting to form the final node features ensure the effective fusion of features at each layer by combining node features at different levels. This method effectively solves the defect of ignoring the importance of features at different levels in traditional methods, and can ensure that the final node features have good representativeness and distinguishing ability. The weighted fusion process not only retains the baseline information of the first-layer node features, but also can fully reflect the details of high-level node features, making the final node features more accurate and rich, and improving the accuracy of node classification.
[0036] S3, display the node classification annotations in real time, store the classification records in the database, and perform continuous reinforcement learning optimization on the graph convolutional network based on the reinforcement learning strategy; Specifically, displaying node classification annotations in real time means displaying the node annotation labels and node data synchronously, and appending classification timestamp data.
[0037] Furthermore, forming classified records and storing them in a database means classifying and labeling all nodes to form classified records and storing them in the database, and the database synchronously generates a backup for each classified record and uploads it to the cloud blockchain storage.
[0038] Furthermore, continuous reinforcement learning optimization of the graph convolutional network based on the reinforcement learning strategy refers to using the Q-learning algorithm to perform reinforcement learning training on the graph convolutional network and regularly evaluate the accuracy of the graph convolutional network.
[0039] This embodiment also provides a computer device, which is suitable for the case of an automatic data labeling and classification method for an artificial intelligence model, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the automatic data labeling and classification method for an artificial intelligence model as proposed in the above embodiment.
[0040] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0041] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for automatically labeling and classifying data for an artificial intelligence model as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, disk or optical disk.
[0042] In summary, the present invention performs clustering analysis on data through a clustering algorithm, and combines the mutual information loss function and the graph regularization loss function to effectively improve the accuracy of node feature extraction, introduces a generalized Voronoi diagram to represent the hierarchical structure of nodes, and effectively improves the accuracy of node classification through the interactive update of edge features by graph regularization and node features, thereby achieving continuous optimization of real-time data.
[0043] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for automatic data labeling and classification for an artificial intelligence model, characterized by: include, Based on the database, we collect structured numerical data, use clustering algorithms to perform clustering analysis on the data to form data nodes, extract local features of the nodes through graph convolutional networks, and train the discriminator through mutual information loss functions to filter local features of the nodes, and perform dimensionality reduction processing simultaneously; The hierarchical structure of nodes is represented by a generalized Voronoi diagram, and a graph regularization loss function is defined. Graph convolution operations are used in each layer to update node features. Edge features are updated through the interaction of graph regularization and node features, and node features are updated again through nonlinear transformation using the updated edge features. A graph structure is formed based on the updated node features and edge features, and node classification and labeling is performed. The node classification annotations are displayed in real time, the classification records are stored in the database, and the graph convolutional network is continuously optimized through reinforcement learning based on the reinforcement learning strategy.
2. The method for automatic data labeling and classification for an artificial intelligence model according to claim 1, characterized in that: The method uses a clustering algorithm to perform clustering analysis on the data to form data nodes, extracts local features of the nodes through a graph convolutional network, and trains a discriminator through a mutual information loss function to screen local features of the nodes, and simultaneously performs dimensionality reduction processing, which means that the data is clustered by a K-means clustering algorithm, and each cluster in the clustering result forms a data node, extracts the basic statistics of all data in the data node as an initial feature vector, connects all data nodes to form connecting edges, calculates the Euclidean distance between data nodes as an initial edge feature, defines nodes with an Euclidean distance less than a set distance threshold as adjacent nodes, and uses a graph convolutional neural network to extract local features of each data node. ; Get the local features of each data node Then the maximum pooling operation is used to aggregate the local features of all nodes to form the global feature s; Using the depth map information maximization method to build the discriminator ; The discriminator is iteratively trained through the loss function until the loss function value converges; According to the trained discriminator Match the local features of each node, when the discriminator When the output node local feature score is less than the set matching threshold, it is judged as a match failure, the node local feature is screened out from the data node, and the node local feature is regenerated through the graph convolutional network for matching until the match is successful; The principal component analysis method is used to reduce the dimension of the local features of the matched data nodes.
3. The method for automatic data labeling and classification for an artificial intelligence model according to claim 2, characterized in that: The hierarchical structure of the nodes is represented by a generalized Voronoi diagram, a graph regularization term loss function is defined, a graph convolution operation is used in each layer to update the node feature value to define the number of layers L, the local feature cosine similarity between the first layer data nodes is calculated, the node j with the highest similarity to the node i is selected to merge to generate the next layer of nodes, and the local features of the nodes i and j are simultaneously merged to form the local features of the nodes in the next layer , traverse all nodes in the first layer to generate nodes in the next layer and obtain local features of nodes in the next layer; Recursively generate L-layer node hierarchy, connect nodes in each layer to form connection edges, and calculate the data integrity of each layer of nodes by combining the similarities between all nodes in each layer ; The data integrity of each layer of nodes is used as the layer weight, and the local features of each layer of nodes are updated according to the layer weight. .
4. The method for automatic data labeling and classification for an artificial intelligence model according to claim 3, characterized in that: The updating of edge features through the interaction of graph regularization and node features and the updating of node features through the updated edge features through nonlinear transformation, forming a graph structure based on the updated node features and edge features and performing node classification and labeling, refers to updating edge features from the first layer through regularization calculation to obtain a positive update value and update the value in reverse ; The ratio of the forward update value to the reverse update value is used as the edge feature update value , update the node local features again based on the edge feature update value; Taking the updated node local features of the first layer as the benchmark, calculating the similarity between the updated node local features in the remaining layers and the updated node local features in the first layer, selecting the updated node local features with the highest similarity and traversing all layers, weighting all the updated node local features with the highest similarity in the remaining layers and the updated node local features in the first layer to form the final node features, and forming a graph structure based on the final node features and edge feature update values; A convolutional neural network is constructed based on the final node features, and the convolutional neural network is trained based on the training data. The final node features are input into the trained convolutional neural network to obtain node classification, and labels are generated to mark the data in the nodes.
5. The method for automatic data labeling and classification for artificial intelligence models according to claim 4, characterized in that: The structured numerical data is collected based on the database and then preprocessed.
6. The method for automatic data labeling and classification for artificial intelligence models according to claim 5, characterized in that: The real-time display of node classification annotations refers to synchronously displaying the node annotation labels and node data, and appending classification timestamp data.
7. The method for automatic data labeling and classification for artificial intelligence models according to claim 6, characterized in that: The forming of classification records and storing them in the database refers to classifying and labeling all nodes to form classification records and storing them in the database, and the database synchronously generates a backup for each classification record and uploads it to the cloud blockchain storage.
8. The method for automatic data labeling and classification for artificial intelligence models according to claim 7, characterized in that: The continuous reinforcement learning optimization of the graph convolutional network based on the reinforcement learning strategy refers to using the Q-learning algorithm to perform reinforcement learning training on the graph convolutional network and regularly evaluating the accuracy of the graph convolutional network.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for automatic data labeling and classification for an artificial intelligence model described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for automatic data labeling and classification for an artificial intelligence model described in any one of claims 1 to 8 are implemented.