Target identification and model training method and device based on global perception graph convolution
By combining the generation of label adjacency matrix and the global perceptual graph convolutional network, the problem of difficulty in taking into account the accuracy and efficiency of traditional image classification technology in complex scenarios is solved, and the accuracy and generalization ability of target recognition are improved.
Patent Information
- Application Number
- CN202510732936.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Traditional image classification technology is difficult to take into account both classification accuracy and efficiency in complex scenarios, especially in real-time classification tasks. Complex models have a long time to reason, and simple methods have weak generalization capabilities in complex scenarios, resulting in insufficient recognition accuracy and reliability.
By generating the label adjacency matrix, establishing the association relationship between category labels, using the global perceptual graph convolution network to extract local features of the image and perform graph convolution operations, combining visual and semantic information for feature fusion, and training the target recognition model.
It improves the accuracy and generalization performance of the model in complex image scenarios, achieving faster processing speed and stronger generalization capabilities.
Smart Images

Figure CN120259785A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of target recognition, and particularly to a target recognition and model training method and device based on global perception graph convolution. Background Art
[0002] With the development of deep learning technology, the picture classification technology based on graph neural network has been paid more and more attention by people, and its application scope has gradually expanded. Traditional picture classification technology only considers that the whole picture belongs to a certain category, without considering the connection between each part of a picture. This will lead to many limitations and cannot make good use of the useful information brought by the picture.
[0003] Since the rise of deep learning technology, image classification technology has made remarkable progress, with a large number of high-performance classification models emerging and being widely used in industrial production and real life. Currently, the mainstream methods mostly use Convolutional Neural Network (CNN) to extract image features and complete the classification task through a fully connected layer. Although introducing more complex feature extraction structures and more parameters can improve the classification accuracy, in many actual application scenarios, the requirement for accuracy is not extreme, but the detection speed is more emphasized, especially in real-time classification tasks. At this time, although complex models have powerful performance, due to the long inference time, it is difficult to meet the efficiency requirements. And using a simple feature extraction method can achieve a faster classification speed and achieve good results in scenarios with simple structures and little change in targets. However, when the image scene is complex and the target size difference is large, the generalization ability of traditional classification methods is weak, and classification errors are likely to occur, ultimately affecting the accuracy and reliability of recognition. Therefore, it is necessary to provide a target recognition and model training method and device based on global perception graph convolution. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of this application is to provide a target recognition and model training method and device based on global perception graph convolution, which improves the problem that traditional classification network methods are difficult to balance classification accuracy and efficiency in complex scenarios.
[0005] To achieve the above and other related objectives, the present application provides a method for training an object recognition model based on global perception graph convolution. The training method includes: obtaining sample images and corresponding class labels, and statistically calculating the co-occurrence probability of one or two identical class labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one class label, and the class label is used to characterize the class of the target object in the sample image; inputting the sample image into the feature extraction network of the object recognition model to extract features from the sample image to obtain a first image feature, and constructing a graph structure corresponding to the sample image and generating a corresponding first adjacency matrix based on each local feature of the first image feature; wherein, the first image feature includes multiple local features, and each local feature corresponds to a preset region of the sample image, and the feature extraction network is a convolutional neural network; inputting the first image feature and the corresponding first adjacency matrix into the first graph convolution network of the object recognition model for graph convolution operation to obtain a second image feature; inputting the first image feature and the second image feature into the fusion network of the object recognition model for feature fusion to obtain a fusion feature; inputting the fusion feature and the label adjacency matrix into the second graph convolution network of the object recognition model for graph convolution operation to identify the target object in the sample image and obtain the predicted class of the target object; calculating the difference degree between the predicted class and the class label, and adjusting the parameters of the object recognition model based on the difference degree to obtain a trained object recognition model.
[0006] In an embodiment of the present application, obtaining sample images and corresponding class labels, and statistically calculating the co-occurrence probability of one or two identical class labels in all sample images to generate a label adjacency matrix includes: obtaining sample images and corresponding class labels; statistically calculating the co-occurrence probability of one or two identical class labels in all sample images to generate an initial label adjacency matrix; performing thresholding processing on the initial label adjacency matrix to obtain a binary label adjacency matrix; adding self-loops to the diagonal of the binary label adjacency matrix to obtain a self-loop label adjacency matrix; calculating a degree matrix based on the self-loop label adjacency matrix; and performing normalization processing on the self-loop label adjacency matrix according to the degree matrix to obtain a final label adjacency matrix.
[0007] In an embodiment of the present application, the second graph convolutional network includes at least two cascaded graph convolutional layers and a classifier. Inputting the fused feature and the label adjacency matrix into the second graph convolutional network of the target recognition model for graph convolutional operation to identify the target object in the sample image and obtain the predicted category, which includes: inputting the fused feature and the label adjacency matrix into the first graph convolutional layer of the second graph convolutional network, performing feature propagation on the fused feature based on the label adjacency matrix to obtain the graph convolutional feature of the first layer; for each remaining graph convolutional layer: inputting the graph convolutional feature generated by the previous graph convolutional layer and the label adjacency matrix into the current graph convolutional layer, performing feature propagation on the input graph convolutional feature based on the label adjacency matrix to generate the graph convolutional feature of the current layer; inputting the graph convolutional feature generated by the last graph convolutional layer into the classifier, and identifying each target object in the image based on the input graph convolutional feature to obtain the predicted category.
[0008] In an embodiment of the present application, the second graph convolutional network includes a graph convolutional layer and a classifier. Inputting the fused feature and the label adjacency matrix into the second graph convolutional network of the target recognition model for graph convolutional operation to identify the target object in the sample image and obtain the predicted category, which includes: inputting the fused feature and the label adjacency matrix into the graph convolutional layer, performing feature propagation on the fused feature based on the label adjacency matrix to generate the graph convolutional feature; inputting the graph convolutional feature into the classifier, and identifying each target object in the image based on the input graph convolutional feature to obtain the predicted category.
[0009] In an embodiment of the present application, inputting the fused feature and the label adjacency matrix into the graph convolutional layer, performing feature propagation on the fused feature based on the label adjacency matrix to generate the graph convolutional feature, which includes: inputting the label adjacency matrix into the graph convolutional layer, performing matrix power operation on the label adjacency matrix for a preset number of times, and saving the label adjacency matrix at this order after each matrix power operation; adding up the label adjacency matrices at all orders to obtain a fused matrix; performing feature propagation on the fused feature based on the fused matrix to generate the graph convolutional feature.
[0010] In an embodiment of the present application, adding up the label adjacency matrices at all orders to obtain a fused matrix, which includes: performing weighted summation on the label adjacency matrix at the corresponding order based on the attenuation factor to obtain the fused matrix; wherein, the attenuation factor is obtained through initialization and updated during the training process of the target recognition model, and the initial attenuation factor is set to be inversely proportional to the corresponding order.
[0011] In one embodiment of the present application, a target recognition method is further provided. The target recognition method includes: obtaining a target image; inputting the target image into a target recognition model to recognize the target object in the target image and obtain a predicted category; wherein, the target recognition model is trained by the target recognition model training method based on global perception graph convolution in any one of the above.
[0012] In one embodiment of the present application, a target recognition model training system based on global perception graph convolution is further provided. The system includes: a data acquisition module, configured to obtain sample images and corresponding category labels, and count the co-occurrence probability of having one or two identical category labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one category label, and the category label is used to represent the category of the target object in the sample image; a first feature extraction module, configured to input the sample image into the feature extraction network of the target recognition model to extract features from the sample image to obtain a first image feature, and construct a graph structure corresponding to the sample image and generate a corresponding first adjacency matrix based on each local feature of the first image feature; wherein, the first image feature includes multiple local features, and each local feature corresponds to a preset area of the sample image, and the feature extraction network is a convolutional neural network; a second feature extraction module, configured to input the first image feature and the corresponding first adjacency matrix into the first graph convolution network of the target recognition model for graph convolution operation to obtain a second image feature; a fusion module, configured to input the first image feature and the second image feature into the fusion network of the target recognition model for feature fusion to obtain a fusion feature; a graph convolution module, configured to input the fusion feature and the label adjacency matrix into the second graph convolution network of the target recognition model for graph convolution operation to recognize the target object in the sample image and obtain the predicted category of the target object; a parameter update module, configured to calculate the difference degree between the predicted category and the category label, and adjust the parameters of the target recognition model based on the difference degree to obtain a trained target recognition model.
[0013] In one embodiment of the present application, an electronic device is further provided, including: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to implement the target recognition model training method or the target recognition method based on global perception graph convolution in any one of the above.
[0014] In one embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, and when the computer program is executed by a processor of a computer, enable the computer to execute the target recognition model training method or the target recognition method based on global perception graph convolution in any one of the above.
[0015] As described above, a method and apparatus for object recognition and model training based on global perception graph convolution of the present application have the following beneficial effects: By statistically generating a label adjacency matrix based on the co-occurrence probability between various category labels in the sample images, the correlation relationship between different category labels is established, so as to improve the semantic recognition ability of the model in the case of multi-label classification or uneven category distribution. By using a convolutional neural network to extract multiple local features of the image and construct a graph structure, and using the first graph convolutional network for structured propagation, the correlation perception ability of the model for each region in the image is improved. After fusing the first image feature representing the visual feature extracted by the convolutional neural network and the second image feature representing the region correlation feature extracted by the first graph convolutional network, using the second graph convolutional network, under the guidance of the label adjacency matrix, the deep fusion of the visual feature and the label semantic relationship is realized. By combining the visual information and semantic information of the image, the discrimination ability of the model for the target categories in complex images is improved, making the final prediction result more accurate, and the label adjacency matrix is only constructed in the training stage and can be directly reused in the inference stage, so that the model has stronger generalization performance and processing speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. 6 is a schematic flowchart of a method for training an object recognition model based on global perception graph convolution provided by an embodiment of the present application; Figure 2 FIG. 9 is a schematic flowchart of an object recognition method provided by an embodiment of the present application; Figure 3 FIG. 12 shows a structural block diagram of a system for training an object recognition model based on global perception graph convolution provided by an embodiment of the present application; Figure 4 FIG. 15 shows a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following specific examples illustrate the embodiments of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0018] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present application. Therefore, only the components related to the present application are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0019] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0020] The present application provides a method for training an object recognition model based on global perception graph convolution. The present application uses algorithms in the fields of computer vision and deep learning, including image enhancement, graph convolution network construction, feature extraction, etc., to implement an object classification and recognition method based on an infinite-nearest neighbor graph convolution network, improving the capabilities of traditional classification networks and solving the problem that traditional classification network methods are difficult to effectively classify in complex scenarios. The overall feature representation of the image is obtained by extracting features through a traditional convolutional neural network and denoted as the first image feature. The first adjacency matrix required is constructed using the information in the image to establish the connection between each part of the image. The constructed first adjacency matrix and the first image feature are input into the first graph convolution network for calculation to obtain the second image feature. The first image feature and the second image feature are fused to obtain the final feature information, and the category to which the target object appears in the image is obtained using the second graph convolution network based on the obtained feature information. The present application is based on the gradient backpropagation algorithm in deep learning. According to the final loss function of the model, the loss of each iteration is automatically calculated during training. Through the chain rule of differentiation, the update gradients of all learnable parameters in the model are calculated, thereby completing the update of the model parameters and realizing an end-to-end training process, avoiding manual intervention and manual calculation of the parameters of the feature extractor and classifier, improving the usability of the system, and enabling the learned model parameters to better adapt to the classification task. Among them, when updating the parameters through gradient backpropagation, the model parameters can be updated more directly and efficiently, avoiding gradient disappearance. According to the principles of digital image processing, the present application performs various data augmentations on the training images, including image flipping, color space conversion, image scaling, etc., improving the utilization rate of the training images, increasing the diversity of samples, reducing the need for data annotation to a certain extent, and enhancing the robustness and generalization ability of the model.
[0021] As Figure 1 shown, the method for training an object recognition model based on global perception graph convolution includes the following steps: S11. Obtain a sample image and its corresponding class label, and count the co-occurrence probability of one or two identical class labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one class label, and the class label is used to characterize the class of the target object in the sample image.
[0022] When the target recognition model needs to be trained, first, it is necessary to obtain sample images containing the target object and label each obtained sample image to obtain its corresponding class label. Herein, the class label is used to characterize the class of the target object included in the corresponding sample image. It can be understood that a sample image can correspond to one class label or two or more class labels to support single-label classification or multi-label classification. Those skilled in the art can adaptively select the label format based on the classification requirements, which are not limited herein. Divide all the obtained sample images into a training set, a test set, and a validation set according to a preset ratio. Exemplarily, the training set includes M sample images to be trained as , where represents the i-th sample image, and each sample image corresponds to a class label combination, denoted as { }, and n , N being the total number of classes.
[0023] Count the co-occurrence probability of one or two identical class labels in all sample images. Herein, the co-occurrence probability refers to the number of times any two class labels and appear in the same sample image in the training set, and calculate the conditional probability P(i|j) that class label appears on the premise that class label appears, as shown in formula (1):
[0024] Generate a label adjacency matrix according to the conditional probability by calculating the conditional probability between all class labels in the training set. Herein, the label adjacency matrix is used to characterize the association relationship between class labels. It can be understood that the class labels and can be different class labels (i.e., i is not equal to j) or the same class label (i.e., i is equal to j), which are not limited herein.
[0025] In an optional embodiment of the present application, step S11 includes steps S111 to S116: S111. Obtain a sample image and its corresponding class label.
[0026] After obtaining the original sample image, the image can be normalized according to the preset pixel mean and standard deviation. Since the image is in RGB format, each pixel has three pixel values corresponding to the three color channels. Therefore, each color channel has a pixel mean and standard deviation. The sample image is normalized at the pixel level in each color channel according to formula (2): (2) where is the normalized sample image, is the preset pixel mean, is the preset pixel standard deviation, is the original sample image.
[0027] In addition, in order to make the image meet the input requirements of the model, the sample image will be uniformly scaled to a preset size (such as 320 320). It should be noted that after the sample image is scaled, the annotation positions of the target objects in the sample image also need to be adjusted accordingly, otherwise there will be a mismatch. The preset size can also be 512 512 or 640 640. Since higher-resolution sample images can improve the recognition accuracy of the network, but will reduce the classification speed of the network, those skilled in the art can adaptively set the size of the preset size based on actual needs and will not be limited here as long as the sample images can be unified to the same size.
[0028] Furthermore, during the model training stage, data augmentation will also be performed on the sample image. Among them, the methods of data augmentation include but are not limited to randomly changing the brightness and saturation of the sample image and randomly horizontally flipping the sample image. The random changes in the brightness and saturation of the sample image are respectively performed in the RGB color space and after converting the sample image to the HSV space. The random horizontal flipping of the sample image uses horizontal random flipping. The flipping and random cropping of the sample image need to consider the annotation positions of the objects in the sample image and be adjusted synchronously. It can be understood that during the training, validation, and testing stages of the target recognition model, the data input to the target recognition model in each batch can include multiple sample images. Since each sample image has the same processing process, for the convenience of description, this application takes any one sample image as an example for illustration, and those skilled in the art can understand and generalize it to the processing operations of batch sample images accordingly.
[0029] S112. Statistically calculate the co-occurrence probability of having one or two identical class labels in all sample images to generate an initial label adjacency matrix.
[0030] Assume that the class label set corresponding to the training set is \(C = { \}\), where \(N\) is the total number of classes. For any two class labels and , count the number of sample images in the training set that contain both of these two class labels, denoted as , and the number of sample images that contain the class label (regardless of whether it includes ), denoted as . Then the conditional probability \(P(i|j)\) of the class label under the condition that the class label appears is \(P(i|j)=\frac{ }{ }\). By calculating \(P(i|j)\) for all combinations of \(i \), \(j \), an initial label adjacency matrix = is obtained based on the conditional probability, where = \(P(i|j)\). The label adjacency matrix reflects the semantic association degree between different class labels.
[0031] S113. Threshold the initial label adjacency matrix to obtain a binary label adjacency matrix.
[0032] To improve the sparsity of the label graph structure and reduce noise interference, a probability threshold \(t\) can be preset. Compare all elements in the initial label adjacency matrix with this probability threshold \(t\), and generate a binary label adjacency matrix = through thresholding, as shown in formula (3): (3) Regard the class labels and as the corresponding nodes \(i\) and \(j\). Among them, = 1 indicates that there is a connection between nodes \(i\) and \(j\). Through thresholding, the association relationships of class labels with relatively strong semantic correlations can be effectively retained, and the association relationships with relatively low correlations can be removed, so that the subsequent first graph convolutional network can focus more on the main classes, which not only improves the processing speed of the model but also enhances the generalization performance of the model.
[0033] S114. Add self-loops to the diagonal of the binary label adjacency matrix to obtain a self-loop label adjacency matrix.
[0034] To ensure that each class label can retain its own feature information during the graph convolution process, add the identity matrix to the diagonal of the binary label adjacency matrix to obtain a self-loop label adjacency matrix , as shown in formula (4): (4) Through this self-loop connection, the feature information of the category label can be retained during the graph convolution propagation process, thereby improving the situation where the feature information of itself decays due to over-reliance on the neighbor category label.
[0035] S115. Calculate the degree matrix based on the self-loop label adjacency matrix.
[0036] The degree matrix D is a diagonal matrix, and its diagonal represents the sum of the connection edges between node i and other nodes, and can be obtained through formula (5): (5) Among them, represents the connection relationship between node i and node j, taking values of 1 or 0, and N is the total number of categories. And by normalizing the degree matrix, the normalized degree matrix can be obtained, where the normalization method is as shown in formula (6): (6) S116. Normalize the self-loop label adjacency matrix according to the degree matrix to obtain the final label adjacency matrix.
[0037] As shown in formula (7), according to the normalized degree matrix normalize the self-loop label adjacency matrix to obtain the final label adjacency matrix L: L = (7) Through this normalization operation, the aggregation of the feature of each category label node to its neighbor nodes during the graph convolution propagation process can have the same scale weight, effectively avoiding the feature offset or over-smoothing problem caused by the difference in the number of node connections.
[0038] S12. Input the sample image into the feature extraction network of the target recognition model, extract the features of the sample image to obtain the first image feature, and based on each local feature of the first image feature, construct a graph structure corresponding to the sample image and generate a corresponding first adjacency matrix; wherein, the first image feature includes multiple local features, and each local feature corresponds to a preset area of the sample image.
[0039] Input the sample image into the feature extraction network. Through multiple layers of convolution and feature mapping, obtain the first image feature with a spatial structure, where the first image feature is used to reflect visual features such as edges and textures in each region of the sample image. Specifically, the first image feature includes multiple local features, and each local feature corresponds to a preset spatial region in the sample image (such as the region covered by the convolution receptive field), and is used to characterize information such as the texture, edge, or object composition of this region. The multiple local features are arranged in sequence according to their positions in the sample image to form the first image feature. Take each local feature as a graph node, establish edge connections between the nodes, and form the graph structure corresponding to the sample image. Generate the first adjacency matrix based on this graph structure, which is used to describe the connection relationship between the nodes in the graph structure. Among them, the edges between the nodes can be constructed according to the feature similarity or spatial adjacency relationship between the local features. It can be understood that the feature extraction network can be any convolutional neural network capable of realizing image feature extraction, including but not limited to ResNet series networks or VGG series networks, where the ResNet network includes but not limited to ResNet50, ResNet101, and ResNet152; the VGG network includes but not limited to VGG16 and VGG19. Exemplarily, the ResNet50 network can be used as the feature extraction network, and its pre-trained model comes from the classification model on the ImageNet dataset, and the learning rate of the first two residual structures of ResNet50 is set to 0 so that it does not participate in the training, which can reduce the overfitting risk during the model training process.
[0040] S13. Input the first image feature and the corresponding first adjacency matrix into the first graph convolutional network of the target recognition model for graph convolutional operation to obtain the second image feature.
[0041] Input the first image feature and the corresponding first adjacency matrix into the first graph convolutional network of the target recognition model together. Obtain the second image feature through graph convolutional operation. Since the second image feature can characterize the relationship between each region in the sample image, it can better handle some complex scenarios, thereby improving the recognition ability of the model.
[0042] S14. Input the first image feature and the second image feature into the fusion network of the target recognition model for feature fusion to obtain the fusion feature.
[0043] Since the first image feature is a visual representation extracted by a convolutional neural network, which can characterize the texture and edge features of the image, and the second image feature is a structure-enhanced semantic feature generated by graph structure propagation, which can characterize the association information between local regions in the image. The fusion network can deeply fuse the first image feature and the second image feature through methods such as feature splicing, weighted summation, and attention mechanism, so as to obtain a fusion feature containing spatial perception ability and graph structure information. Exemplarily, the first image feature and the second image feature are fused by dot multiplication to obtain a fusion feature.
[0044] S15. Input the fusion feature and the label adjacency matrix into the second graph convolutional network of the target recognition model for graph convolutional operation to identify the target object in the sample image and obtain the predicted category.
[0045] Input the fusion feature and the label adjacency matrix into the second graph convolutional network of the target recognition model. Through graph convolutional operation, the target recognition model can effectively fuse the image content information and the label semantic structure information, thereby greatly improving the recognition and detection ability of the target category in the image and obtaining the predicted category.
[0046] In an optional embodiment of the present application, the second graph convolutional network includes at least two cascaded graph convolutional layers and a classifier. At this time, the generation of the predicted category includes the following process: First, input the fusion feature and the label adjacency matrix into the first graph convolutional layer of the second graph convolutional network, and perform feature propagation on the fusion feature based on the label adjacency matrix to obtain the graph convolutional feature of the first layer.
[0047] The second graph convolutional network includes at least two cascaded graph convolutional layers and a classifier. The number of graph convolutional layers is not limited. It can be understood that those skilled in the art can adaptively select the number of graph convolutional layers based on actual task requirements. To improve the processing speed, optionally, there are two graph convolutional layers, forming a semantic propagation module with multi-layer graph structure perception. Specifically, input the fusion feature and the label adjacency matrix into the first graph convolutional layer of the second graph convolutional network. The graph convolutional layer performs the first round of graph structure perception feature propagation operation on the fusion feature based on the connection relationship of various category label nodes in the label adjacency matrix, so as to introduce the context association of other category label semantics while retaining the original category information, and generate the graph convolutional feature of the first layer , as shown in formula (8): (8) Among them, is the fusion feature, are the learnable parameters in the first-layer graph convolution, which will change during model training, It is a non-linear activation function. Optionally, in order to unify the dimensions of the input features, the dimensions of the fused features are also unified to a preset length (such as 200 dimensions). The graph convolution operation first completes the propagation of features within the entire graph in the graph structure, and then completes the feature update through linear transformation and non-linear activation, thereby generating the first-layer graph convolution features.
[0048] Then for each remaining graph convolution layer: The graph convolution features generated by the previous graph convolution layer and the label adjacency matrix are input into the current graph convolution layer. Based on the label adjacency matrix, the input graph convolution features are propagated to generate the graph convolution features of the current layer.
[0049] For each remaining graph convolution layer in the second graph convolution network except the first graph convolution layer, the following process is executed: The graph convolution features output by the previous graph convolution layer and the label adjacency matrix are input into the current graph convolution layer. In this graph convolution layer, still based on the connection relationship between the class labels defined by the label adjacency matrix, the input graph convolution features are propagated, and the graph convolution features of the current layer are output. , as shown in formula (9): (9) Among them, is the graph convolution feature generated by the previous graph convolution layer, are the learnable parameters of the current layer (i.e., the i-th layer) graph convolution layer. Through this multi-level cascaded graph convolution propagation mechanism, the target recognition model can accumulate the high-order dependency relationships between label semantics layer by layer, thereby enhancing the feature expression ability of the model and improving the recognition accuracy.
[0050] Finally, the graph convolution features generated by the last graph convolution layer are input into the classifier, and each target object in the image is recognized based on the input graph convolution features to obtain the predicted category.
[0051] The graph convolution features output by the last graph convolution layer are input into the classifier in the target recognition model. Based on the representation ability of the graph convolution features in the label dimension, the classifier outputs the predicted values corresponding to each category, and can obtain the target objects appearing in the image and generate the predicted category according to the preset prediction threshold or sorting rule. Among them, the predicted category is to identify the target objects from the image and label the categories of the target objects. The classifier of the present application can be a multi-layer perceptron, a fully connected layer plus an activation function, or other network structures for multi-label classification, etc., which are not specifically limited as long as they can perform non-linear transformation and category score calculation on the graph convolution features of various categories.
[0052] Considering the Laplacian smoothing phenomenon in the graph, when the number of graph convolution layers is too large, the node features will be over-smoothed, resulting in nodes becoming difficult to distinguish. To address the above issues, in another optional embodiment of the present application, the second graph convolution network includes a graph convolution layer and a classifier. At this time, the generation process of the predicted category includes steps S151 to S152 (not shown in the figure): S151. Input the fused feature and the label adjacency matrix into the graph convolution layer, and perform feature propagation on the fused feature based on the label adjacency matrix to generate graph convolution features.
[0053] The fused feature and the label adjacency matrix are respectively used as the input feature and the graph structure, and input into the graph convolution layer. The graph convolution layer performs a feature propagation operation on the fused feature based on the label adjacency matrix to generate graph convolution features. These graph convolution features fuse visual semantic information and label structure information and are used as the classification basis for the subsequent classifier.
[0054] In order to expand the receptive field of the second graph convolution network, in an optional embodiment of the present application, the generation process of the graph convolution features includes steps S1511 to S1513 (not shown in the figure): S1511. Input the label adjacency matrix into the graph convolution layer, perform matrix power operations on the label adjacency matrix for a preset number of times, and save the label adjacency matrix at this order after each matrix power operation.
[0055] To expand the receptive field of the second graph convolution and enable each category node to aggregate more semantic information, matrix power operations are performed on the label adjacency matrix for a preset number of times. Specifically, let the label adjacency matrix be L and the power be K, then calculate successively , …, , which represents the information transfer relationship from directly adjacent nodes to K-order neighbor nodes. After each matrix power operation is completed, the label adjacency matrix at this order is saved for subsequent construction of a multi-order information fusion structure. In this way, the model can integrate label semantic information in different-order neighborhoods in one propagation, thereby enhancing the graph structure's ability to model long-distance category associations and contributing to improving the recognition accuracy in complex scenarios. Preferably, an infinite-nearest neighbor matrix (i.e., K is infinite) is used to expand the receptive field of the second graph convolution and aggregate information of more nodes.
[0056] S1512. Accumulate all the label adjacency matrices at all orders to obtain a fusion matrix.
[0057] Accumulate all the label adjacency matrices at different orders element by element to obtain a fusion matrix , as shown in formula (10): (10) Among them, is the k-th order adjacency matrix. This fusion matrix integrates the propagation relationships between label nodes at different orders in terms of structure, enabling the graph convolutional network to simultaneously perceive the multi-order neighborhood information of label nodes in a single feature propagation. It should be noted that at this time, K can take infinity, that is = .
[0058] Considering that the high-level information will cause interference to the information due to the existence of smoothing, to improve the above problems, in an optional embodiment of the present application, the fusion matrix is implemented through the following process: based on the attenuation factor, the label adjacency matrices at the corresponding orders are weighted and summed to obtain the fusion matrix; among them, the attenuation factor is obtained through initialization and updated during the training process of the target recognition model, and the initial attenuation factor is set to be inversely proportional to the corresponding order.
[0059] Optionally, to control the influence intensity of the label adjacency matrices at different orders on the final feature propagation when constructing the fusion matrix, the present application also adopts an attenuation mechanism to weight each order of adjacency matrix, and by introducing an attenuation factor corresponding to the order, each order of label adjacency matrix is weighted, and the obtained fusion matrix is shown in formula (11): (11) Among them, is the attenuation factor corresponding to the order k. For the attenuation factor, its initial value is set to be inversely proportional to the order to reduce the interference of the high-order adjacency matrix on the fusion matrix. During the model training, the attenuation factor is continuously updated as a learnable parameter, so as to realize the dynamic adjustment of the propagation range of different orders of labels.
[0060] S1513. Perform feature propagation on the fused features based on the fusion matrix to generate graph convolutional features.
[0061] To enhance the expression ability of node features, a fully connected layer is introduced after the graph convolution operation to further transform and enhance the output features. Specifically, after performing a graph convolution operation on the fused features and the fusion matrix, a non-linear activation function is introduced for processing, and feature mapping is performed through a multi-layer perceptron MLP. Finally, the obtained graph convolutional features are shown in formula (12): (12) Among them, H is the graph convolutional feature, X is the fused feature, W is the learnable parameter of the graph convolutional layer, is the non-linear activation function, such as the ReLU activation function, etc., is the fusion matrix, which can be obtained through any of the above methods, is the partial derivative of the label adjacency matrix L. During model training, the parameters W of the graph convolutional layer are updated and optimized with continuous training iterations, thereby gradually improving the model's ability to express and discriminate the features of target objects in images. S152. Input the graph convolutional features into a classifier, and identify each target object in the image based on the input graph convolutional features to obtain predicted categories.
[0062] Input the graph convolutional features as the final semantic representation into the classifier in the target recognition model. Since the graph convolutional features fuse the visual features of the original image and the structural semantic information between class labels and can fully express the possibility of each class existing in the image, a more accurate predicted category can be obtained through the classifier.
[0063] S16. Calculate the difference degree between the predicted category and the class label, and adjust the parameters of the target recognition model based on the difference degree to obtain a trained target recognition model.
[0064] The predicted category is output by the graph convolutional network and its subsequent classifier, representing the predicted value of whether each class target exists in the image. By comparing this prediction result with the true class label corresponding to the training sample, the loss value is calculated as a measure of the model performance. Specifically, by calculating the cross-entropy loss between the predicted category and the corresponding true class label, the recognition error of the target recognition model can be measured. By calculating the regularization term loss of the target recognition model, the problem of overfitting of the model can be prevented, and the generalization ability of the model can be improved. Specifically, the cross-entropy loss and the regularization term loss are calculated as shown in formulas (13) and (14): (13) (14) where N is the total number of categories, is the output value corresponding to the c-th category predicted by the model, is the true class label of the corresponding sample image in the c-th category, is the non-linear activation function, is the i-th row vector in the classifier parameter matrix, corresponding to the weight vector of the i-th category, is the matrix norm. Add the above cross-entropy loss and the regularization term loss weightedly to obtain the total difference degree, and update the parameters of the entire target recognition model based on the total difference degree. Finally, a trained target recognition model is obtained.
[0065] As Figure 2 shown, this application also provides a target recognition method. The target recognition method includes the following process: S21. Obtain a target image; S22. Input the target image into a target recognition model to identify the target object in the target image and obtain a predicted category; wherein, the target recognition model is trained by the target recognition model training method based on global perception graph convolution according to any one of the above.
[0066] When it is necessary to identify the target object in the target image, the target image can be input into the target recognition model, the target image is subjected to feature extraction through a feature extraction network to obtain a first image feature, and a graph structure corresponding to the sample image is constructed and a corresponding first adjacency matrix is generated based on each local feature of the first image feature. The first image feature and the corresponding first adjacency matrices are input into the first graph convolution network of the target recognition model for graph convolution operation to obtain a second image feature. The first image feature and the second image feature are input into the fusion network of the target recognition model for feature fusion to obtain a fusion feature. The fusion feature and the label adjacency matrix are input into the second graph convolution network of the target recognition model for graph convolution operation to identify the target object in the sample image and obtain the predicted category of the target object.
[0067] Such as Figure 3As shown in the figure, the training system 300 for the object recognition model based on global perception graph convolution includes: a data acquisition module 310, a first feature extraction module 320, a second feature extraction module 330, a fusion module 340, a graph convolution module 350, and a parameter update module 360. The data acquisition module 310 is configured to acquire sample images and corresponding class labels, and calculate the co-occurrence probability of one or two identical class labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one class label, and the class label is used to characterize the class of the target object in the sample image. The first feature extraction module 320 is configured to input the sample image into the feature extraction network of the object recognition model, extract features from the sample image to obtain first image features, and construct a graph structure corresponding to the sample image and generate a corresponding first adjacency matrix based on each local feature of the first image features; wherein, the first image features include multiple local features, each local feature corresponds to a preset area of the sample image, and the feature extraction network is a convolutional neural network. The second feature extraction module 330 is configured to input the first image features and the corresponding first adjacency matrix into the first graph convolution network of the object recognition model for graph convolution operations to obtain second image features. The fusion module 340 is configured to input the first image features and the second image features into the fusion network of the object recognition model for feature fusion to obtain fusion features. The graph convolution module 350 is configured to input the fusion features and the label adjacency matrix into the second graph convolution network of the object recognition model for graph convolution operations to identify the target object in the sample image and obtain the predicted class of the target object. The parameter update module 360 is configured to calculate the difference degree between the predicted class and the class label, and adjust the parameters of the object recognition model based on the difference degree to obtain a trained object recognition model.
[0068] For the specific limitations of the training system for the object recognition model based on the graph convolution network, reference can be made to the limitations of the training method for the object recognition model based on the graph convolution network in the above text, which will not be elaborated here. Each module in the above training system for the object recognition model based on the graph convolution network can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in a hardware format or independent of it, or stored in the memory of the computer device in a software format, so that the processor can call the corresponding operations of the above modules.
[0069] It should be noted that in order to highlight the innovative part of this application, modules that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.
[0070] Such as Figure 4As shown, the electronic device 4 may include a memory 41, a processor 42, and a bus. It may also include a computer program stored in the memory 41 and executable on the processor 42, such as a target recognition model training program based on global perception graph convolution.
[0071] Among them, the memory 41 includes at least one type of readable storage medium, which includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 41 may be an internal storage unit of the electronic device 4, such as the mobile hard disk of the electronic device 4. In some other embodiments, the memory 41 may also be an external storage device of the electronic device 4, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 4. Further, the memory 41 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 41 can be used not only to store application software installed on the electronic device 4 and various types of data, such as the code for training the target recognition model based on global perception graph convolution, etc., but also to temporarily store data that has been output or will be output.
[0072] In some embodiments, the processor 42 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 42 is the control core (Control Unit) of the electronic device 4, connecting various components of the entire electronic device 4 through various interfaces and lines. By running or executing programs or modules stored in the memory 41 (such as the target recognition model training program based on global perception graph convolution, etc.), and by calling the data stored in the memory 41, it executes various functions of the electronic device 4 and processes data.
[0073] The processor 42 executes the operating system of the electronic device 4 and various installed application programs. The processor 42 executes the application programs to implement the steps in the above-mentioned target recognition model training method based on global perception graph convolution.
[0074] Exemplarily, a computer program can be divided into one or more modules. One or more modules are stored in the memory 41 and executed by the processor 42 to complete the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 4. For example, the computer program can be divided into a data acquisition module 310, a first feature extraction module 320, a second feature extraction module 330, a fusion module 340, a graph convolution module 350, and a parameter update module 360.
[0075] The units implemented in the form of the above software function modules can be stored in a computer-readable storage medium. The computer-readable storage medium can be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the method for training an object recognition model based on global perception graph convolution in various embodiments of the present application.
[0076] In summary, for a method and device for object recognition and model training based on global perception graph convolution disclosed in the present application, the present application uses a graph convolution network method based on deep learning to implement object recognition (such as the recognition of truck types), so as to satisfy the chain rule of differentiation, enabling the gradient to be propagated normally, achieving end-to-end training, and improving the performance of the network. By fusing the first image features extracted by the feature extraction network and the second image features extracted by the first graph convolution network, and then using the second graph convolution network to conduct the final classification and recognition research, the present application can be better extended to complex scenarios, making the model have better scalability. Therefore, the present application effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0077] The above embodiments are only illustrative of the principles and effects of the present application and are not used to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in the present application should still be covered by the claims of the present application.
Claims
1. A method for training an object recognition model based on global perception graph convolution, characterized in that, The training method includes: Obtain sample images and corresponding class labels, and count the co-occurrence probabilities of one or two identical class labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one class label, and the class label is used to characterize the class of the target object in the sample image; Input the sample image into the feature extraction network of the target recognition model, extract features from the sample image to obtain a first image feature, and construct a graph structure corresponding to the sample image and generate a corresponding first adjacency matrix based on each local feature of the first image feature; wherein, the first image feature includes multiple local features, each local feature corresponds to a preset area of the sample image, and the feature extraction network is a convolutional neural network; Input the first image feature and the corresponding first adjacency matrix into the first graph convolutional network of the target recognition model for graph convolution operation to obtain a second image feature; Input the first image feature and the second image feature into the fusion network of the target recognition model for feature fusion to obtain a fusion feature; Input the fusion feature and the label adjacency matrix into the second graph convolutional network of the target recognition model for graph convolution operation to identify the target object in the sample image and obtain the predicted class of the target object; Calculate the difference degree between the predicted class and the class label, and adjust the parameters of the target recognition model based on the difference degree to obtain a trained target recognition model.
2. The method for training an object recognition model based on global perception graph convolution according to claim 1, wherein The obtaining the sample images and corresponding class labels, and counting the co-occurrence probabilities of one or two identical class labels in all sample images to generate a label adjacency matrix includes: Obtain the sample images and corresponding class labels; Count the co-occurrence probabilities of one or two identical class labels in all sample images to generate an initial label adjacency matrix; Perform thresholding on the initial label adjacency matrix to obtain a binary label adjacency matrix; Add self-loops to the diagonal of the binary label adjacency matrix to obtain a self-loop label adjacency matrix; Calculate the degree matrix based on the self-loop label adjacency matrix; Normalize the self-loop label adjacency matrix according to the degree matrix to obtain the final label adjacency matrix.
3. The method for training an object recognition model based on global perception graph convolution according to claim 1, wherein The second graph convolutional network includes at least two cascaded graph convolutional layers and a classifier. The inputting the fusion feature and the label adjacency matrix into the second graph convolutional network of the target recognition model for graph convolution operation to identify the target object in the sample image and obtain the predicted class of the target object includes: Input the fusion feature and the label adjacency matrix into the first graph convolutional layer of the second graph convolutional network, and perform feature propagation on the fusion feature based on the label adjacency matrix to obtain the graph convolutional feature of the first layer; For each of the remaining graph convolutional layers: input the graph convolutional feature generated by the previous graph convolutional layer and the label adjacency matrix into the current graph convolutional layer, and perform feature propagation on the input graph convolutional feature based on the label adjacency matrix to generate the graph convolutional feature of the current layer; Input the graph convolution features generated by the last graph convolution layer into the classifier, and identify each target object in the sample image based on the input graph convolution features to obtain the predicted categories.
4. The method for training an object recognition model based on global perception graph convolution according to claim 1, characterized in that, The second graph convolution network includes a graph convolution layer and a classifier. Input the fusion features and the label adjacency matrix into the second graph convolution network of the target recognition model for graph convolution operations to identify the target objects in the sample image and obtain the predicted categories of the target objects, including: Input the fusion features and the label adjacency matrix into the graph convolution layer, and perform feature propagation on the fusion features based on the label adjacency matrix to generate graph convolution features; Input the graph convolution features into the classifier, and identify each target object in the sample image based on the input graph convolution features to obtain the predicted categories.
5. The training method of the object recognition model based on global perception graph convolution according to claim 4, characterized in that Input the fusion features and the label adjacency matrix into the graph convolution layer, and perform feature propagation on the fusion features based on the label adjacency matrix to generate graph convolution features, including: Input the label adjacency matrix into the graph convolution layer, perform matrix power operations on the label adjacency matrix for a preset number of times, and save the label adjacency matrix at this order after each matrix power operation; Accumulate the label adjacency matrices at all orders to obtain a fusion matrix; Perform feature propagation on the fusion features based on the fusion matrix to generate graph convolution features.
6. The method for training an object recognition model based on global perception graph convolution according to claim 5, wherein The step of accumulating the label adjacency matrices at all orders to obtain a fusion matrix includes: performing weighted summation on the label adjacency matrix at the corresponding order based on an attenuation factor to obtain a fusion matrix; wherein, the attenuation factor is obtained through initialization and updated during the training process of the target recognition model, and the initial attenuation factor is set to be inversely proportional to the corresponding order.
7. A target recognition method, characterized in that, The target recognition method includes: Obtain a target image; Input the target image into the target recognition model to identify the target objects in the target image and obtain the predicted categories; wherein, the target recognition model is trained by the target recognition model training method based on global perception graph convolution according to any one of claims 1-6.
8. A training system for an object recognition model based on global perception graph convolution, characterized in that, The system includes: A data acquisition module, configured to acquire a sample image and a corresponding category label, and count the co-occurrence probability of having one or two identical category labels in all sample images to generate a label adjacency matrix; wherein, each sample image corresponds to at least one category label, and the category label is used to characterize the category of the target object in the sample image; A first feature extraction module, configured to input the sample image into the feature extraction network of the target recognition model, extract features from the sample image to obtain first image features, and construct a graph structure corresponding to the sample image and generate a corresponding first adjacency matrix based on each local feature of the first image features; wherein, the first image features include multiple local features, each local feature corresponds to a preset area of the sample image, and the feature extraction network is a convolutional neural network; A second feature extraction module, configured to input the first image feature and the corresponding first adjacency matrix into a first graph convolutional network of the target recognition model to perform graph convolutional operations, so as to obtain a second image feature; A fusion module, configured to input the first image feature and the second image feature into a fusion network of the target recognition model to perform feature fusion, so as to obtain a fused feature; A graph convolutional module, configured to input the fused feature and the label adjacency matrix into a second graph convolutional network of the target recognition model to perform graph convolutional operations, identify a target object in the sample image, and obtain a predicted category of the target object; A parameter update module, configured to calculate a difference degree between the predicted category and a category label, and adjust parameters of the target recognition model based on the difference degree, so as to obtain a trained target recognition model.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, which when executed by the one or more processors, cause the electronic device to implement the method for training a target recognition model based on global perception graph convolution according to any one of claims 1 to 6 or the target recognition method according to claim 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by a processor of a computer, causes the computer to execute the method for training a target recognition model based on global perception graph convolution according to any one of claims 1 to 6 or the target recognition method according to claim 7.
Citation Information
Patent Citations
Hash code generation method and system for multi-label image
CN112395438A
Efficient width graph convolutional neural network model and training method thereof
CN112633482A
Image processing method and device and electronic equipment
CN114283109A
Multi-label image classification method fusing strong correlation between labels
CN114648635A
Safety helmet wearing identification model training method, identification method and storage medium
CN115035472A
Cited By
Vaccine clinical test quality management system optimization method and system
CN120452645A
Training method and system for multi-modal large model in automobile field
CN120875082A
Rock particle classification method and device based on cross-modal dynamic graph attention collaboration
CN121074882A