AI auxiliary image annotation method based on semi-supervised learning
By adopting a semi-supervised learning-based AI-assisted method in image annotation, using graph convolutional networks and deformable convolutional networks, adaptive matching between vertices and global feature fusion is achieved, which solves the problems of insufficient contour accuracy and error accumulation in traditional methods, and significantly improves the accuracy and automation of the annotation.
Patent Information
- Application Number
- CN202510184989.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
AI Technical Summary
When traditional image annotation methods deal with complex or irregular objects, the difference between the initial contour and the real contour is too large, resulting in increased difficulty in model optimization and unreasonable deformation paths, which affects the labeling accuracy. At the same time, traditional methods ignore global information, lack local features, and make it difficult to correct large errors.
Using AI-assisted image annotation method based on semi-supervised learning, a dynamic graph structure is constructed through a graph convolution network to realize adaptive matching and dynamic pairing strategies between vertices. Combining global and local features, the initial contour is progressively optimized through a deformable convolution network to update the vertex position information in real time.
It significantly improves the accuracy and automation of object contour marking, solves the problem of insufficient contour accuracy, and avoids the problems of error accumulation and unreasonable deformation paths.
Smart Images

Figure CN120182974A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image annotation, and particularly relates to an AI-assisted image annotation method based on semi-supervised learning. Background Art
[0002] In image annotation tasks, traditional methods often rely on manually designed initial contours. When faced with complex or irregular object shapes, this approach is prone to the problem that the initial contour differs greatly from the true contour. This difference not only increases the difficulty of model optimization but may also lead to an unreasonable deformation path, thereby affecting the final annotation accuracy. Especially when dealing with objects with complex edges or rich details, manually designed initial contours often fail to cover all details, resulting in the model being unable to accurately capture the true boundary of the target during the deformation process. In addition, traditional methods usually adopt a fixed vertex pairing strategy. When dealing with objects with complex or irregular shapes, this method is prone to error accumulation due to improper vertex matching, further reducing the performance of the model. The problem of incorrect matching is particularly prominent because the fixed pairing strategy cannot dynamically adjust the vertex correspondence according to the specific shape of the object, resulting in poor performance of the model when faced with local shape changes. This limitation is particularly obvious when the object edge is complex or the number of vertices does not match, ultimately affecting the segmentation effect of the model.
[0003] Another technical problem is that traditional methods usually rely on local features for contour optimization while ignoring the integration of global information. Although local features can provide detailed information, they cannot fully capture the overall shape and context relationship of the object, resulting in difficulty in correcting large errors during the contour optimization process. Especially when dealing with scenarios with complex backgrounds or multiple object interactions, the inadequacy of local features will significantly affect the fitting ability of the model, making the contour unable to accurately fit the true boundary of the object. These problems together constitute the technical challenges in image annotation tasks and need to be solved by a more intelligent mechanism. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an AI-assisted image annotation method based on semi-supervised learning. The present invention introduces a graph convolutional network to construct a dynamic graph structure, realizes adaptive matching and dynamic pairing strategies between vertices, and through a deformable convolutional network, can progressively optimize the initial contour, update the vertex position information in real time, fuse global and local features, and effectively solve the problem of insufficient contour accuracy in traditional image annotation.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: An AI-assisted image annotation method based on semi-supervised learning, the steps are: S1. Obtain the image to be annotated, extract multi-layer features of the image using a deep learning network, and generate a feature map; S2. Based on the feature map, construct a global attention mechanism, calculate the pixel-level correlation weights, and obtain the global context information; S3. Fuse the global context information with the local feature information to generate an enhanced feature representation for contour prediction; S4. Based on the enhanced feature representation, use a graph convolutional network to model the relationships between vertices and construct a dynamic graph structure; S5. According to the dynamic graph structure, calculate the adaptive matching weights between vertices to implement a dynamic vertex pairing strategy; S6. Through the dynamic vertex pairing strategy, combine the global and local features to predict the initial contour position; In this step, a contour initialization module is used to directly regress the offset of the contour vertices by learning the center point features. Assume that the target contour contains N vertices, and the center point coordinates are , and the network generates the center point features based on the input image . The offset of each vertex is predicted by the network, denoted as , where i = 1, 2, ..., N. Subsequently, the vertex coordinates of the initial contour are obtained through the following formula: ; Different from the manual design of traditional methods, the initial contour generated in this way is closer to the real target, avoiding the problem of unreasonable deformation paths. At the same time, it ensures that the deformation direction of each vertex always starts from the center point, guaranteeing the uniqueness and rationality of the path. In addition, since this method directly generates the initial contour from network learning without relying on artificially designed shapes or complex post-processing steps, the training complexity is effectively reduced.
[0006] However, the initial contour regressed only by local features may still have a large error from the real contour. For this reason, the present invention introduces a global deformation module to further optimize the initial contour, refine and optimize the preliminary contour to make it more accurately fit the edge of the object.
[0007] Based on the initial contour generated by the contour initialization module, the global contour deformation module starts to work. This module will use the features of all contour vertices in the image for adjustment to ensure that the position of each vertex is optimized. The model combines the context information of the image (such as the edges and shapes of the object) with the local information of the contour vertices and adaptively adjusts the contour through a neural network. This step makes the contour of each instance more accurate and can handle complex object shapes and details.
[0008] The core idea of the global deformation module is to integrate the global features of all vertices and the center point, and generate a rough contour through a one-time deformation. Specifically, this module combines all vertex features of the initial contour Connect with the center point feature to form a feature vector of length (N + 1) × C , where C is the number of channels. Subsequently, this feature vector is processed by a multi-layer perceptron to output the offset of each vertex , expressed as: ; Next, the vertex coordinates of the rough contour are calculated by the following formula: ; In this process, the global deformation module solves the problem of insufficient local features by integrating global information, avoids the computational overhead of multiple iterations required in traditional methods, and effectively corrects large errors. Compared with traditional local feature aggregation methods (such as the recurrent convolution in Deep Snake), the global deformation module achieves more efficient vertex adjustment through a one-time operation and significantly improves the overall fitting quality of the contour. Experiments show that this global deformation mechanism can not only converge quickly but also generate high-quality rough contours.
[0009] S7. Based on the predicted initial contour, use a deformable convolutional network to achieve progressive optimization of the contour; S8. During the progressive optimization process, update the vertex position information in real-time and adjust the graph structure relationship; S9. Finally, output the accurate object contour annotation result to complete the image annotation task.
[0010] The sub-steps of S1 are: S1.1 Use a convolutional neural network to extract multi-layer features from the acquired image and generate a feature map through a preset network structure; S1.2 Determine the annotation data of the image according to the content of the feature map; S1.3 If the generation of the feature map meets the preset threshold, use image processing technology to optimize the feature layer; S1.4 Use a deep learning algorithm to extract the optimized feature layer again to obtain a more accurate feature map; S1.5 Judge whether the annotation data of the image meets the preset standard according to the accurate feature map; S1.6 If the annotation data meets the standard, generate the final annotation result; S1.7 If the annotation data does not meet the standard, re-perform feature extraction and map generation.
[0011] The sub-steps of S2 are: S2.1 For the feature map, use a global attention mechanism to calculate the pixel-level correlation weight; S2.2 Obtain global context information according to the weight value distribution; S2.3 Determine the information content of the feature map through mapping processing; S2.4 If the information content is greater than the preset threshold, perform mechanism optimization; S2.5 Adopt a computational complexity evaluation method to judge the correlation strength; S2.6 Determine the application range of the attention mechanism according to the global distribution; S2.7 Obtain the final global feature representation through context information.
[0012] The sub-steps of S3 are: S3.1 Obtain the multi-scale feature map of the input image through a convolutional neural network, and use the self-attention mechanism to extract global context information; S3.2 Extract local feature information at different levels in the feature pyramid network, and use a channel attention module to screen and optimize the local features; S3.3 Set a feature fusion weight matrix, and adaptively weight and fuse the global context information and local feature information according to the feature importance; S3.4 In the feature fusion module, use bilinear interpolation to align feature information at different scales to generate a unified enhanced feature representation; The present invention also proposes a multi-directional alignment module, which aims to solve problems such as high learning difficulty, unreasonable deformation path, and limited performance caused by the mismatch between the initial contour and the true contour in the image annotation method. Its core idea is to fix vertices in several directions around the center point and uniformly sample the true contour between the fixed directions, reduce the learning difficulty of the model, and optimize the deformation path at the same time to ensure the accuracy of image annotation.
[0013] Specifically, the multi-directional alignment module starts from the center point and generates several vertices in fixed directions at a fixed angular interval. Suppose there are M fixed directions, and the angles of these directions are calculated by the formula: , and the coordinates of the fixed-direction vertices are obtained through the following formula: ; where is the center point coordinate of the target, is the radius from the vertex to the center point. Through this formula, the multi-directional alignment module ensures that the vertices in the fixed directions are evenly distributed on the initial contour of the target at uniform angles. Between the fixed directions, the module uniformly samples the remaining N - M vertices on the real contour, making the distribution of the sampled vertices more regular. The vertices in the fixed directions reduce the degree of freedom of vertex pairing, while the uniform sampling method ensures that the distribution of the remaining vertices can reasonably cover the real contour. The key parameter of this module is the number M of fixed directions. When the value of M is low (e.g., M = 0), no direction is fixed, which is equivalent to completely free vertex pairing. Although the performance ceiling is high, the learning difficulty is very great; when M = N, all vertices are fixed in directions, the learning difficulty is the lowest, but the performance ceiling is also significantly reduced. Through experiments, it is found that M = 4 is an optimal choice, achieving a balance between learning difficulty and performance.
[0014] The advantage of the multi-directional alignment module is that it optimizes the deformation path, enabling the deformation of each vertex to start from the center point and be adjusted through the fixed directions, greatly simplifying the vertex pairing problem in training, and at the same time eliminating the problems of path crossing and unreasonable deformation paths commonly seen in traditional methods. In addition, due to the more regular distribution between vertices, the error accumulation in the training process is reduced, and the fitting accuracy of the model is improved.
[0015] S3.5 Input the enhanced feature representation into the fully convolutional network, and predict the boundary contour of the target object through pixel-level classification; S3.6 If there are discontinuous regions in the predicted contour, use the morphological processing algorithm to connect and smooth the contour; S3.7 Finally, output the complete closed boundary contour of the target object to complete the contour prediction task.
[0016] The sub-steps of S4 are: S4.1 Obtain the dynamic graph data, extract the initial feature information of vertices and edges, and construct the initial adjacency matrix of the graph; S4.2 Use the graph convolutional network to perform multi-layer propagation on the vertex features, and combine the information transfer mechanism to enhance the feature representation ability of vertices; S4.3 Based on the enhanced vertex features, calculate the similarity between vertices and update the adjacency relationship of the dynamic graph; S4.4 According to the updated adjacency relationship, reconstruct the dynamic structure model of the graph to capture the evolution trend between vertices; S4.5 Through the connection relationship of the vertex frame, judge the weight of feature fusion and optimize the output result of the graph convolutional network; S4.6 Combine the feature fusion structure model to generate the final feature representation of the dynamic graph as the basic input for the model construction; S4.7 Introduce data evolution information in the network layer, iteratively train the features of the dynamic graph, and improve the representation accuracy of the model.
[0017] In this step, a dynamic matching loss function is adopted to solve the limitations of traditional fixed vertex pairing strategies in contour fitting. These methods often make it difficult for the model to be optimized due to the rigid rules of vertex matching, especially performing poorly in complex or irregular contour shapes. The core idea of this loss function is to dynamically adjust the pairing relationship between predicted vertices and real contour vertices, optimize the loss function through the optimal matching strategy, enable the model to more flexibly fit the real contour, and avoid error accumulation at the same time.
[0018] Specifically, assume that the target contour contains real vertices and predicted vertices. The goal of this loss function is to minimize the point-to-point distance between the real contour and the predicted contour. The definition of the dynamic matching loss function includes two parts: the matching loss from the predicted vertices to the real contour and the matching loss from the real contour to the predicted vertices. Its loss function formula is as follows: ; P is the set of predicted contour vertices; T is the set of real contour vertices; represents the distance loss between each predicted vertex and the nearest vertex on the real contour; represents the distance loss between each real vertex and the nearest vertex on the predicted contour.
[0019] The dynamic matching strategy ensures that the predicted vertices and real vertices can dynamically adjust their pairing relationship according to their positions, rather than being forced to be fixed. This optimal matching scheme can effectively solve the problem of vertex mis-matching in traditional methods, especially being particularly important when the contour shape is complex or the number of vertices does not match.
[0020] Through the dynamic matching loss function, the model can flexibly adapt to the details and local shapes of the real contour during training, and at the same time smoothly optimize the overall fitting quality of the entire contour. Experiments show that this loss function is significantly better than the L1 loss and Chamfer distance on targets with complex boundaries, with the boundary quality (Boundary AP) improving by 1.1, ensuring the detail accuracy of contour fitting.
[0021] The sub-steps of S5 are: S5.1 Obtain the eigenvalue changes of vertices in the time dimension by constructing the vertex feature matrix of the dynamic graph; S5.2 For the edge connection relationship, use the cosine similarity algorithm to calculate the dynamic similarity between vertices to obtain the preliminary matching degree; S5.3 Design an adaptive weight update mechanism according to the changes in the time dimension to determine the weight value adjustment rule; S5.4 If the vertex eigenvalue changes significantly, increase the similarity weight; otherwise, decrease the weight. S5.5 Based on the updated weight value, generate a vertex pairing strategy to mark the vertex pairs with high matching degrees. S5.6 Integrate the pairing strategy with the dynamic graph structure to construct an adaptive matching relationship between vertices.
[0022] The sub-steps of S6 are as follows: S6.1 Obtain the target scene data and extract multi-dimensional feature information. S6.2 Use the eigenvalue calculation method to calculate the weight distribution of the global feature and the local feature. S6.3 Determine the dynamic vertex pairing label generation strategy according to the weight distribution. S6.4 Establish the mapping relationship between the global feature and the local feature through the pairing label generation strategy. S6.5 Based on the mapping relationship, calculate the set of position points of the target initial position and the contour line. S6.6 Input the set of position points into the prediction model and output the initial contour position of the target. S6.7 Adjust the feature weight and the pairing strategy according to the difference between the predicted value and the actual data.
[0023] The sub-steps of S7 are as follows: S7.1 Obtain the initial contour structure corresponding to the predicted value to get the preliminary contour information. S7.2 Use the preset deformable convolutional network structure to extract the contour edge features. S7.3 Perform multi-scale convolutional operations on the extracted eigenvalues to obtain the contour detail information. S7.4 According to the contour detail information, establish a progressive optimization model to determine the optimization parameters. S7.5 If the contour detail information meets the preset threshold, stop the optimization process. S7.6 Output the final contour structure through the optimization model. S7.7 Store the final contour structure at the preset target position.
[0024] The sub-steps of S8 are as follows: S8.1 Obtain the vertex set and its initial position value, and calculate the current error amount. S8.2 If the error amount is greater than the preset threshold, call the optimizer to generate an update amount. S8.3 Adjust the vertex position value according to the update amount to obtain the new position value. S8.4 Recalculate the weight value of the relationship edge using the new position value; S8.5 Update the relationship edges in the structure diagram according to the weight value; S8.6 Determine whether the number of iterations reaches the upper limit or the convergence value meets the requirements; S8.7 If not satisfied, repeat the above process until the constraint terms and target values are met.
[0025] The sub-steps of S9 are as follows: S9.1 Process the image data through a preset edge detection algorithm to obtain an initial contour map; S9.2 For the initial contour map, adopt a segmentation method based on feature points to extract the candidate region of the target object; S9.3 Judge the pixel value distribution characteristics of the candidate region. If it meets the preset threshold judgment, it is determined as a valid contour region; S9.4 Classify the valid contour region through a classification algorithm to obtain a preliminary annotation result; S9.5 Calculate the error rate between the preliminary annotation result and the standard data set. If the error rate is higher than the preset value, adjust the threshold judgment for re-judgment; The process of calculating the error rate between the preliminary annotation result and the standard data set in this step is as follows: First is data preparation to ensure that the data set meets the input requirements of the model. The annotation file needs to contain polygon or mask information for instance segmentation. Usually, the COCO or Pascal VOC format is recommended. Divide the data set into a training set, a validation set, and a test set for model training, hyperparameter tuning, and final evaluation respectively. To enhance the model's adaptability to new data, data augmentation techniques can be introduced, including operations such as random flipping, rotation, scaling, and cropping to increase data diversity.
[0026] Next, ensure the preparation of the running environment. Install the necessary deep learning frameworks and install relevant dependencies such as numpy, opencv-python, and matplotlib for data processing and visualization. In the model fine-tuning stage, the configuration file of the model needs to be modified according to the characteristics of the data set, including the number of classes, the input image size, and training hyperparameters. Monitor the training log, observe whether the loss function converges normally, and adjust the learning rate or other hyperparameters according to the results of the validation set.
[0027] After training is completed, use the test script to load the model weights and run inference on the test set to generate the predicted contour results. The test process will output common evaluation metrics, including the average segmentation precision, the average boundary precision, and the inference efficiency. These metrics can intuitively reflect whether the performance of the model meets the expectations. To visually understand the model's effect, it is recommended to visually compare the true contours and the predicted contours. For example, use Matplotlib to draw the true and predicted polygon contours, or overlay the image annotation results on the original image to check the model's performance in terms of boundary details.
[0028] S9.6 Evaluate the accuracy of the final annotation results according to the precision evaluation metrics; S9.7 Output the final contour annotation results that meet the precision requirements.
[0029] The present invention can achieve the following beneficial effects: The present invention extracts multi-layer features of an image through a deep learning network, and combines a global attention mechanism to obtain context information to generate an enhanced feature representation.
[0030] The present invention introduces a graph convolutional network to construct a dynamic graph structure, realizes adaptive matching and dynamic pairing strategies between vertices. Through a deformable convolutional network, the initial contour can be progressively optimized, and the vertex position information can be updated in real time. By fusing global and local features, the problem of insufficient contour precision in traditional image annotation is effectively solved.
[0031] The technical effect of the present invention is reflected in significantly improving the accuracy and automation of object contour annotation, and providing a new solution for image segmentation and object detection tasks in the field of computer vision. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 is the flow chart of the present invention; Figure 2 is the learnable contour initialization framework diagram of the present invention; Figure 3 is the convolutional neural network diagram of the present invention; Figure 4 is the experimental comparison effect diagram of the present invention Figure 1 ; Figure 5 is the experimental comparison effect diagram of the present invention Figure 2 . DETAILED DESCRIPTION OF THE INVENTION
[0033] The preferred solution is as Figures 1 to 5 shown. An AI-assisted image annotation method based on semi-supervised learning specifically includes the following steps: Step S1: Obtain the image to be annotated, extract multi-level features of the image using a deep learning network, and generate a feature map.
[0034] Extract multi-level features from the obtained image using a convolutional neural network, and generate a feature map through a preset network structure. Determine the annotation data of the image according to the content of the feature map. If the generation of the feature map meets the preset threshold, optimize the feature layer using image processing techniques. Extract the optimized feature layer again through a deep learning algorithm to obtain a more accurate feature map. Judge whether the annotation data of the image meets the preset standard according to the accurate feature map. If the annotation data meets the standard, generate the final annotation result. If the annotation data does not meet the standard, re-perform feature extraction and map generation.
[0035] Specifically, when obtaining the image to be annotated, it can be batch-collected from a public dataset or a business system through an automated tool. For example, randomly extract 1000 images with a resolution of 224x224 from the ImageNet dataset. When extracting multi-level features of the image, use a VGG16 deep convolutional neural network, which contains 13 convolutional layers and 3 fully connected layers, and fine-tune it based on the ImageNet pre-trained model. After adjusting the input image size to 224x224, successively pass through the convolutional layer, pooling layer, and activation function to extract features at different levels. Among them, the size of the first convolutional kernel is 3x3, the stride is 1, and 64 feature maps are output; the second convolutional kernel is still 3x3, and 128 feature maps are output; as the network depth increases, the third layer outputs 256 feature maps, the fourth layer outputs 512 feature maps, and the finally generated feature map contains 512 channels. During the feature extraction process, set the learning rate to 0.01, use the ReLU activation function, and use the maximum pooling layer for downsampling, with a pooling window of 2x2 and a stride of 2. In this way, feature information at different scales can be extracted from the original image, providing effective input data for subsequent tasks such as image classification or object detection.
[0036] Step S2: Based on the feature map, construct a global attention mechanism, calculate pixel-level correlation weights, and obtain global context information.
[0037] For the feature map, adopt a global attention mechanism to calculate pixel-level correlation weights. Obtain global context information according to the weight value distribution. Determine the information volume of the feature map through mapping processing. If the information volume is greater than the preset threshold, perform mechanism optimization. Use a computational complexity evaluation method to judge the correlation strength. Determine the application scope of the attention mechanism according to the global distribution. Obtain the final global feature representation through the context information.
[0038] Specifically, based on the feature map, first, the features of the input image are extracted through a convolutional neural network to generate a feature map with a size of 256×64×64. To construct the global attention mechanism, a self-attention module is used to calculate the pixel-level correlation weights. Specifically, the feature map is flattened into a 256×4096 matrix, and three linear transformation layers are respectively used to generate the query matrix Q, the key matrix K, and the value matrix V. The dimensions of Q and K are both 64×4096, and the dimension of V is 256×4096. The transpose of Q and K is calculated through matrix multiplication to obtain a 4096×4096 similarity matrix, which represents the correlation strength between each pixel and other pixels. To normalize the similarity, the Softmax function is used to process each row to obtain the normalized attention weight matrix. Then, the attention weight matrix is multiplied by V to obtain the weighted feature representation with a dimension of 256×4096. Finally, the weighted features are reshaped into a size of 256×64×64 and added to the original feature map to obtain the final feature map that integrates global context information. For example, when processing an image with a resolution of 64×64, through the above process, the correlations between pixels that are far apart in the image can be captured, such as the color gradient relationship between the sky and the ground, thereby enhancing the model's understanding ability of complex scenes.
[0039] Step S3: Integrate the global context information and the local feature information to generate an enhanced feature representation for contour prediction.
[0040] Obtain the multi-scale feature maps of the input image through a convolutional neural network, and use the self-attention mechanism to extract the global context information. Extract the local feature information at different levels in the feature pyramid network, and use the channel attention module to screen and optimize the local features. Set the feature fusion weight matrix, and adaptively weight and fuse the global context information and the local feature information according to the feature importance. In the feature fusion module, use bilinear interpolation to align the feature information at different scales to generate a unified enhanced feature representation. Input the enhanced feature representation into the fully convolutional network, and predict the boundary contour of the target object through pixel-level classification. If there are discontinuous regions in the predicted contour, use the morphological processing algorithm to connect and smooth the contour. Finally, output the complete closed boundary contour of the target object to complete the contour prediction task.
[0041] Specifically, assume we have a medical imaging picture for predicting the contour of a tumor. The global context information can be extracted through a deep convolutional neural network (e.g., ResNet-50, whose convolutional layer outputs a feature map with a size of 2048x7x7). Specifically, the original image (e.g., 224x224x3) is input into ResNet-50, its last global average pooling layer and fully connected layer are removed, and the output of the convolutional layer is retained. In this way, we obtain a global feature map, where each 7x7 spatial position corresponds to a 2048-dimensional feature vector, representing the global context information of different regions of the image. Then, to obtain local features, we can perform a sliding window operation on the original image using a smaller convolutional kernel (e.g., 3x3, with a stride of 1) to extract the local information around each pixel. Assume we extract 64 local features for each pixel, then each pixel has a 64-dimensional local feature vector. To fuse these two types of features, we can adopt the method of feature concatenation. First, the global feature map is upsampled to the same size as the original image (224x224) through bilinear interpolation, obtaining a global feature tensor of 2048x224x224. Then, for each pixel position, the corresponding 2048-dimensional global feature vector and 64-dimensional local feature vector are concatenated to obtain a 2112-dimensional fused feature vector. In this way, each pixel has an enhanced feature representation that fuses global and local information. To further optimize the feature representation, we introduce an attention mechanism. A small convolutional network (e.g., containing two 3x3 convolutional layers, with an output channel number of 1) is applied to the fused feature tensor (2112x224x224) to generate an attention weight map (1x224x224). This weight map represents the importance of each pixel. Then, the attention weight map is multiplied element-wise with the fused feature tensor to obtain the weighted fused feature. Finally, the weighted fused feature is input into a fully convolutional network (FCN) for contour prediction. The FCN can adopt the structure of U-Net, including an encoder and a decoder path, and fuses the features in the encoder stage with the features in the decoder stage through skip connections. The output of the FCN is a probability map (1x224x224) with the same size as the input image, where the value of each pixel represents the probability that the pixel belongs to the tumor contour. The entire network is trained using the cross-entropy loss function, and the network parameters are updated through the backpropagation algorithm, finally obtaining a model that can accurately predict the tumor contour.
[0042] Step S4, based on the enhanced feature representation, use a graph convolutional network to model the relationships between vertices and construct a dynamic graph structure.
[0043] Obtain dynamic graph data, extract the initial feature information of vertices and edges, and construct the initial adjacency matrix of the graph. Use a graph convolutional network to perform multi-layer propagation on vertex features, and combine the information transfer mechanism to enhance the feature representation ability of vertices. Based on the enhanced vertex features, calculate the similarity between vertices and update the adjacency relationship of the dynamic graph. According to the updated adjacency relationship, reconstruct the dynamic structure model of the graph to capture the evolution trend between vertices. Through the connection relationship of the vertex frame, judge the weight of feature fusion and optimize the output result of the graph convolutional network. Combine the feature fusion structure model to generate the final feature representation of the dynamic graph as the basic input for model construction. Introduce data evolution information in the network layer to iteratively train the features of the dynamic graph and improve the representation accuracy of the model.
[0044] Specifically, in terms of enhancing feature representation, first extract the initial features of vertices through feature engineering. For example, use the TF-IDF algorithm to extract features from text data to obtain a 128-dimensional feature vector for each vertex. Then, use an autoencoder to reduce the dimension of these features and compress the dimension to 64 to reduce noise and retain key information. Next, use a graph convolutional network (GCN) to model the relationship between vertices. Through a two-layer GCN structure, the ReLU activation function is used in the first layer, and the Softmax function is used in the second layer to finally output the classification probability of each vertex. When constructing the dynamic graph structure, introduce time series analysis and use the LSTM network to capture the dynamic changes of vertex features. Set the time window to 10 and the step size to 1 to capture short-term and long-term dependencies. By combining GCN and LSTM, the dynamic relationship between vertices can be modeled more accurately, and the prediction performance of the graph structure can be improved. For example, in social network analysis, the change trend of user behavior can be predicted through the above method with an accuracy rate of over 85%.
[0045] Step S5: Calculate the adaptive matching weight between vertices according to the dynamic graph structure to implement the dynamic vertex pairing strategy.
[0046] By constructing the vertex feature matrix of the dynamic graph, obtain the change of vertex eigenvalue in the time dimension. For the edge connection relationship, use the cosine similarity algorithm to calculate the dynamic similarity between vertices to obtain the preliminary matching degree. According to the change in the time dimension, design an adaptive weight update mechanism to determine the weight value adjustment rule. If the vertex eigenvalue changes significantly, increase the similarity weight; otherwise, decrease the weight. Based on the updated weight value, generate a vertex pairing strategy and mark the vertex pairs with high matching degrees. Integrate the pairing strategy with the dynamic graph structure to construct an adaptive matching relationship between vertices.
[0047] Specifically, assume there is a dynamically changing social network graph where vertices represent people and edges represent the interaction relationships between people, and the interaction frequency changes over time. We need to calculate the adaptive matching weights between vertices to achieve dynamic vertex pairing for friend recommendation. At the initial moment, the interaction relationships of users A, B, and C are as follows: The number of interactions between A and B in the past week is 10 times, between A and C is 2 times, and between B and C is 5 times. We can represent this using an adjacency matrix W, where W[i][j] represents the number of interactions between vertices i and j. Initially, W = [[0, 10, 2], [10, 0, 5], [2, 5, 0]]. To calculate the adaptive matching weights, we introduce a time decay factor λ = 9, indicating that the influence of past interactions decays over time. Then, we update the weights using a weighted average method. For example, at the next moment, A and B interacted 5 times, A and C interacted 8 times, and B and C interacted 3 times. The new interaction matrix is W_new = [[0, 5, 8], [5, 0, 3], [8, 3, 0]]. The updated weight matrix W_updated = λ * W + (1 - λ) * W_new = 9 * [[0, 10, 2], [10, 0, 5], [2, 5, 0]] + 1 * [[0, 5, 8], [5, 0, 3], [8, 3, 0]] = [[0, 5, 6], [5, 0, 8], [6, 8, 0]]. This matrix represents the latest interaction intensities between A - B, A - C, and B - C, and takes into account the influence of historical interactions. To achieve dynamic vertex pairing, we can set a threshold, such as 0. By traversing the W_updated matrix, we find that the weights of A - B (5) and B - C (8) exceed the threshold. Then, in the friend recommendation system, we can recommend B to A and C to B. To form a chain of thought, we introduce a feedback mechanism for friend recommendation. Assume that the probability of a user accepting a recommendation is proportional to the weight. We can adjust the λ value in the weight update formula based on whether the user accepts the recommendation. If users generally accept the recommendation, it indicates that the current weight calculation is relatively accurate, and we can appropriately increase λ to enhance the influence of historical data; if users generally reject the recommendation, we decrease λ to give more consideration to the latest interaction data. This feedback mechanism makes the weight calculation more adaptive and the pairing strategy more dynamic and accurate.
[0048] Step S6, through the dynamic vertex pairing strategy, combine the global and local features to predict the initial contour position.
[0049] Obtain target scenario data and extract multi-dimensional feature information. Adopt an eigenvalue calculation method to calculate the weight distribution of global features and local features. Determine the dynamic vertex pairing label generation strategy according to the weight distribution. Through the pairing label generation strategy, establish the mapping relationship between global features and local features. Based on the mapping relationship, calculate the set of position points of the target initial position and the contour line. Input the set of position points into the prediction model to output the initial contour position of the target. Adjust the feature weights and pairing strategies according to the difference between the predicted value and the actual data.
[0050] Specifically, assume we have a medical image segmentation task with the goal of segmenting the initial contour of the heart. The input image is a grayscale image of 256x256 pixels. First, a pre-trained U-Net network (encoder-decoder structure) is used to extract the global features of the image. In the encoder part of U-Net, assuming 4 times of downsampling, a feature map of 16x16x1024 is obtained, representing the global semantic information of the image. At the same time, in order to capture local details, we apply the Sobel operator on the original image to calculate the edge gradient. Specifically, 3x3 Sobel kernels in the horizontal and vertical directions are used respectively: [-1, 0, 1; -2, 0, 2; -1, 0, 1] and [-1, -2, -1; 0, 0, 0; 1, 2, 1], obtaining two 256x256 gradient maps, representing the edge intensities in the horizontal and vertical directions respectively. Then, dynamic vertex pairing is performed. We first uniformly sample 64 vertices on the global feature map as the initial control points. To achieve dynamic pairing, we calculate the similarity between each vertex and its surrounding 8 neighbors (including diagonals). The similarity metric uses cosine similarity, and the calculation formula is: Sim(A, B) = (A · B) / (||A|| * ||B||), where A and B are the feature vectors of the two vertices (extracted from the global feature map). Then, according to the similarity ranking, the 4 most similar neighbors are selected for each vertex to form pairs. After pairing, a graph structure is formed, where vertices represent control points and edges represent pairing relationships. Then, a graph convolutional network (GCN) is used to further fuse the global and local features. The input of GCN is the global feature of each vertex (a 1024-dimensional vector extracted from U-Net) and the local feature (Sobel gradient values, normalized to the [0, 1] interval, and then the number of channels is expanded to the same as the global feature, i.e., 1024 dimensions through a 1x1 convolutional layer). GCN aggregates the features of neighbor nodes through the adjacency matrix (constructed by the pairing relationship) to update the feature representation of each vertex. For example, two layers of GCN are used, and each layer outputs 512-dimensional features. Finally, the features of each vertex output by GCN are input into a fully connected layer to predict the offset (Δx, Δy) of each vertex relative to its initial position. For example, the output of the fully connected layer is a 64x2 matrix, representing the offsets of 64 vertices in the x and y directions respectively. Adding these offsets to the initial vertex positions, the predicted initial contour can be obtained. To form a closed contour, the predicted vertices are connected in sequence.
[0051] Step S7, based on the predicted initial contour, use a deformable convolutional network to achieve progressive optimization of the contour.
[0052] Obtain the initial contour structure corresponding to the predicted value to get the preliminary contour information. Adopt a preset deformable convolutional network structure to extract the contour edge features. Perform multi-scale convolutional operations on the extracted eigenvalue to obtain the contour detail information. According to the contour detail information, establish a progressive optimization model to determine the optimization parameters. If the contour detail information meets the preset threshold, stop the optimization process. Output the final contour structure through the optimization model. Store the final contour structure at the preset target location.
[0053] Specifically, assume we have a medical image dataset containing lung CT scan images and corresponding lung nodule contour annotations. Our goal is to segment the lung nodules. First, use a basic U-Net network. Input a CT slice with a size of 256x256, and through a series of convolutional and downsampling operations, extract multi-scale features. Finally, output a probability map of the same size as the input image. Each pixel value in the probability map represents the probability that the pixel belongs to a lung nodule. The threshold is set to 5. Connect the pixel points with a probability greater than 5 to form the initial lung nodule contour. This initial contour may be relatively rough, with problems such as inaccurate edges and missing details. Subsequently, introduce a deformable convolutional network (DCN) to optimize this initial contour. Specifically, in the decoder part of the U-Net, we replace the standard 3x3 convolution with a deformable convolution. The deformable convolution dynamically adjusts the sampling position of the convolution kernel by learning an offset field. For example, for a 3x3 deformable convolution, it will additionally learn a 2x3x3 offset field, where 2 represents the offsets in the x and y directions, and 3x3 corresponds to the 9 sampling points of the convolution kernel. These offsets can be decimals, indicating that the sampling position can deviate from the standard grid points. The initial value of the offset field is set to 0, indicating that the initial sampling position is the same as the standard convolution. Through the backpropagation algorithm, use the contour annotation information as the supervision signal to optimize the offset field and the weights of the convolution kernel. The loss function uses Dice Loss to calculate the overlap degree between the predicted contour and the true contour. During the training process, the offset field gradually learns how to adjust the sampling position according to the image content, so that the convolution kernel can more accurately capture the edges and details of the lung nodules. After the first optimization, an improved contour is obtained. Then, use this improved contour as the input of the DCN again for the second optimization to further improve the accuracy of the contour. This progressive optimization process can be repeated multiple times, such as 3 times. Each optimization is based on the result of the previous time, gradually approaching the true lung nodule contour. To ensure that the offset does not become too large and cause sampling beyond the image boundary, the offset can be clipped, for example, limited within the range of [-8, 8] pixels. By generating the initial contour based on U-Net and then using the deformable convolutional network for multiple progressive optimizations, the accuracy of lung nodule segmentation can be effectively improved.
[0054] Step S8, during the progressive optimization process, the vertex position information is updated in real time to adjust the graph structure relationship.
[0055] Obtain the vertex set and its initial position values, and calculate the current error amount. If the error amount is greater than the preset threshold, call the optimizer to generate an update amount. Adjust the vertex position values according to the update amount to obtain new position values. Recalculate the weight values of the relationship edges using the new position values. Update the relationship edges in the structure graph according to the weight values. Determine whether the number of iterations reaches the upper limit or the convergence value meets the requirements. If not, repeat the above process until the constraint terms and target values are satisfied.
[0056] Specifically, assume that there is an initial triangular mesh in a three-dimensional space, with vertex coordinates A(0, 0, 0), B(0, 0, 0), and C(0, 0, 0). Initially, these three vertices form a triangular patch. Now, we use a physics-based spring-mass model for progressive optimization. First, calculate the initial length of each edge. For example, the length of the AB edge is √((0 - 0)²+(0 - 0)²+(0 - 0)²)=√11. Set the same spring constant k = 10 for each edge and set a target length. Here, assume that the target length of all edges is equal to their initial length, so as to maintain the initial shape. Then, in each iteration, calculate the resultant force on each vertex. Taking vertex A as an example, it is subject to the spring forces from the AB and AC edges. According to Hooke's law, the force exerted by the AB edge on point A is F_AB = k * (||AB|| - target length_AB) * (B - A) / ||AB||. Substituting the values, F_AB = 10 * (√11 - √11) * (0, -0, -0) / √11 = (0, 0, 0). Similarly, calculate the force exerted by the AC edge on point A. Add all the force vectors to get the resultant force on point A. Use the explicit Euler method to update the vertex position. Set the time step Δt = 01, then the new position of point A is A_new = A + F_A * Δt. Since the initial state is in equilibrium, all the forces are 0 and the position remains unchanged after the update. To trigger the optimization process, assume that vertex B is subject to an external force F_ext = (5, 2, 1), then the position update of vertex B will become B_new = B + (the reaction force of F_AB + F_ext) * Δt. This will cause the position of point B to change, which in turn affects the lengths and forces of the AB and BC edges, triggering the deformation of the entire mesh. At the same time, after the vertex positions are updated, we can dynamically adjust the graph structure. For example, if the length of the AB edge exceeds a certain threshold (such as 5 times the initial length), we can insert a new vertex D at the midpoint of the AB edge, with coordinates (A + B) / 2 = (5, 5, 5), and divide the original AB edge into two edges AD and DB. At the same time, update the relevant triangular patch information, splitting the original triangle ABC into triangles ACD and BCD. To ensure the correct calculation of the force on vertex D, we need to set the spring constant and target length for the AD and DB edges. For example, also set k = 10 and the target length to half of the initial length of the AB edge. In this way, the newly added vertices and edges also participate in the subsequent spring-mass model calculations, realizing the dynamic adjustment of the graph structure. Through continuous iteration of vertex position updates and graph structure adjustments, the progressive optimization of the mesh is achieved.
[0057] Step S9, finally output the accurate object contour annotation result to complete the image annotation task.
[0058] Process the image data through a preset edge detection algorithm to obtain an initial contour map. For the initial contour map, adopt a segmentation method based on feature points to extract the candidate regions of the target object. Judge the pixel value distribution characteristics of the candidate regions. If the preset threshold judgment is satisfied, it is determined as a valid contour region. Classify the valid contour region through a classification algorithm to obtain a preliminary annotation result. Calculate the error rate between the preliminary annotation result and the standard data set. If the error rate is higher than the preset value, adjust the threshold judgment for re-judgment. Evaluate the accuracy of the final annotation result according to the accuracy evaluation index. Output the final contour annotation result that meets the accuracy requirements.
[0059] Specifically, assume that the input image is a color RGB image of 800x600 pixels, and the goal is to label the exact contours of all apples in the image. First, convert the input RGB image data into the HSV color space. Utilize the more intuitive expression of color in the HSV space. Through the hue (H) component, set the threshold range to [0, 10] and [170, 180] (the red region is at both ends of the HSV color wheel), and initially screen out the reddish regions in the image. Then, use the Canny edge detection algorithm, set the high threshold to 150 and the low threshold to 50, and perform edge detection on the screened red regions to obtain candidate edge pixels. The specific steps of the Canny algorithm include: first, smooth the image with a 5x5 Gaussian filter (σ = 4) to reduce noise; then calculate the gradient intensity and direction of the image in the horizontal and vertical directions; then perform non-maximum suppression to only retain the local maximum points in the gradient direction; finally, use double thresholds and hysteresis techniques to connect the edges. Pixels above the high threshold are considered strong edges, pixels below the low threshold are discarded, and pixels between the two, if connected to strong edge pixels in the 8-neighborhood, are retained, otherwise discarded. Next, use the Suzuki85 contour finding algorithm to extract all closed contours from the results of the Canny edge detection. For each found contour, calculate its Hu invariant moments, which is a set of 7 numerical values used to describe the shape characteristics of the contour and has translational, rotational, and scale invariance. Compare these Hu moments with the pre-trained Hu moment feature library of apple contours, use the Euclidean distance as the similarity measure, and set the threshold to 1. If the Euclidean distance between the Hu moments of a certain contour and the Hu moments of the apples in the library is less than 1, then this contour is considered likely to be an apple contour. Finally, for the screened candidate contours, adopt the polygon approximation algorithm (Douglas-Peucker algorithm), set an approximation accuracy ε = 3, simplify the curve contour into a series of vertices, obtain the final exact apple contour annotation, and output the x and y coordinates of these vertices in coordinate form to complete the annotation.
[0060] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. An AI-assisted image annotation method based on semi-supervised learning, characterized in that The following steps are involved: Obtain the image to be annotated, use the deep learning network to extract the multi-layer features of the image, and generate a feature map; Based on the feature map, a global attention mechanism is constructed to calculate pixel-level association weights to obtain global context information. The global context information is fused with the local feature information to generate enhanced feature representation for contour prediction; Based on enhanced feature representation, a graph convolutional network is used to model the relationship between vertices and construct a dynamic graph structure. According to the dynamic graph structure, the adaptive matching weights between vertices are calculated to implement a dynamic vertex pairing strategy. Through the dynamic vertex pairing strategy, the initial contour position is predicted by combining global and local features. Based on the predicted initial contour, a deformable convolutional network is used to achieve progressive optimization of the contour. During the progressive optimization process, the vertex position information is updated in real time and the graph structure relationship is adjusted. Finally, accurate object contour annotation results are output to complete the image annotation task.
2. According to claim 1, an AI-assisted image annotation method based on semi-supervised learning is characterized in that: The method of obtaining an image to be annotated, extracting multi-layer features of the image using a deep learning network, and generating a feature map includes: A convolutional neural network is used to extract multi-layer features from the acquired image, and a feature map is generated through a preset network structure; Determine the annotation data of the image according to the content of the feature map; If the generated feature map meets the preset threshold, the feature layer is optimized using image processing technology; The optimized feature layer is extracted again through the deep learning algorithm to obtain a more accurate feature map; Based on the precise feature map, determine whether the image annotation data meets the preset standards; If the labeled data meets the standards, the final labeling result is generated; If the labeled data does not meet the standards, feature extraction and atlas generation will be performed again.
3. According to claim 1, an AI-assisted image annotation method based on semi-supervised learning is characterized in that: Based on the feature map, a global attention mechanism is constructed to calculate pixel-level association weights to obtain global context information, including: For the feature map, a global attention mechanism is used to calculate the pixel-level association weights; Obtain global context information based on weight value distribution; Determine the information content of the feature graph through graph processing; If the amount of information is greater than the preset threshold, the mechanism is optimized; The computational evaluation method is used to determine the strength of the correlation; Determine the application scope of the attention mechanism based on the global distribution; Through the context information, the final global feature representation is obtained.
4. According to claim 1, an AI-assisted image annotation method based on semi-supervised learning is characterized in that: The method of fusing global context information with local feature information to generate enhanced feature representation for contour prediction includes: The multi-scale feature map of the input image is obtained through a convolutional neural network, and the global context information is extracted using the self-attention mechanism; Extract local feature information at different levels in the feature pyramid network, and use the channel attention module to screen and optimize the local features; Set the feature fusion weight matrix to perform adaptive weighted fusion of global context information and local feature information according to feature importance; In the feature fusion module, bilinear interpolation is used to align feature information of different scales to generate a unified enhanced feature representation; The enhanced feature representation is input into the fully convolutional network to predict the boundary contour of the target object through pixel-level classification; If there are discontinuous areas in the predicted contour, the morphological processing algorithm is used to connect and smooth the contour; Finally, the complete closed boundary contour of the target object is output to complete the contour prediction task.
5. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: Based on the enhanced feature representation, the graph convolutional network is used to model the relationship between vertices and construct a dynamic graph structure, including: Get dynamic graph data, extract initial feature information of vertices and edges, and construct the initial adjacency matrix of the graph; The graph convolutional network is used to propagate vertex features in multiple layers, and the feature representation ability of vertices is enhanced by combining the information transmission mechanism; Based on the enhanced vertex features, the similarity between vertices is calculated and the adjacency relationship of the dynamic graph is updated; According to the updated adjacency relationship, the dynamic structure model of the graph is reconstructed to capture the evolution trend between vertices; By using the connection relationship of the vertex frame, the weight of feature fusion is determined to optimize the output results of the graph convolutional network; Combine feature fusion and structural model to generate the final feature representation of the dynamic graph as the basic input for model construction; Data evolution information is introduced into the network layer, and the features of dynamic graphs are iteratively trained to improve the representation accuracy of the model.
6. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: The method of calculating the adaptive matching weights between vertices according to the dynamic graph structure and implementing the dynamic vertex pairing strategy includes: By constructing the vertex feature matrix of the dynamic graph, the feature value changes of the vertex in the time dimension are obtained; For edge connection relationships, the cosine similarity algorithm is used to calculate the dynamic similarity between vertices to obtain a preliminary matching degree; According to the changes in the time dimension, an adaptive weight update mechanism is designed to determine the weight value adjustment rules; If the vertex eigenvalue changes significantly, the similarity weight is increased, otherwise the weight is reduced; Based on the updated weight values, a vertex pairing strategy is generated to mark vertex pairs with high matching degrees; The pairing strategy is integrated with the dynamic graph structure to construct adaptive matching relationships between vertices.
7. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: The method predicts the initial contour position by combining global and local features through a dynamic vertex pairing strategy, including: Acquire target scene data and extract multi-dimensional feature information; The eigenvalue calculation method is used to calculate the weight distribution of global features and local features; According to the weight distribution, determine the dynamic vertex pairing label generation strategy; Through the matching mark generation strategy, the mapping relationship between global features and local features is established; Based on the mapping relationship, calculate the target initial position and the location point set of the contour line; Input the position point set into the prediction model and output the initial contour position of the target; Adjust feature weights and pairing strategies based on the difference between predicted values and actual data.
8. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: The method of implementing progressive optimization of the contour based on the predicted initial contour by using a deformable convolutional network comprises: Obtain the initial contour structure corresponding to the predicted value and obtain preliminary contour information; Use the preset deformable convolutional network structure to extract contour edge features; Perform multi-scale convolution operations on the extracted feature values to obtain contour detail information; According to the detailed information of the profile, a progressive optimization model is established to determine the optimization parameters; If the contour detail information meets the preset threshold, the optimization process is stopped; By optimizing the model, the final contour structure is output; The final contour structure is stored in the preset target location.
9. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: In the progressive optimization process, the vertex position information is updated in real time and the graph structure relationship is adjusted, including: Get the vertex set and its initial position value, and calculate the current error; If the error amount is greater than the preset threshold, the optimizer is called to generate the update amount; Adjust the vertex position value according to the update amount to obtain a new position value; Recalculate the weight value of the relationship edge using the new position value; Update the relationship edges in the structure graph according to the weight values; Determine whether the number of iterations reaches the upper limit or whether the convergence value meets the requirements; If not, repeat the above process until the constraints and target values are met.
10. The AI-assisted image annotation method based on semi-supervised learning according to claim 1, characterized in that: The final output is an accurate object contour annotation result to complete the image annotation task, including: Process the image data through a preset edge detection algorithm to obtain an initial contour map; For the initial contour map, a segmentation method based on feature points is used to extract the candidate area of the target object; Determine the pixel value distribution characteristics of the candidate area, and if it meets the preset threshold, it is determined to be a valid contour area; The effective contour area is classified by classification algorithm to obtain the preliminary annotation result; Calculate the error rate between the preliminary annotation result and the standard data set. If the error rate is higher than the preset value, adjust the threshold and make a new judgment. According to the accuracy evaluation index, the accuracy of the final annotation result is evaluated; Output the final contour annotation knot that meets the accuracy requirements.
Citation Information
Cited By
Unit type curtain wall component matching method and system combined with AI identification
CN120599307A
Interactive scanning image intelligent labeling method and device based on panorama
CN122155944A