A method and device for training a semantic segmentation model
The method automates semantic segmentation model training by using street image data to generate pseudo-labels and adjust weights, addressing the cost and sample availability issues of manual labeling, ensuring accurate segmentation without human intervention.
Patent Information
- Application Number
- CN202510422118.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing semantic segmentation methods require manual labeling of each image, increasing labor costs, and labeling samples may be difficult to obtain, resulting in limited model performance.
By acquiring street image data, setting up the initial semantic segmentation model, using initial masks, pseudo-labels and key edge area sets for weight calculations, expanding the image data for training until the total loss of the model is not greater than the preset threshold, reducing manual marking steps and improving model accuracy.
This enables no need to manually label each image, reduces labor costs, and enhances the accuracy of the model and the ability to adapt to complex scenarios through pseudo-label expansion.
Smart Images

Figure CN119919669B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital image technology, and in particular, to a method and device for training a semantic segmentation model. Background Art
[0002] The application of semantic segmentation technology in the automatic parking system is not limited to the detection of parking spaces and obstacles, but can also combine other sensor data (such as lidar, ultrasonic sensors, etc.) to achieve the fusion of multi-modal data, providing richer and more accurate environmental information for the system. For example, in low-light or night-time environments, semantic segmentation combined with the image data of infrared sensors can effectively improve the object detection accuracy in the parking scenario, thus ensuring that the vehicle can work stably under different lighting conditions. With the continuous progress of autonomous driving technology, the application of semantic segmentation in autonomous driving and automatic parking systems is becoming more and more extensive and important. It can not only help autonomous vehicles achieve accurate target recognition in complex environments, but also ensure the safety and efficiency of automatic parking operations.
[0003] In the prior art, when performing a semantic segmentation task, it is necessary for annotators to manually mark each image and collect accurately marked data for training a deep neural network, which increases a great deal of labor costs. In addition, for some specific semantic categories, it may be difficult to obtain annotated samples, resulting in the problem of limited model performance.
[0004] Therefore, there is an urgent need to propose a method and device for training a semantic segmentation model to solve the technical problems in the prior art that the semantic segmentation method requires manual marking of each image, and the annotated samples may be difficult to obtain, resulting in an increase in labor costs and labor costs. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and device for training a semantic segmentation model to solve the technical problems in the prior art that the semantic segmentation method requires manual marking of each image, and the annotated samples may be difficult to obtain, resulting in an increase in labor costs and labor costs.
[0006] To solve the above problems, the present invention provides a method for training a semantic segmentation model, including:
[0007] Obtain image data on the street and set an initial semantic segmentation model;
[0008] Train the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result;
[0009] Calculate weights according to the initial mask, the pseudo labels, and the set of key edge regions to obtain the total model loss;
[0010] When the total loss of the model is greater than a preset total loss threshold, augment the image data according to the pseudo-labels, and train the target semantic segmentation model according to the augmented image data until the total loss of the model is not greater than the preset total loss threshold, and determine that the model training is completed.
[0011] In a possible implementation manner, the training of the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result includes:
[0012] Annotate the image data to obtain a data set;
[0013] Divide the data set to obtain a training set;
[0014] Input the training set into the initial semantic segmentation model for training to obtain a target semantic segmentation model and an output result.
[0015] In a possible implementation manner, the initial semantic segmentation model includes a feature extraction network, a perception layer, an action layer, and a feedback layer. The training of the initial semantic segmentation model by inputting the training set to obtain a target semantic segmentation model and an output result includes:
[0016] Input the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain a feature map;
[0017] Input the feature map into the perception layer for key region extraction to obtain key region features;
[0018] Input the feature map and the key region features into the action layer for mask extraction and dynamic optimization to obtain an initial mask and a dynamic weight map;
[0019] Input the initial mask and the dynamic weight map into the feedback layer for high-confidence region screening to obtain an output result.
[0020] In a possible implementation manner, the output result includes an initial mask, pseudo-labels, and a set of key edge regions. The calculation of the total loss of the model according to the initial mask, the pseudo-labels, and the set of key edge regions includes:
[0021] Perform standard forward propagation and backward propagation calculations on the initial mask and the pseudo-labels to obtain a first weight;
[0022] Obtain a second weight and a third weight according to the set of key edge regions and the pseudo-labels;
[0023] Calculate the first weight, the second weight, and the third weight to obtain the total loss of the model.
[0024] In a possible implementation manner, obtaining a second weight and a third weight according to the set of key edge regions and the pseudo-label includes:
[0025] Performing mask extraction on the set of key edge regions to obtain an edge region mask;
[0026] Calculating an edge region loss for the edge region mask and the pseudo-label to obtain a second weight;
[0027] Performing balance optimization on the edge region mask, the pseudo-label, and a preset weight to obtain a third weight.
[0028] In a possible implementation manner, the feature extraction network includes a hybrid frequency domain convolutional layer and a multi-scale attention residual module; inputting the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain a feature map includes:
[0029] Inputting the image data into the hybrid frequency domain convolutional layer for weight calculation to obtain an adaptive weight for each feature channel;
[0030] Determining a convolution mode corresponding to each feature channel according to the adaptive weight; the convolution mode with the adaptive weight greater than a preset weight threshold is a traditional convolution; the convolution mode with the adaptive weight not greater than the preset weight threshold is a frequency domain convolution;
[0031] Performing feature extraction on the image data according to the convolution mode to obtain a feature image for each feature channel;
[0032] Performing multi-scale feature fusion and adjustment on the feature image according to the multi-scale attention residual module to obtain a feature map.
[0033] In a possible implementation manner, inputting the feature map into the perception layer to extract key regions to obtain key region features includes:
[0034] Inputting the output features of each feature channel of the feature map into the perception layer for global statistics to obtain a global response intensity for each feature channel;
[0035] Determining the feature channels with the global response intensity greater than a preset intensity threshold as key feature channels, and determining the features of the key feature channels as local features;
[0036] Calculating global features according to the global context information and the local features by a global average pooling layer;
[0037] Fuse the local feature and the global feature to obtain an enhanced feature;
[0038] Based on a dynamic gating mechanism, screen the enhanced feature to obtain a key region feature.
[0039] In a possible implementation manner, the calculation process of the first weight is as follows:
[0040]
[0041]
[0042]
[0043]
[0044] In the formula, is the initial mask and the mask data of the i th one in it, is the pseudo-label of the i th one, is the initial mask and the pixel propagation data of the i th one is obtained through the pixel propagation method, is the i th pixel, is the first weight.
[0045] In a possible implementation manner, the calculation process of the second weight is as follows:
[0046]
[0047] In the formula, is the total number of edge region pixels in the edge region mask; is the cross-entropy loss function; is the th edge value in the edge region mask; is the th pseudo-label.
[0048] On the other hand, the present invention also provides a semantic segmentation model training device, including:
[0049] A data acquisition module, configured to acquire image data on a street and set an initial semantic segmentation model;
[0050] A model training module, configured to train the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result;
[0051] A weight calculation module, configured to perform weight calculation according to the initial mask, the pseudo label, and the set of key edge regions to obtain the total model loss;
[0052] A result judgment module, configured to, when the total model loss is greater than a preset total loss threshold, expand the image data according to the pseudo label, and train the target semantic segmentation model according to the expanded image data until the total model loss is not greater than the preset total loss threshold, and determine that the model training is completed.
[0053] The beneficial effects of the present invention are as follows: obtaining image data on the street and setting an initial semantic segmentation model; training the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, a pseudo label, and a set of key edge regions; performing weight calculation according to the initial mask, the pseudo label, and the set of key edge regions to obtain the total model loss; when the total model loss is greater than a preset total loss threshold, expanding the image data according to the pseudo label, and training the target semantic segmentation model according to the expanded image data until the total model loss is not greater than the preset total loss threshold, and determining that the model training is completed; the present invention provides an initial semantic segmentation model, which can process each image without manually marking and processing each image; further, calculation and judgment can be performed on the output result, ensuring the accuracy of the semantic segmentation model and reducing the labor cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic flowchart of an embodiment of the semantic segmentation model training method provided by the present invention;
[0055] Figure 2 For the present invention Figure 1 It is a schematic flowchart of an embodiment of step S102 in the present invention;
[0056] Figure 3 For the present invention Figure 2 It is a schematic flowchart of an embodiment of step S203 in the present invention;
[0057] Figure 4 For the present invention Figure 3 It is a schematic flowchart of an embodiment of step S301 in the present invention;
[0058] Figure 5 For the present invention Figure 3 It is a schematic flowchart of an embodiment of step S302 in the present invention;
[0059] Figure 6 For the present invention Figure 1 It is a schematic flowchart of an embodiment of step S103 in the present invention;
[0060] Figure 7Schematic structural diagram of an embodiment of the model visualization result provided by the present invention;
[0061] Figure 8 Schematic structural diagram of an embodiment of the semantic segmentation model training device provided by the present invention;
[0062] Figure 9 Schematic structural diagram of an embodiment of the electronic device provided by the present invention. Detailed implementation manners
[0063] The preferred embodiments of the present invention will be specifically described below with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0064] As Figure 1 shown, a specific embodiment of the present invention discloses a semantic segmentation model training method, including:
[0065] S101. Obtain image data on the street and set an initial semantic segmentation model;
[0066] S102. Train the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, pseudo-labels, and a set of key edge regions;
[0067] S103. Calculate weights according to the initial mask, pseudo-labels, and the set of key edge regions to obtain the total model loss;
[0068] S104. When the total model loss is greater than a preset total loss threshold, expand the image data according to the pseudo-labels, and train the target semantic segmentation model according to the expanded image data until the total model loss is not greater than the preset total loss threshold, and determine that the model training is completed.
[0069] It should be understood that: the manner of obtaining the image data in step S101 can be images of each street or other places obtained by an image acquisition device, or historical stored images called from a storage medium.
[0070] In a specific embodiment of the present invention, image acquisition devices can be used to take pictures at various corners of the street, covering various different environmental conditions. By taking pictures at different time periods and in different environments (such as day, night, sunny, rainy, etc.), the diversity and representativeness of the data are ensured. The captured data will be used for subsequent semantic segmentation analysis. An initial semantic segmentation model can also be set up. The initial semantic segmentation model can be the PAF-MSM model, that is, a model constructed by the Perception-Action-Feedback loop (PAF) and a traditional neural network (MSM). To improve the accuracy of semantic segmentation of the initial semantic segmentation model, the initial semantic segmentation model can be trained with image data to obtain the output results of the initial semantic segmentation model and the trained target semantic segmentation model. Among them, the output results include the initial mask, pseudo-labels, and a set of key edge regions. To train the model separately, the initial mask is processed to screen out high-confidence regions for generating pseudo-labels . For each pixel i , if its predicted maximum confidence Max( ) is higher than the set threshold τ , then its predicted category is used as the pseudo-label , otherwise the pixel is ignored. The gradient of the training set can also be calculated to obtain the set of key edge regions Edge
[0071] . To determine whether the target semantic segmentation model meets the requirements, weight calculations can be performed based on the initial mask, pseudo-labels, and the set of key edge regions to obtain the weights for each calculation, so that the total model loss can be calculated. Then, based on the judgment result of the total model loss and the preset total loss threshold, it can be determined whether the target semantic segmentation model meets the training requirements. Specifically, if the total model loss is greater than the preset total loss threshold, the image data can be augmented according to the pseudo-labels, and the target semantic segmentation model can be trained with the augmented image data until the total model loss is not greater than the preset total loss threshold, determining that the model training is completed. Then, the image data can be semantically segmented through the trained semantic segmentation model
[0072] Compared with the prior art, the present embodiment provides a method for obtaining image data on a street and setting an initial semantic segmentation model; training the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, pseudo-labels, and a set of key edge regions; calculating weights according to the initial mask, pseudo-labels, and the set of key edge regions to obtain a total model loss; when the total model loss is greater than a preset total loss threshold, augmenting the image data according to the pseudo-labels, and training the target semantic segmentation model according to the augmented image data until the total model loss is not greater than the preset total loss threshold, determining that the model training is completed; the present invention provides an initial semantic segmentation model, which can process each image without manually marking and processing each image; further, it can also perform calculation and judgment on the output result, ensuring the accuracy of the semantic segmentation model and reducing the labor cost.
[0073] In some embodiments of the present invention, as Figure 2 shown, step S102 includes:
[0074] S201. Label the image data to obtain a data set;
[0075] S202. Divide the data set to obtain a training set;
[0076] S203. Input the training set into the initial semantic segmentation model for training to obtain a target semantic segmentation model and an output result.
[0077] In a specific embodiment of the present invention, after the image data is collected, the image data can be labeled. For example, label road scene images such as pedestrians, cars, traffic lights, and zebra crossings, so that the labeled image data can be obtained. All the image data can be combined into a data set, and then the data set can be divided according to a preset ratio. For example, if the data set has 20,000 road scene pictures and the preset ratio can be 8:1:1, then the corresponding number of training sets, validation sets, and test sets can be obtained. Then, the image data in the training set can be input into the initial semantic segmentation model to train the initial semantic segmentation model, so as to obtain a target semantic segmentation model and an output result.
[0078] In some embodiments of the present invention, the initial semantic segmentation model includes a feature extraction network, a perception layer, an action layer, and a feedback layer. As Figure 3 shown, step S203 includes:
[0079] S301. Input the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain a feature map;
[0080] S302. Input the feature map into the perception layer for key region extraction to obtain key region features;
[0081] S303. Input the feature map and the key region features into the action layer for mask extraction and dynamic optimization to obtain an initial mask and a dynamic weight map;
[0082] S304. Input the initial mask and the dynamic weight map into the feedback layer for high-confidence region screening to obtain the output result.
[0083] In a specific embodiment of the present invention, the initial semantic segmentation model may include a feature extraction network, a perception layer, an action layer, and a feedback layer. The image data of the training set can be input into the feature extraction network for layer-by-layer sampling to obtain a feature map. Specifically, in some embodiments of the present invention, the feature extraction network includes a hybrid frequency-domain convolutional layer and a multi-scale attention residual module; as Figure 4 shown, step S301 includes:
[0084] S401. Input the image data into the hybrid frequency-domain convolutional layer for weight calculation to obtain the adaptive weight of each feature channel;
[0085] S402. Determine the convolution mode corresponding to each feature channel according to the adaptive weight; the convolution mode with an adaptive weight greater than the preset weight threshold is a traditional convolution; the convolution mode with an adaptive weight not greater than the preset weight threshold is a frequency-domain convolution;
[0086] S403. Extract features from the image data according to the convolution mode to obtain the feature images of each feature channel;
[0087] S404. Perform multi-scale feature fusion and adjustment on the feature images according to the multi-scale attention residual module to obtain the feature map.
[0088] In a specific embodiment of the present invention, the feature extraction network may include a hybrid frequency-domain convolutional layer and a multi-scale attention residual module. The feature extraction network adopts a hybrid frequency-domain convolution (HFDC) neural network structure and simultaneously integrates a multi-scale attention residual module (MSARM) with custom different convolution scale features to enhance the feature extraction ability. Among them, the hybrid frequency-domain convolution (HFDC) optimizes the model performance by adaptively selecting traditional convolution and frequency-domain convolution according to the feature complexity and hierarchical requirements. In the shallow layer, traditional convolution is suitable for processing local features and has high computational efficiency; while in the deep layer, frequency-domain convolution effectively compresses the computational amount through frequency-domain transformation and captures long-range dependencies, overcoming the limitations of traditional convolution. The HFDC module can adaptively switch the convolution mode at different layers, improving the feature expression while optimizing the computational efficiency. Specifically, before each convolution, the HFDC module calculates the adaptive weight according to the information of the input image data , and the calculation is as shown in formula (1):
[0089] = (Net( ))(1)
[0090] In the formula, Net( ) is the network calculation result based on the input image data , and is the sigmoid function that limits the output within the range of [0, 1].
[0091] This weight is obtained through network learning and determines whether to use traditional convolution or frequency-domain convolution. If the complexity of the feature map is low (such as shallow features), traditional convolution is selected; if the complexity of the feature map is high (such as deep features), then frequency-domain convolution is selected. Through this mechanism, the HFDC module can automatically select the convolution mode at different levels to ensure that each feature channel of the image data can perform feature extraction in the best way, so that the feature images of each feature channel can be obtained. Then, the multi-scale attention residual module can perform multi-scale feature fusion and adjustment on the feature images to obtain the feature map. Among them, the MSARM module (Multi-Scale Attention Residual Module) aims to enhance the feature extraction ability through multi-scale feature fusion and dynamic weighting mechanism and is suitable for processing the feature expression of complex scenes. The input feature and the multi-scale features a1 , a2 , a3 are used as the inputs of the module, providing semantic information under different receptive fields respectively. In the first step, the multi-scale features are integrated through the Feature Fusion module to obtain the preliminary fused feature . This fusion process can be achieved by concatenation or weighted summation to ensure the synergistic effect of information at different scales in subsequent processing. The fused feature then enters the pooling layer for spatial compression to enhance the perception ability of global features. After pooling, the feature is reduced in dimension through a 1×1 convolutional layer to reduce the computational cost and introduce preliminary non-linearity. Then, the feature passes through another 1×1 convolution layer again to adapt to the processing requirements of the next stage. An important link is the Weighted Feature Fusion mechanism, which fuses the processed feature with the original input feature . Through the attention mechanism or dynamic weight adjustment, this module can identify the importance of features, thereby optimizing the fusion effect. The fused feature is further normalized and non-linearly enhanced through Batch Normalization and the activation function (ReLU). Finally, the module introduces a Residual Connection, directly connecting the input feature Introduce the output to alleviate the vanishing gradient problem and retain the information of the original features. The final output features after fusion synthesize multi-scale, global context and the original input features, thus obtaining a feature map that not only enriches the feature expression but also maintains the data flow and consistency. Through this process, the MSARM module has strong adaptability and expressiveness in complex feature extraction tasks.
[0092] Furthermore, the MSM network consists of a feature extraction network MS and a weighted fusion operation, where all convolutions are replaced by a mixed-frequency convolution module. The MS extraction network is a network that fuses multi-scale and multi-dimensions. The cross-feature fusion part performs multiple dilated convolution operations on the deep information, greatly optimizing and balancing the receptive field, and finally performing a concat operation with the shallow features to output the result. For example, the size of the input image data is 1024×1024×C, where C represents the number of channels, and for an RGB image, C = 3. First, the input is fed into three parallel branches, which respectively use 1x1, 3x3, and 5x5 mixed-frequency convolutions for feature extraction. The 1×1 convolution retains the spatial resolution of the image and at the same time realizes the linear combination of channel features, and the output is ; The 3×3 and 5×5 convolutions provide supplements in extracting small-scale and large-scale features, and the outputs are and . Then, these outputs together with the input of the previous layer are input into the multi-convolution scale feature attention module (MSARM), where the weighted fusion of multi-scale features is performed through the attention mechanism to obtain the output , and at this time, the spatial resolution of 1024×1024 is still maintained. Subsequently, the backbone structure of the network starts to perform multi-level processing on the features. The features after the first layer of MSARM are downsampled (convolution and pooling operations with a 2×2 stride), and the resolution becomes 512×512. These downsampled features are extracted by another set of parallel 1×1, 3×3, and 5×5 convolution kernels and then fed into the MSARM module for feature fusion to obtain medium-scale features . Next, these features are downsampled again (the resolution is halved to 256×256), and the foregoing operations are repeated to extract features to obtain deeper multi-scale fusion results , , .
[0093] After the deep feature extraction is completed, the embodiments of the present invention adopt a weighted fusion method for feature maps at different levels, and the size of the weight reflects the contribution degree of the features at this layer. The weighted fusion is carried out in the following manner:
[0094] (2)
[0095] Among them, is the feature map of the i th layer, N is the number of pixel points, is the adaptive weight of the feature map of this layer, is the finally fused feature map. Through this weighted summation, the model can dynamically adjust the importance of different semantic levels.
[0096] In some embodiments of the present invention, as Figure 5 shown, step S302 includes:
[0097] S501. Input the output features of each feature channel of the feature map into the perception layer for global statistics to obtain the global response intensity of each feature channel;
[0098] S502. Determine the feature channels with global response intensity greater than the preset intensity threshold as key feature channels, and determine the features of the key feature channels as local features;
[0099] S503. Calculate the global features according to the global average pooling layer for the global context information and local features;
[0100] S504. Fuse the local features and the global features to obtain enhanced features;
[0101] S505. Screen the enhanced features based on the dynamic gating mechanism to obtain key region features.
[0102] In the specific embodiments of the present invention, the perception layer is a key part of the entire MSM network, aiming to extract key regions from the high-dimensional features output by the feature extraction network to support the operations of the subsequent action layer and feedback module. The specific steps can be: for the output features of each feature channel of the output feature map of the feature extraction network MS F∈ carry out global statistical analysis to calculate the importance of each feature channel. Through global average pooling (Global Average Pooling, GAP), calculate the global response intensity of each feature channel , and the calculation is as shown in formula (2):
[0103] (2)
[0104] In the formula, is the importance score of channel c , reflecting the contribution degree of this feature channel to the target task, C is the number of feature channels, and H and W are respectively the height and width of the feature map of the feature channel.
[0105] Then, the global response intensity of each feature channel can be compared with a preset intensity threshold T and, if it is higher than the preset intensity threshold T , then this feature channel is considered to contain information about the key region, that is, the feature channels satisfying c ∈{ c | >T} are retained. Feature channels with global response intensity greater than the preset intensity threshold can be determined as key feature channels, completing the selection of key feature channels. Then, the perception layer can further analyze the spatial distribution of features, highlighting the key region and suppressing the background by calculating the attention weight of each pixel point N . The weight calculation of spatial attention is shown in Equation (3):
[0106] (3)
[0107] wherein, and b are learnable parameters, is an activation function (Sigmoid), and represents the features of the key feature channel. If , it is considered that the corresponding pixel belongs to the key region, that is, the local feature, and the pixel selection condition is .
[0108] To further enhance the recognition of the key region, the perception layer combines global context information. The global feature is calculated through the global average pooling layer GAP operation, and the calculation is shown in Equation (4):
[0109] (4)
[0110] wherein, is the feature map of the local feature.
[0111] Then, the local feature and the global feature can be fused to obtain the enhanced feature , and the calculation is shown in Equation (5):
[0112] (5)
[0113] This fusion process ensures that the local feature is more accurate under the guidance of global semantic consistency, further optimizing the feature expression ability. Finally, the perception layer screens the enhanced feature through a dynamic gating mechanism, adjusting the information flow in the feature channels and spatially. The calculation of the dynamic gating weight is shown in Equation (6):
[0114] (6)
[0115] The feature representation after the dynamic gating mechanism is as follows: ,
[0116] The dynamic gating mechanism dynamically adjusts the degree of feature selection according to the task requirements. If , the corresponding features are filtered out. Through this multi-level feature analysis and screening, the sensing layer effectively locates the key regions and obtains the key region features after filtering , which lays a foundation for generating a high-quality initial mask for the subsequent action layer.
[0117] The design of the action layer is closely combined with the output of the sensing layer. By fusing the key region features and global features, generating an initial mask, and dynamically optimizing the segmentation decision, it provides support for the improvement of the overall model performance. First, the action layer receives the key region features generated by the sensing layer and the features of the original feature extraction network . The preliminary fusion features are generated through weighted fusion calculation, as shown in Equation (7):
[0118] (7)
[0119] where α ∈[0,1] is the weight parameter, which is used to control the balance between the key features and global features. This feature fusion method can not only highlight the key region information screened by the sensing layer but also ensure that the global background information is not ignored. Based on the fusion features , the action layer generates an initial mask through the weighted fusion operation of the MSM network. The decoding branch uses convolutional and upsampling modules to convert into a mask prediction with the same size as the input image, as shown in Equation (8):
[0120] (8)
[0121] In the formula, and are the convolutional weight and bias term respectively.
[0122] The action layer not only generates an initial mask , but also dynamically adjusts the segmentation decision. Using the key region mask output by the sensing layer, a weight map w ( x , y ) is generated through the dynamic weight balance formula, as shown in Equation (9):
[0123] (9)
[0124] In the formula, γ∈[0,1] is a dynamic adjustment parameter. If γ is larger, the model pays more attention to the key area , if γ is smaller, the model pays more attention to the initial global prediction , which is calculated as , where t is the current training round, is the step size adjusted in each round. Finally, the output result of the action layer is the initial mask and the dynamic weight map w ( x , y ). The initial mask provides guidance for the subsequent optimization module, and the dynamic weight map ensures the balance between the key area and other areas.
[0125] In some embodiments of the present invention, as Figure 6 shown, step S103 includes:
[0126] S601. Perform standard forward propagation and backward propagation calculations on the initial mask and the pseudo-label to obtain the first weight;
[0127] S602. Obtain the second weight and the third weight according to the key edge area set and the pseudo-label;
[0128] S603. Calculate the first weight, the second weight and the third weight to obtain the total model loss.
[0129] In a specific embodiment of the present invention, during the first calculation process, standard forward propagation and backward propagation can be performed on the input initial mask and pseudo-label, as shown in formulas (10)-(13):
[0130] (10)
[0131] (11)
[0132] (12)
[0133] (13)
[0134] In the formula, is the mask data of the th i in the initial mask is the i th pseudo-label, is the th pixel propagation data obtained by the pixel propagation method from the initial mask i , is the i th pixel, is the first weight.
[0135] In some embodiments of the present invention, step S602 includes:
[0136] Performing mask extraction on the set of key edge regions to obtain an edge region mask;
[0137] Calculating an edge region loss for the edge region mask and the pseudo-label to obtain a second weight;
[0138] Performing balance optimization on the edge region mask, the pseudo-label, and a preset weight to obtain a third weight.
[0139] In a specific embodiment of the present invention, in the second calculation, focus on optimizing the edge region loss, extract the set of key edge regions Edge, and obtain the edge region mask through the gradient method or the morphological method , and then calculate the second weight based on the pixels in the edge region through formula (14):
[0140] (14)
[0141] In the formula, is the total number of edge region pixels in the edge region mask; is the cross-entropy loss function; is the th edge value in the edge region mask; is the th pseudo-label, is the second weight.
[0142] In the third weight calculation, re-regress the balance optimization of the global loss, and calculate the third weight by combining the loss weights of ordinary pixels and key regions, as shown in formula (15):
[0143] (15)
[0144] In the formula, is the weight, calculated as: , if i / 1, is the weight gain coefficient of the edge region (e.g., = 0.2), and the weight of the loss function balances the influence of edge pixels and ordinary pixels to ensure the global optimization effect.
[0145] Then, the first weight, the second weight, and the third weight can be calculated to obtain the total model loss, as calculated in formula (16):
[0146] L= (16)
[0147] Among them, are respectively shown in Formulas (17)-(19):
[0148] (17)
[0149] (18)
[0150] (19)
[0151] In the formula, L is the total loss of the model, T is the total number of training calculations, t is the current round.
[0152] Among them, the above three weight calculations can be carried out simultaneously. Through the three weight calculations, the optimization of the edge region and the global loss is refined layer by layer, the learning rate and the weight are adjusted formulaically, so that the key region receives more attention, and at the same time, the overall performance of the model is balanced and improved. The combination of the Perception-Action-Feedback loop (PAF) and the traditional neural network (MSM) jointly promotes the improvement of the semantic segmentation performance through their interaction. The PAF loop enables the network to make more intelligent decisions and optimizations on the basis of the traditional neural network, overcomes the dependence on manually designed data augmentation and training steps in the traditional method, and provides a new solution for the automated image segmentation task.
[0153] Furthermore, it can be judged whether the total loss of the model is greater than the preset total loss threshold. If not, it means that the semantic segmentation model has achieved an ideal effect, and the target semantic segmentation model can be determined as the trained model. If so, the pseudo-labels can be expanded to the training set to obtain a new training set, and the target semantic segmentation model can be trained according to the new training set until the total loss of the model is not greater than the preset total loss threshold, and it is determined that the model training is completed. Specifically, it can be achieved by dynamically adjusting the data samples and the training set weights through the initial mask so as to improve the segmentation ability of the model for complex scenes. After determining the pseudo-labels it is possible to store the feature maps of the pseudo-labels in the training set for expanding the training set, which significantly improves the data utilization rate. It is also possible to calculate the loss weights of the pixels in these regions in the key edge region set Edgez After that, ensure that these regions account for a larger proportion in the total loss. Embodiments of the present invention also feedback these difficult samples back to the training set to obtain a new training set. In the next round of training, increase the appearance frequency of these "pictures with blurred edges" and reduce the training of simple pictures. Through multiple calculations and optimizations of key samples, the loop optimization gradually improves the model's processing ability for edge regions and complex backgrounds, and balances the loss at the global level.
[0154] Further, after the target semantic segmentation model is trained, the target semantic segmentation model can be evaluated. For example, a comparative experiment is conducted between the target semantic segmentation model PAF-MSM (ours) and classical semantic segmentation networks UNet, Deeplabv3+, and Segnet to prove the effectiveness of the proposed network. The visualization results of the model are as Figure 7 shown. According to Figure 7 it can be seen that the IoU, Recall, and Precision of the PAF-MSM model are all higher than those of other segmentation methods. The PAF-MSM model has achieved a very good segmentation effect, and the PAF-MSM method is superior to UNet, Deeplabv3+, and Segnet in the segmentation accuracy of industrial road scenes.
[0155] To better implement the semantic segmentation model training method in the embodiments of the present invention, correspondingly, embodiments of the present invention also provide a semantic segmentation model training device, as Figure 8 shown. The semantic segmentation model training device 800 includes:
[0156] A data acquisition module 801, configured to acquire image data on the street and set an initial semantic segmentation model;
[0157] A model training module 802, configured to train the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, a pseudo label, and a set of key edge regions;
[0158] A weight calculation module 803, configured to calculate weights according to the initial mask, the pseudo label, and the set of key edge regions to obtain a total model loss;
[0159] A result judgment module 804, configured to, when the total model loss is greater than a preset total loss threshold, expand the image data according to the pseudo label, and train the target semantic segmentation model according to the expanded image data until the total model loss is not greater than the preset total loss threshold, and determine that the model training is completed.
[0160] The semantic segmentation model training device 800 provided by the above embodiments can implement the technical solutions described in the above embodiments of the semantic segmentation model training method. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above embodiments of the semantic segmentation model training method, which will not be elaborated here.
[0161] As Figure 9 shown, the present invention also correspondingly provides an electronic device 900. The electronic device 900 includes a processor 901, a memory 902, and a display 903. Figure 9 Only some components of the electronic device 900 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0162] In some embodiments, the memory 902 can be an internal storage unit of the electronic device 900, such as the hard disk or memory of the electronic device 900. In other embodiments, the memory 902 can also be an external storage device of the electronic device 900, such as a plug-in hard disk equipped on the electronic device 900, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0163] Furthermore, the memory 902 can also include both the internal storage unit and the external storage device of the electronic device 900. The memory 902 is used to store the application software installed on the electronic device 900 and various types of data.
[0164] In some embodiments, the processor 901 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 902 or process data, such as the semantic segmentation model training method in the present invention.
[0165] In some embodiments, the display 903 can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 903 is used to display the information of the electronic device 900 and to display a visual user interface. The components 901 - 903 of the electronic device 900 communicate with each other through a system bus.
[0166] In some embodiments of the present invention, when the processor 901 executes the semantic segmentation model training program in the memory 902, the following steps can be implemented:
[0167] Obtain image data on the street and set an initial semantic segmentation model;
[0168] Train an initial semantic segmentation model based on the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, pseudo-labels, and a set of key edge regions;
[0169] Calculate the weights based on the initial mask, pseudo-labels, and the set of key edge regions to obtain the total model loss;
[0170] When the total model loss is greater than the preset total loss threshold, augment the image data according to the pseudo-labels, and train the target semantic segmentation model based on the augmented image data until the total model loss is not greater than the preset total loss threshold, and determine that the model training is completed.
[0171] It should be understood that when the processor 901 executes the semantic segmentation model training program in the memory 902, in addition to the above functions, other functions can also be implemented. For specific details, please refer to the description of the corresponding method embodiments above.
[0172] Furthermore, the embodiments of the present invention do not specifically limit the type of the electronic device 900 mentioned. The electronic device 900 can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, or other portable electronic devices. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running IOS, android, microsoft, or other operating systems. The above-mentioned portable electronic devices can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (such as a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 900 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (such as a touch panel).
[0173] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the method steps or functions provided by the above-mentioned method embodiments for training a semantic segmentation model can be implemented.
[0174] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0175] The above has introduced in detail the method and device for training a semantic segmentation model provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for training a semantic segmentation model, characterized in that Including: Obtain image data on the street and set an initial semantic segmentation model; Train the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, pseudo-labels, and a set of key edge regions; Perform weight calculation according to the initial mask, the pseudo-labels, and the set of key edge regions to obtain a total model loss; When the total model loss is greater than a preset total loss threshold, augment the image data according to the pseudo-labels, and train the target semantic segmentation model according to the augmented image data until the total model loss is not greater than the preset total loss threshold, and determine that the model training is completed; The initial semantic segmentation model includes a feature extraction network, a perception layer, an action layer, and a feedback layer; The feature extraction network includes a hybrid frequency-domain convolutional layer and a multi-scale attention residual module; Input the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain feature maps, including: Input the image data into the hybrid frequency-domain convolutional layer for weight calculation to obtain an adaptive weight for each feature channel; Determine the convolutional mode corresponding to each feature channel according to the adaptive weight; the convolutional mode with the adaptive weight greater than the preset weight threshold is a traditional convolution; the convolutional mode with the adaptive weight not greater than the preset weight threshold is a frequency-domain convolution; Extract features from the image data according to the convolutional mode to obtain feature images for each feature channel; Perform multi-scale feature fusion and adjustment on the feature images according to the multi-scale attention residual module to obtain feature maps; The process of layer-by-layer sampling by the feature extraction network includes: Send the image data into three parallel branches, and use hybrid frequency convolution for feature extraction in the parallel branches to obtain multi-scale features; Perform weighted fusion of the multi-scale features through an attention mechanism to obtain a weighted fusion feature; The weighted fusion feature after passing through the first layer of MSARM is downsampled, and the features after this downsampling are extracted by another group of parallel convolutional kernels and then sent into the MSARM module for feature fusion to obtain medium-scale features; Repeat the above feature extraction operations to obtain a deeper multi-scale fusion result.
2. The semantic segmentation model training method according to claim 1, wherein The training of the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result includes: Label the image data to obtain a data set; Divide the data set to obtain a training set; Input the training set into the initial semantic segmentation model for training to obtain a target semantic segmentation model and an output result.
3. The semantic segmentation model training method according to claim 2, characterized in that, The inputting the training set into the initial semantic segmentation model for training to obtain a target semantic segmentation model and an output result includes: Input the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain feature maps; Input the feature maps into the perception layer for key region extraction to obtain key region features; Input the feature map and the key region features into the action layer for mask extraction and dynamic optimization to obtain an initial mask and a dynamic weight map; Input the initial mask and the dynamic weight map into the feedback layer for high-confidence region screening to obtain an output result.
4. The semantic segmentation model training method according to claim 1, characterized in that The weight calculation based on the initial mask, the pseudo-label, and the set of key edge regions to obtain the total model loss includes: Perform standard forward and backward propagation calculations on the initial mask and the pseudo-label to obtain a first weight; Obtain a second weight and a third weight according to the set of key edge regions and the pseudo-label; Calculate the first weight, the second weight, and the third weight to obtain the total model loss.
5. The semantic segmentation model training method according to claim 4, wherein The obtaining of the second weight and the third weight according to the set of key edge regions and the pseudo-label includes: Perform mask extraction on the set of key edge regions to obtain an edge region mask; Calculate the edge region loss for the edge region mask and the pseudo-label to obtain a second weight; Perform balance optimization on the edge region mask, the pseudo-label, and a preset weight to obtain a third weight.
6. The semantic segmentation model training method according to claim 1, wherein The inputting of the feature map into the perception layer for key region extraction to obtain key region features includes: Input the output features of each feature channel of the feature map into the perception layer for global statistics to obtain the global response intensity of each feature channel; Determine the feature channels with the global response intensity greater than a preset intensity threshold as key feature channels, and determine the features of the key feature channels as local features; Calculate the global features based on the global average pooling layer for global context information and the local features; Fuse the local features and the global features to obtain enhanced features; Screen the enhanced features based on a dynamic gating mechanism to obtain key region features.
7. The semantic segmentation model training method according to claim 4, wherein The calculation process of the first weight is as follows: In the formula, is the initial mask in the i th mask data, is the i th pseudo label, is the initial mask obtained the i th pixel propagation data through the pixel propagation method, is the i th pixel, is the first weight, is the cross-entropy loss function, N is the number of pixels.
8. The semantic segmentation model training method according to claim 5, characterized in that The calculation process of the second weight is as follows: Wherein, is the total number of edge region pixels in the edge region mask; is the cross-entropy loss function; is the th edge value in the edge region mask; is the th pseudo-label.
9. A semantic segmentation model training device, characterized in that, Including: A data acquisition module for acquiring image data on the street and setting an initial semantic segmentation model; A model training module for training the initial semantic segmentation model according to the image data to obtain a target semantic segmentation model and an output result; the output result includes an initial mask, a pseudo-label, and a set of key edge regions; A weight calculation module for calculating weights according to the initial mask, the pseudo-label, and the set of key edge regions to obtain the total model loss; A result judgment module for, when the total model loss is greater than a preset total loss threshold, expanding the image data according to the pseudo-label and training the target semantic segmentation model according to the expanded image data until the total model loss is not greater than the preset total loss threshold, and determining that the model training is completed; The initial semantic segmentation model includes a feature extraction network, a perception layer, an action layer, and a feedback layer; The feature extraction network includes a hybrid frequency-domain convolutional layer and a multi-scale attention residual module; Input the image data of the training set into the feature extraction network for layer-by-layer sampling to obtain a feature map, including: Input the image data into the hybrid frequency-domain convolutional layer for weight calculation to obtain the adaptive weight of each feature channel; Determine the convolutional mode corresponding to each feature channel according to the adaptive weight; the convolutional mode with the adaptive weight greater than the preset weight threshold is traditional convolution; the convolutional mode with the adaptive weight not greater than the preset weight threshold is frequency-domain convolution; Extract features from the image data according to the convolutional mode to obtain the feature images of each feature channel; Perform multi-scale feature fusion and adjustment on the feature images according to the multi-scale attention residual module to obtain a feature map; The process of the feature extraction network performing layer-by-layer sampling includes: Send the image data into three parallel branches, and the parallel branches use hybrid frequency convolution for feature extraction to obtain multi-scale features; Perform weighted fusion of the multi-scale features through an attention mechanism to obtain weighted fusion features; The weighted fusion features after the first layer of MSARM are downsampled, and the features after this downsampling are extracted by another group of parallel convolutional kernels and then sent into the MSARM module again for feature fusion to obtain medium-scale features; Repeat the above feature extraction operations to obtain a deeper multi-scale fusion result.
Citation Information
Patent Citations
Weak supervision image semantic segmentation method, system and device and storage medium
CN116309653A
Colorectal cancer image segmentation method and system, and medium
CN119741499A