Semantic Segmentation Method, Device, Electronic Device and Storage Medium

By introducing decoder network and auxiliary head network in neural network model training, focusing on training results and losses in the process, the model is optimized to solve the problem of ignoring intermediate states and changes in neural network models during training, improving the accuracy, stability and generalization capabilities of the model.

CN118608781BActive Publication Date: 2025-06-10TIANYUN RONGCHUANG DATA TECH BEIJING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410626387.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-06-10
Estimated Expiration
2044-05-20

AI Technical Summary

Technical Problem

During the training process, existing neural network models only focus on the loss value of the final output, ignore intermediate states and changes, resulting in the inability to discover and correct potential problems in time, reducing the accuracy, stability and generalization capabilities of the model.

Method used

By obtaining the training data set, the decoder network and auxiliary head network in the initial neural network model are trained, and the result loss function and the process loss function are obtained respectively, and the model is optimized to obtain the target semantic segmentation network model.

Benefits of technology

By paying attention to the training results and losses in the process, the model is comprehensively optimized, and the accuracy, stability and generalization capabilities of the neural network model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608781B_ABST
    Figure CN118608781B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a semantic segmentation method, apparatus, electronic device, and storage medium, which are applied to the technical field of semantic segmentation, and can solve the problem that in the training process of a neural network model, only the loss value of the final output is concerned, while the intermediate state and changes of the model during the training process are ignored. The method includes: obtaining a training data set; training a decoder network in an initial neural network model through the training data set to obtain a first result loss function; training an auxiliary head network in the initial neural network model through the training data set to obtain a first process loss function; optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model; inputting the data to be measured into the target semantic segmentation network model to obtain a target semantic segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of semantic segmentation, and in particular, to a semantic segmentation method, apparatus, electronic device, and storage medium. Background Art

[0002] Semantic segmentation is an important task in the field of computer vision. Its goal is to label each pixel in an image with a label corresponding to a certain object category. That is to say, the task of semantic segmentation is to identify different objects in an image and distinguish them from the background, so as to achieve a detailed understanding of the image content. Currently, in the training process of most neural network models, they often only focus on the final output loss value, while ignoring the intermediate states and changes of the model during the training process, lacking monitoring and evaluation of the intermediate process. This may lead to the inability to timely discover and correct potential problems that occur during the training of the model, thus unable to perform necessary interventions and adjustments in the early stage of training, reducing the accuracy, stability, and generalization ability of the neural network model. Summary of the Invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present application provide a semantic segmentation method, apparatus, electronic device, and storage medium to solve the problem that in the training process of a neural network model, it often only focuses on the final output loss value while ignoring the intermediate states and changes of the model during the training process.

[0004] To achieve the above object, the technical solutions provided by the embodiments of the present application are as follows:

[0005] In a first aspect, embodiments of the present application provide a semantic segmentation method. The semantic segmentation method includes: obtaining a training data set, where the training data set at least includes: a plurality of training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling;

[0006] Training a decoder network in an initial neural network model through the training data set to obtain a first result loss function;

[0007] Training an auxiliary head network in the initial neural network model through the training data set to obtain a first process loss function;

[0008] Optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model;

[0009] Inputting the data to be measured into the target semantic segmentation network model to obtain a target semantic segmentation result, where the target semantic segmentation result is used to indicate the classification result of the data to be measured.

[0010] As an alternative implementation, in the first aspect of the embodiments of the present application, optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain the target semantic segmentation network model includes:

[0011] Calculating the sum of the first result loss function and the first process loss function;

[0012] Based on the sum, optimizing the initial neural network model through the backpropagation algorithm to obtain the target semantic segmentation network model.

[0013] As an alternative implementation, in the first aspect of the embodiments of the present application, optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain the target semantic segmentation network model includes:

[0014] Optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a first neural network model;

[0015] Training the decoder network in the first neural network model through the training data set to obtain a second result loss function;

[0016] Training the auxiliary head network in the first neural network model through the training data set to obtain a second process loss function;

[0017] Optimizing the first neural network model according to the second result loss function and the second process loss function, and repeating the above steps until the second result loss function and the second process loss function are in a stable state to obtain the target semantic segmentation network model.

[0018] As an alternative implementation, in the first aspect of the embodiments of the present application, inputting the data to be measured into the target semantic segmentation network model to obtain the target semantic segmentation result includes:

[0019] Inputting the data to be measured into the encoder network of the target semantic segmentation network model to obtain the feature map corresponding to the data to be measured;

[0020] Inputting the feature map into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result.

[0021] As an alternative implementation, in the first aspect of the embodiments of the present application, inputting the feature map into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result includes:

[0022] Input the feature map into the decoder network of the target semantic segmentation network model to obtain a target vector corresponding to the feature map, where each element in the target vector is used to indicate the classification result of each pixel in the data to be measured;

[0023] Perform masking processing on the target vector through a preset masking matrix to obtain the target semantic segmentation result.

[0024] As an optional implementation manner, in the first aspect of the embodiments of the present application, the obtaining of the training data set includes:

[0025] Obtain the multiple training samples;

[0026] Perform pre-labeling on the multiple training samples respectively to obtain a training segmentation result corresponding to each training sample;

[0027] Extract features from the multiple training samples through the encoder network in the initial neural network model to obtain a feature map corresponding to each training sample;

[0028] Perform dimension and size adjustment on the feature map corresponding to each training sample through the neck connector network in the initial neural network model;

[0029] Determine the multiple training samples, as well as the adjusted feature map and training segmentation result corresponding to each training sample as the training data set.

[0030] As an optional implementation manner, in the first aspect of the embodiments of the present application, the training of the decoder network in the initial neural network model through the training data set to obtain a first result loss function includes:

[0031] Input the adjusted feature map corresponding to each training sample into the decoder network of the initial neural network model to obtain a first result training vector, where the first result training vector is used to indicate the prediction result of each pixel point in the multiple training samples;

[0032] Obtain the first result loss function through the cross-entropy algorithm according to the first result training vector and the training segmentation result in the training data set.

[0033] In a second aspect, an embodiment of the present application provides a semantic segmentation device, where the semantic segmentation device includes: an acquisition module, configured to acquire a training data set, where the training data set at least includes: multiple training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling;

[0034] A processing module, configured to train a decoder network in an initial neural network model through the training data set to obtain a first result loss function;

[0035] The processing module is further configured to train an auxiliary head network in the initial neural network model through the training data set to obtain a first process loss function;

[0036] The processing module is further configured to optimize the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model;

[0037] The processing module is further configured to input the data to be measured into the target semantic segmentation network model to obtain a target semantic segmentation result, where the target semantic segmentation result is used to indicate the classification result of the data to be measured.

[0038] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module is specifically configured to calculate the sum of the first result loss function and the first process loss function;

[0039] The processing module is specifically configured to optimize the initial neural network model based on the sum through a backpropagation algorithm to obtain the target semantic segmentation network model.

[0040] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module is specifically configured to optimize the initial neural network model according to the first result loss function and the first process loss function to obtain a first neural network model;

[0041] The processing module is specifically configured to train a decoder network in the first neural network model through the training data set to obtain a second result loss function;

[0042] The processing module is specifically configured to train an auxiliary head network in the first neural network model through the training data set to obtain a second process loss function;

[0043] The processing module is specifically configured to optimize the first neural network model according to the second result loss function and the second process loss function, and repeat the above steps until the second result loss function and the second process loss function are in a stable state to obtain the target semantic segmentation network model.

[0044] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module is specifically configured to input the data to be measured into an encoder network of the target semantic segmentation network model to obtain a feature map corresponding to the data to be measured;

[0045] The processing module is specifically configured to input the feature map into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result.

[0046] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module is specifically configured to input the feature map into the decoder network of the target semantic segmentation network model to obtain a target vector corresponding to the feature map, and each element in the target vector is used to indicate the classification result of each pixel in the data to be measured;

[0047] The processing module is specifically configured to perform masking processing on the target vector through a preset masking matrix to obtain the target semantic segmentation result.

[0048] As an optional implementation manner, in the second aspect of the embodiments of the present application, the obtaining module is specifically configured to obtain the multiple training samples;

[0049] The processing module is specifically configured to perform pre-labeling on the multiple training samples respectively to obtain a training segmentation result corresponding to each training sample;

[0050] The processing module is specifically configured to perform feature extraction on the multiple training samples through the encoder network in the initial neural network model to obtain a feature map corresponding to each training sample;

[0051] The processing module is specifically configured to perform dimension and size adjustment on the feature map corresponding to each training sample through the neck connector network in the initial neural network model;

[0052] The processing module is specifically configured to determine the multiple training samples, and the adjusted feature map and training segmentation result corresponding to each training sample as the training data set.

[0053] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module is specifically configured to input the adjusted feature map corresponding to each training sample into the decoder network of the initial neural network model to obtain a first result training vector, and the first result training vector is used to indicate the prediction result of each pixel point in the multiple training samples;

[0054] The processing module is specifically configured to obtain the first result loss function according to the first result training vector and the training segmentation result in the training data set through the cross-entropy algorithm.

[0055] In a third aspect, an embodiment of the present application provides an electronic device, and the electronic device includes:

[0056] A memory storing executable program code;

[0057] A processor coupled to the memory;

[0058] The processor calls the executable program code stored in the memory and executes the semantic segmentation method in the first aspect of the embodiments of the present application.

[0059] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium that stores a computer program, and the computer program causes a computer to execute the semantic segmentation method in the first aspect of the embodiments of the present application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, an optical disc, etc.

[0060] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a computer, it causes the computer to execute some or all of the steps of any one of the methods in the first aspect.

[0061] In a sixth aspect, an embodiment of the present application provides an application publishing platform, and the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, it causes the computer to execute some or all of the steps of any one of the methods in the first aspect.

[0062] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0063] The embodiments of the present application provide a semantic segmentation method, apparatus, electronic device, and storage medium. A training data set is obtained, and the training data set at least includes: a plurality of training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling; through the training data set, the decoder network in the initial neural network model is trained to obtain a first result loss function; through the training data set, the auxiliary head network in the initial neural network model is trained to obtain a first process loss function; according to the first result loss function and the first process loss function, the initial neural network model is optimized to obtain a target semantic segmentation network model; the data to be measured is input into the target semantic segmentation network model to obtain a target semantic segmentation result, and the target semantic segmentation result is used to indicate the classification result of the data to be measured. In this solution, a decoder network and an auxiliary head network are introduced. Through training samples, both the decoder network and the auxiliary head network are trained simultaneously. The decoder network focuses on the loss of the training result, and the auxiliary head network focuses on the loss of the training process. It can not only consider the result but also consider the data in the model training process, and optimize the model by comprehensively considering the loss of each layer, which can effectively improve the accuracy, stability, and generalization ability of the neural network model. Description of the Drawings

[0064] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0065] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0066] Figure 1 is a flowchart showing a semantic segmentation method provided by an embodiment of this application Figure One ;

[0067] Figure 2 is a flowchart showing a semantic segmentation method provided by an embodiment of this application Figure Two ;

[0068] Figure 3 is a schematic diagram showing the result of a semantic segmentation method provided by an embodiment of this application;

[0069] Figure 4 is a schematic diagram showing the structure of a semantic segmentation device provided by an embodiment of this application;

[0070] Figure 5 is a schematic diagram showing the structure of an electronic device provided by an embodiment of this application. Detailed Embodiments

[0071] To be able to more clearly understand the above objects, features, and advantages of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. It should be noted that, without conflict, the embodiments of this application and the features in the embodiments can be combined with each other. Obviously, the described embodiments are some, rather than all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0072] The terms "first" and "second" etc. in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of the objects.

[0073] As used in the embodiments of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0074] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0075] Semantic segmentation is an important task in the field of computer vision. Its goal is to label each pixel in an image with a label corresponding to a certain object category. In other words, the task of semantic segmentation is to identify different objects in an image and distinguish them from the background, so as to achieve a detailed understanding of the image content. Semantic segmentation is different from image classification, which only needs to determine the category to which the entire image belongs, while the former needs to classify each pixel in the image. Semantic segmentation is also different from instance segmentation, which not only needs to distinguish different objects, but also needs to distinguish different individuals of the same category.

[0076] The main methods for achieving semantic segmentation include techniques based on deep learning, especially convolutional neural networks (CNNs). Among them, the fully convolutional network (FCN) is a commonly used structure in semantic segmentation. It can gradually restore the low-resolution feature map to the same resolution as the input image through multiple convolution and upsampling operations, so as to achieve pixel-level classification.

[0077] In recent years, with the development of deep learning technology, the accuracy and efficiency of semantic segmentation have been significantly improved. Various new network structures and optimization methods have emerged continuously, making the application of semantic segmentation more and more extensive in fields such as autonomous driving, medical image analysis, and robot vision. Generally speaking, semantic segmentation is an important research direction in the field of computer vision. It provides richer information for image understanding and analysis through pixel-level category division. With the continuous progress of technology, the application prospect of semantic segmentation will be broader.

[0078] Semantic segmentation can provide detailed image understanding: Semantic segmentation can assign labels to each pixel in an image, thus achieving a detailed understanding of the image content. This helps to identify different objects in the image and distinguish them from the background. Moreover, it has a very wide range of application fields: Since semantic segmentation can provide rich image information, it has extensive applications in multiple fields. For example, in autonomous driving, semantic segmentation can help vehicles identify roads, vehicles, pedestrians, etc.; in medical image analysis, it can help doctors identify lesion areas; in robot vision, semantic segmentation can help robots understand the environment and make corresponding decisions. At the same time, it is also combined with deep learning technology: In recent years, with the development of deep learning technology, the accuracy and efficiency of semantic segmentation have been significantly improved. Deep learning models, especially convolutional neural networks, can automatically learn image features and perform pixel-level classification, greatly enhancing the performance of semantic segmentation.

[0079] However, semantic segmentation algorithms usually have a high computational cost: The semantic segmentation task usually needs to process a large amount of pixel data and requires complex calculations to obtain accurate segmentation results. This results in a high computational cost for semantic segmentation, especially when dealing with high-resolution images. Moreover, there are certain limitations for complex scenarios: Although the development of deep learning technology has improved the performance of semantic segmentation, there are still certain limitations when dealing with complex scenarios. For example, when there are multiple occluded or overlapping objects in the image, semantic segmentation may face challenges. During the segmentation process, it is unable to distinguish different instances of the same type of target: Semantic segmentation mainly focuses on pixel-level class division and does not distinguish different instances within the same category. This means it cannot provide information about the specific location or individual identity of the target. If such information is needed, more advanced instance segmentation techniques need to be used.

[0080] Convolutional neural network is a class of feedforward neural networks that contain convolutional calculations and have a deep structure, and it is one of the representative algorithms of deep learning. It has the ability of feature representation and can perform shift-invariant classification on the input information according to its hierarchical structure. Therefore, it is also called Shift-Invariant Artificial Neural Networks (SIANN).

[0081] Convolutional neural networks have a wide range of applications in multiple fields, especially in the field of computer vision. For example, it can be used for tasks such as image recognition, object recognition, and image processing. Its working principle is to extract local features in the image through convolutional operations and combine these features through a hierarchical structure to form an overall understanding of the image. This enables convolutional neural networks to process complex image data and recognize different objects and scenes in the image. In addition, convolutional neural networks can also improve performance through optimization methods. For example, using better activation functions, optimizers, regularization methods, convolutional kernels, and data augmentation methods can all enhance the performance, convergence speed, and generalization ability of the model. In recent years, convolutional neural networks have also demonstrated their application value in other fields, such as medical tasks, autonomous driving, video analysis, game AI, etc.

[0082] Currently, during the training process of most neural network models, they often only focus on the loss value of the final output, while ignoring the intermediate states and changes of the model during the training process, lacking monitoring and evaluation of the intermediate process. This may lead to the inability to timely discover and correct potential problems that occur during the training of the model, thus unable to perform necessary interventions and adjustments in the early stage of training, reducing the accuracy, stability, and generalization ability of the neural network model.

[0083] To solve some or all of the above technical problems, the embodiments of the present application provide a semantic segmentation method, device, electronic device, and storage medium. Obtain a training data set, where the training data set at least includes: multiple training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling; through the training data set, train the decoder network in the initial neural network model to obtain a first result loss function; through the training data set, train the auxiliary head network in the initial neural network model to obtain a first process loss function; according to the first result loss function and the first process loss function, optimize the initial neural network model to obtain a target semantic segmentation network model; input the data to be measured into the target semantic segmentation network model to obtain a target semantic segmentation result, and the target semantic segmentation result is used to indicate the classification result of the data to be measured. In this solution, a decoder network and an auxiliary head network are introduced. Through training samples, both the decoder network and the auxiliary head network are trained simultaneously. The decoder network focuses on the loss of the training result, and the auxiliary head network focuses on the loss of the training process. It can not only consider the result but also the data during the model training process, and optimize the model by comprehensively considering the loss of each layer, which can effectively improve the accuracy, stability, and generalization ability of the neural network model.

[0084] As Figure 1 shown, Figure 1 is a flowchart of a semantic segmentation method provided by the embodiments of the present application. The method may include the following steps:

[0085] 101. Obtain a training data set.

[0086] In the embodiments of the present application, the training data set at least includes: a plurality of training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling.

[0087] It should be noted that when training a neural network model, the training data set used is sample data with pre-annotated results. The sample data can be image data or video data. The training segmentation result is obtained after manual annotation. Semantic annotation is performed on each sample data, that is, each object included in each picture or video is manually annotated.

[0088] 102. Train the decoder network in the initial neural network model through the training data set to obtain a first result loss function.

[0089] In the embodiments of the present application, after obtaining the training data set, the initial neural network model can be trained through the training data set. The initial neural network model includes a decoder network and an auxiliary head network. The decoder network can specifically be an upsampling neural network, which is a fully connected layer network. The model structure type of the decoder network can be a UperNet segmentation head network.

[0090] In some embodiments, UperNet is a deep learning network for image segmentation. It is improved based on the Feature Pyramid Network (FPN) and improves the segmentation performance by fusing feature information of different scales. In UperNet, the segmentation head network plays a crucial role. The segmentation head network is usually located at the top layer of the UperNet architecture and receives multi-scale feature maps from the Feature Pyramid Network as input. These feature maps contain rich context information and detailed information of different scales, which are crucial for accurate segmentation.

[0091] In the segmentation head network, a series of convolutional layers, upsampling layers, and possibly fusion operations are usually used to extract and process these feature maps. The convolutional layers are used to further extract features, and the upsampling layers are used to restore the size of the feature maps to the same size as the original image for pixel-level prediction. The fusion operation can merge feature maps of different scales to make full use of their respective advantages. Finally, the segmentation head network outputs a segmentation map with the same size as the original image, and each pixel is assigned a class label. This segmentation map can be used to represent the segmentation results of different regions in the image, thereby realizing the recognition and positioning of different objects in the image.

[0092] It should be noted that the specific structure of the segmentation head network of UperNet may vary according to different application scenarios and task requirements. In practical applications, it may be necessary to adjust and optimize the structure and parameters of the segmentation head network according to specific tasks and datasets to obtain the best segmentation performance. In summary, the segmentation head network of UperNet is a key component that uses multi-scale feature information for pixel-level prediction and segmentation, providing strong support for image segmentation tasks.

[0093] In the embodiment of the present application, the decoder network can perform upsampling according to the input sample data to obtain an output vector, where the element at each position in the vector represents the prediction result of each pixel in the input sample data. The decoder network is mainly responsible for predicting the final image segmentation result from the sample data. That is to say, the decoder network only focuses on the final segmentation result. Therefore, the loss function corresponding to the decoder network is only used to indicate the loss of the upsampling result, that is, the first result loss function is obtained.

[0094] 103. Train the auxiliary head network in the initial neural network model through the training dataset to obtain the first process loss function.

[0095] In the embodiment of the present application, the initial neural network model includes a decoder network and an auxiliary head network. The auxiliary head network can specifically be a fully convolutional network. A fully convolutional network is a special form of a convolutional neural network that completely replaces the fully connected layers in the convolutional neural network with convolutional layers. This change enables the network to accept input images of any size and output feature maps of corresponding sizes, thereby achieving pixel-level prediction. The fully convolutional network was initially proposed and applied to the image segmentation task. Its working principle is to extract features in the image through convolutional layers and then use transposed convolution (or deconvolution) to upsample the feature maps to restore them to the same size as the original image. In this way, the fully convolutional network can classify each pixel in the image, thereby achieving accurate image segmentation.

[0096] The fully convolutional network has achieved remarkable results in the field of image segmentation and is widely used in various computer vision tasks. For example, it can accurately segment different target regions in the image, such as separating objects from the background or segmenting different objects. In addition, the fully convolutional network can also be used for tasks such as semantic segmentation and instance segmentation, which require each pixel or each instance in the image to be labeled with the corresponding category. Generally speaking, the fully convolutional network is a powerful deep learning tool that has demonstrated excellent performance and broad application prospects in image segmentation and other computer vision tasks.

[0097] The fully convolutional network can be applied to inputs of any size: The fully convolutional network can accept input images of any size without cropping or scaling the images as in traditional convolutional neural networks, thus preserving the spatial information in the original images. Moreover, pixel-level classification can be achieved: The fully convolutional network can classify each pixel in the image, realizing end-to-end pixel-level segmentation, which is crucial for image segmentation tasks. At the same time, it has a certain degree of efficiency compared to convolutional neural networks: The fully convolutional network reduces the computational amount through convolutional and pooling operations, improves the processing speed, and can meet the requirements of real-time segmentation. In addition, it also has good scalability: The structure of the fully convolutional network can be easily extended to other tasks such as image classification and object detection, having good scalability.

[0098] In an embodiment of the present application, the auxiliary head network can perform upsampling based on the input sample data to obtain an output vector. Each element at each position in the vector represents the prediction result of each pixel in the input sample data. Among them, the auxiliary head network is mainly responsible for calculating the loss of the intermediate feature map during training. Since the upsampling process is a process of gradually changing a low-resolution image into a high-resolution image, there are losses in each transformation process. The above decoder network only focuses on the loss between the image before upsampling and the image after upsampling, while the auxiliary head network will focus on the loss of each layer during the upsampling process. That is to say, the auxiliary head network focuses on the segmentation process. Therefore, the loss function corresponding to the auxiliary head network can be used to indicate the loss during the upsampling process, that is, the first process loss function is obtained.

[0099] 104. Optimize the initial neural network model according to the first result loss function and the first process loss function to obtain the target semantic segmentation network model.

[0100] In an embodiment of the present application, after obtaining the first result loss function and the first process loss function through one training, the initial neural network model can be optimized. This step is actually a loop step, that is, continuously optimize the neural network model according to the loss function, then perform training after optimization, and then obtain the loss function. Repeat steps such as model training, obtaining the loss function, and optimizing the model, and finally obtain the target semantic segmentation network model.

[0101] In some embodiments, the sum of the first result loss function and the first process loss function can be calculated; and based on the sum, the initial neural network model can be optimized through the backpropagation algorithm to obtain the target semantic segmentation network model.

[0102] It should be noted that during the process of model training, it is necessary to continuously optimize the model according to the training results to make the output of the model closer and closer to the actual results. At this time, the backpropagation algorithm is needed. The backpropagation algorithm is a commonly used algorithm in neural network training, which is used to calculate the gradient of the loss function with respect to the model parameters. The backpropagation algorithm is the core of the neural network learning process. By continuously adjusting the weights and bias terms of the model, the gap between the predicted output and the actual output of the model is minimized.

[0103] It should be noted that the basic idea of the backpropagation algorithm is to calculate the gradient of the loss function with respect to each parameter through the chain rule, and then use these gradients to update the parameters. This process mainly includes the following steps:

[0104] Forward propagation: Perform forward propagation to calculate the output of the network. The input data passes through each layer of the neural network, from the input layer to the output layer, and the output of each layer is calculated layer by layer; for each layer, the output is a function of the input and the parameters of that layer; finally, the predicted output of the entire network model is obtained. This process involves linearly combining the weights and bias terms with the input data, and then applying an activation function to obtain the output of that layer.

[0105] Calculate the loss: After obtaining the predicted output of the network, a loss function (such as mean squared error, cross-entropy, etc.) can be used to calculate the error between the predicted output and the actual output. The choice of the loss function depends on the specific task. For example, mean squared error is used for regression problems, and cross-entropy loss is used for classification problems.

[0106] Initialize the gradients: For each neuron in the output layer, calculate its corresponding gradient according to the loss function. For other layers, initialize the gradients to 0.

[0107] Backpropagation: Starting from the output layer, propagate the gradients layer by layer in reverse, that is, calculate the gradient of the loss function with respect to each parameter layer by layer. For each neuron, calculate the gradient of its activation value, which usually involves multiplying the gradient passed from the next layer by the derivative of the activation function of the current layer. Then, use the chain rule to calculate the gradients of the corresponding weights and bias terms of that neuron. Specifically, for each neuron in the output layer, its gradient can be directly calculated through the derivative of the loss function with respect to the output. For neurons in the hidden layer, its gradient is the product of the gradient of its output value with respect to the loss function and the gradient passed from the next layer.

[0108] Apply the chain rule: During the backpropagation process, connect the gradient calculations of each layer through the chain rule. This means that the gradient of each neuron is the product of the gradient of its activation function and the gradients of the subsequent layers.

[0109] Parameter Update: Using the calculated gradients, update the weights and bias terms of the model through optimization algorithms (such as gradient descent, Adam, etc.) and the calculated gradients. This typically involves multiplying the learning rate by the gradient and then subtracting this product from the current parameter values.

[0110] Iterative Optimization: Repeat the above steps and gradually adjust the model's parameters through multiple iterations to continuously reduce the value of the loss function until a preset stopping condition is reached (such as the maximum number of iterations, convergence of the loss value, etc.).

[0111] It can be seen that the key to the backpropagation algorithm lies in effectively calculating gradients using the chain rule and updating parameters through optimization algorithms. This enables the neural network to learn complex mapping relationships and achieve excellent performance in various tasks.

[0112] In the embodiments of the present application, it can be seen that after obtaining the first result loss function and the first process loss function, their sum can be calculated, which is the loss situation of the initial neural network model. To reduce this loss function, the initial neural network model can be optimized, that is, adjust the parameters and weights of the initial neural network model according to the sum, and then continue training until the target semantic segmentation network model is obtained.

[0113] Specifically, in some embodiments, optimize the initial neural network model according to the first result loss function and the first process loss function to obtain the first neural network model; train the decoder network in the first neural network model through the training dataset to obtain the second result loss function; train the auxiliary head network in the first neural network model through the training dataset to obtain the second process loss function; optimize the initial neural network model according to the second result loss function and the second process loss function, and repeat the above steps until the second result loss function and the second process loss function are in a stable state to obtain the target semantic segmentation network model.

[0114] It should be noted that the second result loss function and the second process loss function are the results of iteration, which not only indicate one-time repeated training. That is to say, after obtaining the second result loss function and the second process loss function, the first neural network model is further optimized through the sum of the second result loss function and the second process loss function, and then the second neural network model is obtained; again, through the training data set, the decoder network and the auxiliary head network in the second neural network model are trained respectively to obtain the third result loss function and the third process loss function, and then the second neural network model is optimized through the third result loss function and the third process loss function, and then the third neural network model is obtained; then the above steps are continuously repeated until it is detected that the loss function of the neural network model obtained after a certain optimization has converged, then there is no need to optimize anymore, and this neural network model can be determined as the target semantic segmentation network model.

[0115] In some embodiments, optimizing the neural network model is achieved by updating parameters and weights. Specifically, the gradient descent algorithm can be used to update the model parameters to minimize the loss function.

[0116] Specifically, usually the parameters of the early layers of the model are frozen, and then only the parameters of the later layers of the model are updated. This can prevent the parameters of the early layers of the model from overfitting the training data, thereby improving the generalization ability of the model.

[0117] Among them, the frozen parameters are usually the parameters of the early layers of the model, such as: convolutional layers and fully connected layers. The remaining parameters are usually the parameters of the later layers of the model, such as: classification layers and regression layers.

[0118] The method of distinguishing frozen parameters and remaining parameters is usually to use the learning rate to control. For the frozen parameters, the learning rate is usually set to 0, indicating that these parameters are not updated. For the remaining parameters, the learning rate is usually set to a small value, indicating that the update amplitude of these parameters is small.

[0119] In some embodiments, the remaining parameters can be initialized first, then the learning rate is set, and then the loss function and the gradient are iteratively calculated, so as to update the remaining parameters using the gradient descent algorithm, and the steps of iteration and update are continuously repeated until the parameters reach the expected effect. In practical applications, the parameter settings can be adjusted according to specific situations. For example, the number of frozen parameters and the adjustment results can be determined according to the structure of the model and the quality of the training data.

[0120] Among them, the learning rate is a hyperparameter that needs to be set during automated training. The learning rate is one of the important parameters affecting the model training effect. The learning rate controls the amplitude of the model parameter update. If the learning rate is set too large, it will cause the model parameters to be unstable during training and prone to oscillation or divergence. If the learning rate is set too small, it will cause the model training speed to be slow and even unable to converge. Setting the learning rate is to train a better model.

[0121] 105. Input the data to be measured into the target semantic segmentation network model to obtain the target semantic segmentation result.

[0122] In the embodiment of the present application, after obtaining the target semantic segmentation network model through model training, the target semantic segmentation network model can be used to perform semantic segmentation on data such as images or videos. Specifically, the data to be measured can be input into the target semantic segmentation network model, so as to obtain the target semantic segmentation result, and the target semantic segmentation result can be used to indicate the classification result of the data to be measured.

[0123] A semantic segmentation method provided by an embodiment of the present application introduces a decoder network and an auxiliary head network. Through training samples, model training is performed on both the decoder network and the auxiliary head network at the same time. The decoder network focuses on the loss of the training result, and the auxiliary head network focuses on the loss during the training process. It can not only consider the result but also consider the data during the model training process, and optimize the model by comprehensively considering the loss of each layer, which can effectively improve the accuracy, stability and generalization ability of the neural network model.

[0124] As Figure 2 shown, Figure 2 is a flowchart of a semantic segmentation method provided by an embodiment of the present application. The method may further include the following steps:

[0125] 201. Obtain a plurality of training samples.

[0126] In the embodiment of the present application, the training sample is picture data or video data. The training sample can be pre-stored data in a database, data obtained from the cloud, or real-time collected data. This embodiment does not make specific limitations.

[0127] 202. Perform pre-labeling on the plurality of training samples respectively to obtain the training segmentation result corresponding to each training sample.

[0128] In the embodiments of the present application, before model training, manual pre-labeling is required. This pre-labeling is to manually label the content objects corresponding to different pixel points on the pictures or videos of the training samples, so as to obtain the training segmentation result. Since this training segmentation result is the learning reference for subsequent model training, it needs to be ensured to be absolutely correct.

[0129] 203. Through the encoder network in the initial neural network model, feature extraction is performed on multiple training samples to obtain the feature map corresponding to each training sample.

[0130] In the embodiments of the present application, during the model training process, image analysis is based on the features of pixels. Therefore, it is necessary to extract the feature information of each training sample, which can be performed through the encoder network. This encoder network is the backbone network and is a deep convolutional neural network that performs feature extraction on pictures or videos and outputs the corresponding feature map.

[0131] In some embodiments, the model structure of the encoder network is the ConvNeXt network. ConvNeXt is a model of convolutional neural network (CNN) used for image classification and semantic segmentation. ConvNeXt is completely constructed by standard ConvNet modules, which has good characteristics in terms of accuracy and scalability, and exceeds Swin Transformers in COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.

[0132] It should be noted that the ConvNeXt network significantly improves the performance of pure ConvNet in multiple recognition benchmark tests through a series of innovations. It uses the spatial pyramid pooling (SPP) technology to improve the pooling layer in the traditional CNN, captures features at different scales by performing pyramid pooling on the input at different scales, thereby improving the accuracy of the CNN. In addition, ConvNeXt also adopts the idea of the residual network (ResNet) to combine the features of the shallower layer with those of the deeper layer to improve the performance of the network.

[0133] ConvNeXt can also fuse the fully convolutional mask autoencoder framework and the global response normalization (GRN) layer to optimize the original ConvNeXt architecture. At the same time, the generalization ability and efficiency of the model are improved by using self-supervised learning technology. A global response normalization (GRN) layer is newly added in ConvNeXt, and the original LayerScale layer is deleted.

[0134] It should be noted that the feature map corresponding to each training sample extracted by the encoder network can be a feature for characterizing the pixel values of pixel points, or a feature for characterizing pixel boundaries, or a feature for characterizing the background image, etc. Of course, there are many other features, which are not specifically limited here.

[0135] 204. Through the neck connector network in the initial neural network model, the feature map corresponding to each training sample is adjusted in dimension and size.

[0136] In the embodiment of the present application, during subsequent model training, the size requirements of the data for each layer of the model and the encoder network through which the sample data passes may be different. Therefore, the feature map output by the encoder network may not meet the requirements of the decoder network or the auxiliary head network. Therefore, the feature map can be adjusted in dimension and size through the neck connector network.

[0137] In some embodiments, the dimension adjustment is used to indicate adjusting the pixel dimension of the sample data. For example: converting the RGB pixels of the original sample data into grayscale pixels.

[0138] In some embodiments, the size adjustment is used to indicate adjusting the resolution of the sample data. For example: changing the resolution of the original sample data from 1080*1920 to 64*64.

[0139] 205. Multiple training samples, as well as the adjusted feature map and training segmentation result corresponding to each training sample, are determined as the training dataset.

[0140] 206. The adjusted feature map corresponding to each training sample is input into the decoder network of the initial neural network model to obtain the first result training vector.

[0141] In the embodiment of the present application, during the process of model training, by inputting the feature map of each training sample into the decoder network, the first result training vector corresponding to each feature map can be obtained. The first result training vector is used to indicate the prediction result of each pixel point in each feature map, and this prediction result is the semantic segmentation result corresponding to the feature map initially output by the model.

[0142] 207. Through the cross-entropy algorithm, according to the first result training vector and the training segmentation result in the training dataset, the first result loss function is obtained.

[0143] In the embodiment of the present application, the first result training vector is the result output by the model, and the training segmentation result in the training dataset is the result pre-annotated manually. The difference between the first result training vector and the training segmentation result is the loss situation of the model, that is, the first result loss function. The smaller the first result loss function, the more accurate the model is currently.

[0144] In some embodiments, cross entropy is a commonly used loss function in machine learning and deep learning, especially in classification problems. It measures the difference between the predicted probability distribution and the true probability distribution. The cross entropy loss function is often used for the optimization of models such as logistic regression and neural networks.

[0145] It should be noted that the definition of cross entropy is based on two probability distributions p and q, and its formula is: H(p, q) = -∑p(i)*logq(i). Here, p(i) represents the probability of the i-th category in the true distribution, and q(i) represents the probability of the i-th category in the model's predicted distribution. The log here is usually the natural logarithm with base e.

[0146] In classification problems, the true distribution p is usually one-hot encoded. That is, for a certain sample, the probability corresponding to its true category is 1, and the probabilities of other categories are 0. Therefore, the cross entropy loss function can be simplified to: L = -logq(c), where c is the true category of the sample, and q(c) is the probability that the model predicts the sample belongs to the true category.

[0147] In some embodiments, using cross entropy as the loss function is relatively easy to optimize: The cross entropy loss function is a convex function with a unique minimum value. Therefore, optimization algorithms such as gradient descent can be used to find the optimal solution. Also, it is sensitive to the probability distribution: Cross entropy can well reflect the difference between the predicted probability distribution and the true probability distribution, and even a tiny difference can be reflected in the loss value. At the same time, it can be applied to multi-classification problems: The cross entropy loss function can be easily extended to multi-classification problems by simply extending the summation in the formula to all categories.

[0148] In the training of deep learning models, the backpropagation algorithm is usually used to calculate the gradient of the cross entropy loss function with respect to the model parameters, and the model parameters are updated through optimization algorithms to minimize the loss function.

[0149] In the embodiments of the present application, since the encoder network mainly focuses on the loss situation of the feature results, that is to say, during the model training process, the first result loss function corresponding to the encoder network can only represent the loss situation corresponding to the result output by the model. For example: The input feature map for training is 64*64, and during the model training process, upsampling is performed and it is successively converted to 128*128, 256*256, 512*512, 1080*1080, 1080*1920. Finally, the result output by the encoder network is 1080*1920. Then the first result loss function can only represent the loss situation of the 1080*1920 layer.

[0150] 208. Input the adjusted feature map corresponding to each training sample into the auxiliary head network of the initial neural network model to obtain the first-process training vector.

[0151] In the embodiment of the present application, during the model training process, input the feature map of each training sample into the auxiliary head network, and the first-process training vector corresponding to each feature map can be obtained. The first-process training vector is used to indicate the prediction result of each pixel point in each feature map, and this prediction result is the semantic segmentation result corresponding to the feature map initially output by the model.

[0152] 209. Through the cross-entropy algorithm, obtain the first-process loss function according to the first-process training vector and the training segmentation result in the training dataset.

[0153] In the embodiment of the present application, the first-process training vector is the result output by the model, and the training segmentation result in the training dataset is the result pre-annotated manually. The difference between the first-process training vector and the training segmentation result is the loss situation of the model, that is, the first-process loss function. The smaller the first-process loss function, the more accurate the model is currently.

[0154] It should be noted that the process of specifically calculating the first-process loss function through the cross-entropy algorithm is the same as the way of specifically calculating the first-result loss function through the cross-entropy algorithm, and will not be elaborated here.

[0155] In the embodiment of the present application, since the auxiliary head network mainly focuses on the loss situation of the process intermediate state, that is to say, during the model training process, the first-process loss function corresponding to the auxiliary head network can represent the loss situation corresponding to each layer during the model training process. For example: the input feature map for training is 64*64, and during the model training process, upsampling is performed and it is sequentially converted to 128*128, 256*256, 512*512, 1080*1080, 1080*1920, and finally the result output by the auxiliary head network is 1080*1920. Then the first-process loss function can represent the loss situations of the 128*128 layer, 256*256 layer, 512*512 layer, and 1080*1080 layer respectively.

[0156] 210. Optimize the initial neural network model according to the first-result loss function and the first-process loss function to obtain the target semantic segmentation network model.

[0157] In the embodiment of the present application, for the description of step 210, please refer to the detailed description of step 104 in the above embodiment, and the embodiment of the present application will not be elaborated here.

[0158] 211. Input the data to be measured into the encoder network of the target semantic segmentation network model to obtain the feature map corresponding to the data to be measured.

[0159] In the embodiments of the present application, after training is completed, during the process of semantic segmentation of actual data, it is also necessary to extract a feature map from the data, and the extraction method is the same as that in the training process, which will not be elaborated here.

[0160] In some embodiments, for the feature map corresponding to the data to be measured, it also needs to be input into the neck connector for dimension and size adjustment, which will not be elaborated here.

[0161] 212. Input the feature map into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result.

[0162] In the embodiments of the present application, after obtaining the feature map, the feature map can be input into the decoder network of the target semantic segmentation network model, and the auxiliary head network does not need to be input, and the target semantic segmentation result can be obtained.

[0163] In some embodiments, some post-processing can also be performed on the result output by the decoder network. Specifically, the feature map is input into the decoder network of the target semantic segmentation network model to obtain a target vector corresponding to the feature map, and each element in the target vector is used to indicate the classification result of each pixel in the data to be measured; through a preset mask matrix, the target vector is masked to obtain the target semantic segmentation result.

[0164] It should be noted that through the preset mask matrix, the output result can be converted into an image with the same size as the original input image, and the value of each pixel point represents the classification result of the corresponding area.

[0165] Exemplarily, as Figure 3 shown is a schematic diagram of the semantic segmentation result. After the result obtained through the target semantic segmentation network model in Figure 3 is masked, it has different pixel values for each object, and each different object can be accurately distinguished.

[0166] A semantic segmentation method provided by the embodiments of the present application introduces a decoder network and an auxiliary head network. Through training samples, the model training of both the decoder network and the auxiliary head network is carried out simultaneously. The decoder network focuses on the loss of the training result, and the auxiliary head network focuses on the loss during the training process. It can not only consider the result but also consider the data during the model training process, and optimize the model by comprehensively considering the loss of each layer, which can effectively improve the accuracy, stability and generalization ability of the neural network model.

[0167] As Figure 4 shown, the embodiments of the present application provide a semantic segmentation device, and the semantic segmentation device may include:

[0168] An acquisition module 401, configured to acquire a training data set, where the training data set at least includes: a plurality of training samples, and a training segmentation result corresponding to each training sample, and the training segmentation result is obtained after semantic segmentation pre-labeling;

[0169] A processing module 402, configured to train a decoder network in an initial neural network model through the training data set to obtain a first result loss function;

[0170] The processing module 402 is further configured to train an auxiliary head network in the initial neural network model through the training data set to obtain a first process loss function;

[0171] The processing module 402 is further configured to optimize the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model;

[0172] The processing module 402 is further configured to input the data to be measured into the target semantic segmentation network model to obtain a target semantic segmentation result, and the target semantic segmentation result is used to indicate the classification result of the data to be measured.

[0173] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module 402 is specifically configured to calculate the sum of the first result loss function and the first process loss function;

[0174] The processing module 402 is specifically configured to optimize the initial neural network model based on the sum through a backpropagation algorithm to obtain a target semantic segmentation network model.

[0175] As an optional implementation manner, in the second aspect of the embodiments of the present application, the processing module 402 is specifically configured to optimize the initial neural network model according to the first result loss function and the first process loss function to obtain a first neural network model;

[0176] The processing module 402 is specifically configured to train a decoder network in the first neural network model through the training data set to obtain a second result loss function;

[0177] The processing module 402 is specifically configured to train an auxiliary head network in the first neural network model through the training data set to obtain a second process loss function;

[0178] The processing module 402 is specifically configured to optimize the first neural network model according to the second result loss function and the second process loss function, and repeat the above steps until the second result loss function and the second process loss function are in a stable state to obtain a target semantic segmentation network model.

[0179] As an alternative implementation, in the second aspect of the embodiments of the present application, the processing module 402 is specifically configured to input the data to be measured into the encoder network of the target semantic segmentation network model to obtain a feature map corresponding to the data to be measured;

[0180] The processing module 402 is specifically configured to input the feature map into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result.

[0181] As an alternative implementation, in the second aspect of the embodiments of the present application, the processing module 402 is specifically configured to input the feature map into the decoder network of the target semantic segmentation network model to obtain a target vector corresponding to the feature map, and each element in the target vector is used to indicate the classification result of each pixel in the data to be measured;

[0182] The processing module 402 is specifically configured to perform masking processing on the target vector through a preset masking matrix to obtain the target semantic segmentation result.

[0183] As an alternative implementation, in the second aspect of the embodiments of the present application, the acquisition module 401 is specifically configured to acquire a plurality of training samples;

[0184] The processing module 402 is specifically configured to perform pre-labeling on the plurality of training samples respectively to obtain a training segmentation result corresponding to each training sample;

[0185] The processing module 402 is specifically configured to perform feature extraction on the plurality of training samples through the encoder network in the initial neural network model to obtain a feature map corresponding to each training sample;

[0186] The processing module 402 is specifically configured to perform dimension and size adjustment on the feature map corresponding to each training sample through the neck connector network in the initial neural network model;

[0187] The processing module 402 is specifically configured to determine the plurality of training samples, as well as the adjusted feature map and training segmentation result corresponding to each training sample, as a training data set.

[0188] As an alternative implementation, in the second aspect of the embodiments of the present application, the processing module 402 is specifically configured to input the adjusted feature map corresponding to each training sample into the decoder network of the initial neural network model to obtain a first result training vector, and the first result training vector is used to indicate the prediction result of each pixel point in the plurality of training samples;

[0189] The processing module 402 is specifically configured to obtain a first result loss function according to the first result training vector and the training segmentation result in the training data set through the cross-entropy algorithm.

[0190] In the embodiments of the present application, each module can implement the semantic segmentation method provided in the above method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein again.

[0191] As Figure 5 shown, an embodiment of the present application further provides an electronic device, which may include:

[0192] A memory 501 storing executable program code;

[0193] A processor 502 coupled to the memory 501;

[0194] Wherein, the processor 502 invokes the executable program code stored in the memory 501 and executes the semantic segmentation method executed by the electronic device in the above method embodiments.

[0195] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the semantic segmentation method in the above method embodiments and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0196] An embodiment of the present application further provides a computer program product, which stores a computer program. When the computer program is executed by a processor, it implements each process of the semantic segmentation method in the above method embodiments and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0197] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0198] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0199] In the present application, the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0200] In the present application, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0201] In this application, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The computer-readable medium includes permanent and non-permanent, removable and non-removable storage media. The storage medium can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (Parallel Random Access Memory, PRAM), static random access memory (Static Random Access Memory, SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), programmable read-only memory (Programmable Read-only Memory, PROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), other types of random access memory (Random Access Memory, RAM), read-only memory (Read-Only Memory, ROM), one-time programmable read-only memory (One-time Programmable Read-Only Memory, OTPROM), electrically-erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), flash memory or other memory technologies, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0202] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0203] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments. Dividing them into multiple embodiments is only used to highlight the different technical features in different embodiments. Those skilled in the art should know that the above-mentioned multiple embodiments can also be combined arbitrarily.

[0204] In various embodiments of the present application, it should be understood that the magnitude of the serial numbers of the above processes does not necessarily mean the inevitable sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0205] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0206] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0207] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute some or all of the steps of the above methods in various embodiments of this application.

[0208] The above are only the specific implementation manners of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A semantic segmentation method, characterized in that: The method comprises: Obtain a training data set, the training data set comprising at least: a plurality of training samples, and a training segmentation result corresponding to each training sample, the training segmentation result being obtained after semantic segmentation pre-marking; Using the training data set, a decoder network in the initial neural network model is trained to obtain a first result loss function, where the first result loss function is used to indicate the loss of the image sampling result in the model training; Using the training data set, training the auxiliary head network in the initial neural network model to obtain a first process loss function, wherein the auxiliary head network is a fully convolutional network, and the first process loss function is used to indicate the loss in the image upsampling process in the model training; According to the first result loss function and the first process loss function, optimizing the initial neural network model to obtain a target semantic segmentation network model; The data to be tested is input into the target semantic segmentation network model to obtain a target semantic segmentation result, and the target semantic segmentation result is used to indicate the classification result of the data to be tested.

2. The method according to claim 1, characterized in that The step of optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model includes: Calculating the sum of the first result loss function and the first process loss function; The target semantic segmentation network model is obtained by optimizing the initial neural network model based on the sum through a back propagation algorithm.

3. The method according to claim 2, characterized in that The step of optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model includes: Optimizing the initial neural network model according to the first result loss function and the first process loss function to obtain a first neural network model; Using the training data set, training the decoder network in the first neural network model to obtain a second result loss function; Using the training data set, training the auxiliary head network in the first neural network model to obtain a second process loss function; According to the second result loss function and the second process loss function, the first neural network model is optimized, and the above steps are repeated until the second result loss function and the second process loss function are in a stable state, thereby obtaining the target semantic segmentation network model.

4. The method according to claim 1, characterized in that The step of inputting the test data into the target semantic segmentation network model to obtain the target semantic segmentation result includes: Inputting the test data into the encoder network of the target semantic segmentation network model to obtain a feature map corresponding to the test data; The feature map is input into the decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result.

5. The method according to claim 4, characterized in that The step of inputting the feature map into a decoder network of the target semantic segmentation network model to obtain the target semantic segmentation result includes: Inputting the feature map into the decoder network of the target semantic segmentation network model to obtain a target vector corresponding to the feature map, wherein each element in the target vector is used to indicate a classification result of each pixel in the data to be tested; By presetting the mask matrix, the target vector is masked to obtain the target semantic segmentation result.

6. The method according to claim 1, characterized in that The step of obtaining a training data set includes: Acquire the multiple training samples; Pre-marking the multiple training samples respectively to obtain a training segmentation result corresponding to each training sample; Extracting features from the plurality of training samples through an encoder network in the initial neural network model to obtain a feature map corresponding to each training sample; By using the neck connector network in the initial neural network model, adjusting the dimension and size of the feature map corresponding to each training sample; The multiple training samples, and the adjusted feature map and training segmentation result corresponding to each training sample are determined as the training data set.

7. The method according to claim 6, characterized in that The step of training the decoder network in the initial neural network model through the training data set to obtain a first result loss function includes: Inputting the adjusted feature map corresponding to each training sample into the decoder network of the initial neural network model to obtain a first result training vector, where the first result training vector is used to indicate the prediction result of each pixel point in the multiple training samples; The first result loss function is obtained according to the first result training vector and the training segmentation results in the training data set through a cross entropy algorithm.

8. A semantic segmentation device, characterized in that: include: An acquisition module is used to acquire a training data set, wherein the training data set includes at least: a plurality of training samples and a training segmentation result corresponding to each training sample, wherein the training segmentation result is obtained after semantic segmentation pre-marking; A processing module, used to train a decoder network in the initial neural network model using the training data set to obtain a first result loss function, where the first result loss function is used to indicate the loss of the sampling result of the image in the model training; The processing module is further used to train the auxiliary head network in the initial neural network model through the training data set to obtain a first process loss function, wherein the auxiliary head network is a fully convolutional network, and the first process loss function is used to indicate the loss in the image upsampling process in the model training; The processing module is further used to optimize the initial neural network model according to the first result loss function and the first process loss function to obtain a target semantic segmentation network model; The processing module is also used to input the data to be tested into the target semantic segmentation network model to obtain a target semantic segmentation result, and the target semantic segmentation result is used to indicate a classification result of the data to be tested.

9. An electronic device, characterized in that: include: A memory storing executable program code; and a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the semantic segmentation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the semantic segmentation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image processing method and device

    CN110969627A

  • Semantic segmentation network model training method and device, equipment and storage medium

    CN114663659A