A training method and device of a semantic segmentation model, and an active sample screening method

By jointly training the initial semantic segmentation model and the generative adversarial network model, and using a discriminator to filter active samples, the problem of high annotation cost in semantic segmentation model training is solved, segmentation accuracy is improved and the sample selection process is simplified.

CN116883789BActive Publication Date: 2026-06-12NEUSOFT REACH AUTOMOBILE TECH (SHENYANG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NEUSOFT REACH AUTOMOBILE TECH (SHENYANG) CO LTD
Filing Date
2023-07-20
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing technologies suffer from high annotation costs and low annotation efficiency in semantic segmentation model training, making it difficult to accurately select active samples that play an important role in improving model performance.

Method used

By jointly training the initial semantic segmentation model and the adversarial generative network model, image semantic segmentation results are generated, and the first and second discriminators are used to identify the source of the samples in order to filter out active samples, thus simplifying the sample selection process.

Benefits of technology

It improves the segmentation accuracy of semantic segmentation models, reduces annotation costs, and simplifies the sample selection process, directly selecting samples that meet the model's requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883789B_ABST
    Figure CN116883789B_ABST
Patent Text Reader

Abstract

The application provides a semantic segmentation model training method, an active sample screening method and device. The semantic segmentation model training method comprises: training an initial semantic segmentation model according to a training sample and a generated image to obtain a first semantic segmentation result and a generated image semantic segmentation result; and retraining the initial semantic segmentation model according to the first semantic segmentation result and the generated image semantic segmentation result to obtain a target semantic segmentation model, wherein the target semantic segmentation model comprises the initial semantic segmentation model and a generative adversarial network model. The target semantic segmentation model provided by the application adds the generative adversarial network model, so that the segmentation result is more accurate. Meanwhile, the model can simplify the active sample screening process and reduce the labeling cost of samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a training method for a semantic segmentation model, an active sample selection method, and an apparatus. Background Technology

[0002] The main purpose of semantic segmentation of images is to predict what object each pixel in an input image belongs to. When training a semantic segmentation model, labeled images with annotation data are typically used as training samples. However, since these labeled images are pixel-level annotations, this requires a significant amount of manpower and is inefficient.

[0003] Therefore, how to accurately select active samples that play an important role in improving the segmentation effect of semantic segmentation models, so as to reduce the labeling cost of training samples, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a training method for a semantic segmentation model, an active sample selection method and apparatus, which can simplify the sample selection process and reduce the labeling cost of training samples.

[0005] Firstly, a training method for a semantic segmentation model is provided, including:

[0006] The initial semantic segmentation model is trained based on the training samples and the generated image to obtain the first semantic segmentation result and the semantic segmentation result of the generated image; the generated image is obtained by training an adversarial generative network model.

[0007] The initial semantic segmentation model is retrained based on the first semantic segmentation result and the generated image semantic segmentation result to obtain a target semantic segmentation model, which includes the initial semantic segmentation model and an adversarial generative network model.

[0008] Secondly, an active sample screening method is provided, including:

[0009] The samples to be screened are input into the target semantic segmentation model described above;

[0010] Based on the discrimination results of the first discriminator and the second discriminator, it is determined whether the sample to be screened should be used as an active sample.

[0011] Thirdly, a training device for a semantic segmentation model is provided, comprising:

[0012] The first training module is used to train the initial semantic segmentation model based on training samples and generated images to obtain the first semantic segmentation result and the semantic segmentation result of the generated image; the generated image is obtained by training an adversarial generative network model.

[0013] The second training module is used to retrain the initial semantic segmentation model and the adversarial generative network model based on the first semantic segmentation result and the generated image semantic segmentation result, so as to obtain the target semantic segmentation model.

[0014] Fourthly, an active sample screening device is provided, comprising:

[0015] The sample input module is used to input the samples to be screened into the target semantic segmentation model described above, and to obtain the first semantic segmentation result and the generated semantic segmentation result;

[0016] The determining module is used to determine whether to include the sample to be screened as an active sample based on the discrimination results of the first discriminator and / or the second discriminator.

[0017] Fifthly, an electronic device is provided, comprising: a processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory, and performing the methods as described in the first aspect or its various implementations.

[0018] In a sixth aspect, a computer-readable storage medium is provided for storing a computer program that causes a computer to perform the methods described in the first aspect or its various implementations.

[0019] The technical solution provided in this application first trains an initial semantic segmentation model based on training samples and generated images to obtain a first semantic segmentation result and a generated image semantic segmentation result. Then, the initial semantic segmentation model is retrained based on the first semantic segmentation result and the generated image semantic segmentation result to obtain a target semantic segmentation model. Because the target semantic segmentation model incorporates an adversarial generative network model, the obtained target semantic segmentation model can extract more accurate segmentation results, improving the model's segmentation accuracy. Furthermore, this target semantic segmentation model can be directly used for active sample selection, thus simplifying the selection process and ensuring that the selected samples better meet the model's needs.

[0020] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 An application scenario diagram provided for an embodiment of this application;

[0023] Figure 2 A flowchart illustrating a training method for a semantic segmentation model provided in an embodiment of this application;

[0024] Figure 3 A flowchart illustrating another training method for a semantic segmentation model provided in this application embodiment;

[0025] Figure 4 A flowchart illustrating another training method for a semantic segmentation model provided in this application embodiment;

[0026] Figure 5 A flowchart of an active sample screening method provided in this application embodiment;

[0027] Figure 6 A schematic diagram of a training device for a semantic segmentation model provided in an embodiment of this application;

[0028] Figure 7 A schematic diagram of an active sample screening device provided in an embodiment of this application;

[0029] Figure 8 This is a schematic block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0032] Semantic segmentation plays a crucial role in the fields of assisted driving and autonomous driving. Since training a semantic segmentation model requires labeling images pixel-by-pixel, which is costly, selecting samples that are important to the model and contribute to its performance is essential. This allows for more valuable image labeling with limited resources.

[0033] As mentioned above, most active learning methods for semantic segmentation sample problems focus on learning from a specific region or certain pixels in an image, ignoring the overall complexity of the image and its importance to the model, and are not applicable to practical applications. A few algorithms that use the image as a whole for active sample learning calculate the image's complexity and difficulty scores, sort the samples in the pool from highest to lowest score, set a percentage threshold, and select the top-scoring samples as active samples. While this approach considers the overall image, this top-k method is more suitable for static sample pools and performs poorly with dynamic sample pools in projects. Furthermore, the percentage threshold requires more expert experience and multiple trials to determine, and for dynamic sample pools, the percentage threshold needs to be re-determined through repeated trials, making it unsuitable for real-world projects.

[0034] To address the aforementioned technical problems, the inventive concept of this application is as follows: an electronic device can retrain the initial semantic segmentation model and the adversarial generative network model based on the first semantic segmentation result and the generated image semantic segmentation result to obtain a target semantic segmentation model. The target semantic segmentation model provided by this invention can improve the inference accuracy of the segmentation model without increasing inference time performance. At the same time, it can also enable the training of the semantic segmentation model obtained by this application to actively learn from the samples, thereby greatly simplifying the process of screening active samples. It does not require setting a score threshold or a top-k percentage threshold, and can more accurately screen out active samples that play an important role in improving the model performance.

[0035] It should be understood that the technical solution of this application can be applied to the following scenarios, but is not limited to:

[0036] In some possible ways, Figure 1 An application scenario diagram provided for an embodiment of this application, such as Figure 1 As shown, this application scenario may include electronic device 110 and network device 120. Electronic device 110 can establish a connection with network device 120 through a wired network or a wireless network.

[0037] For example, electronic device 110 may be a desktop computer, laptop computer, tablet computer, etc., but is not limited thereto. Network device 120 may be a terminal device or a server, but is not limited thereto. In one embodiment of this application, electronic device 110 may send a request message to network device 120, which may be used to request the acquisition of training samples and the generation of images. Further, electronic device 110 may receive a response message sent by network device 120, which includes training samples and generated images.

[0038] also, Figure 1 An electronic device 110 and a network device 120 are provided as examples, but other numbers of electronic devices and network devices may be included in practice, and this application does not limit this.

[0039] In other possible implementations, the technical solution of this application may also be executed by the aforementioned electronic device 110, or by the aforementioned network device 120, and this application does not impose any restrictions on this.

[0040] After introducing the application scenarios of the embodiments of this application, the technical solution of this application will be described in detail below:

[0041] Figure 2 A flowchart illustrating a training method for a semantic segmentation model provided in this application embodiment is shown. This method can be performed by, for example... Figure 1 The electronic device 110 shown performs, but is not limited to, its functions. For example... Figure 2 As shown, the method may include the following steps:

[0042] S210. Train the initial semantic segmentation model based on the training samples and the generated image to obtain the first semantic segmentation result and the semantic segmentation result of the generated image.

[0043] In this step, the generated image is obtained by training an adversarial generative network model, which includes a generator and a first discriminator. The generator is used to generate a generated image from the feature samples, which is also called a pseudo-original image; the first discriminator is used to determine the source of the feature samples.

[0044] Here, the training samples can be road images containing road condition information collected by the vehicle's onboard camera during autonomous driving. The first semantic segmentation result can be the image semantic segmentation result of the drivable area of ​​the road surface corresponding to the road image and the image semantic segmentation result of the road sign. For example, the image semantic segmentation result of the drivable area of ​​the road surface can include the first sign corresponding to the area in the road image where vehicles can drive. The first sign can include lane lines, arrows, and stop lines, etc.; as another example, the image semantic segmentation result of the road sign can include the second sign corresponding to the area in the road image where vehicles can drive. The second sign can include road signs, traffic restriction signs, and traffic lights, etc.

[0045] In this embodiment of the application, the generated image can be a generated image obtained by inputting feature samples into an adversarial generative network model and training it adversarially. Here, the initial semantic segmentation model includes an encoder and a decoder. By inputting training samples into the initial semantic segmentation model, a first semantic segmentation result can be obtained. By inputting the generated image into the initial semantic segmentation model, a semantic segmentation result of the generated image can be obtained.

[0046] S220. Based on the first semantic segmentation result and the generated image semantic segmentation result, the initial semantic segmentation model is retrained to obtain the target semantic segmentation model.

[0047] In this step, the target semantic segmentation model is used to obtain the second semantic segmentation result based on the input training samples. The target semantic segmentation model may include the initial semantic segmentation model and the adversarial generative network model. Here, the second semantic segmentation result is the semantic segmentation result output by the target semantic segmentation model based on the training samples.

[0048] Here, a discrimination is performed based on the first semantic segmentation result and the generated image semantic segmentation result. Then, a first loss function value is calculated based on the first semantic segmentation result and the generated image semantic segmentation result. At the same time, a second loss function value is calculated based on the training samples and the generated image. Then, the parameters of the initial semantic segmentation model and the generative adversarial network model are adjusted based on the first loss function value and the second loss function value until the loss function values ​​corresponding to the initial semantic segmentation model and the generative adversarial network model are within their respective preset ranges, thus obtaining the target semantic segmentation model. Here, the loss function values ​​corresponding to the initial semantic segmentation model and the generative adversarial network model include the first loss function value and the second loss function value.

[0049] In some feasible implementations, when retraining the initial semantic segmentation model based on the first semantic segmentation result and the generated image semantic segmentation result, the initial semantic segmentation model can be trained multiple times according to preset rules until the discrimination result corresponding to the second discriminator meets the corresponding preset condition, at which point training ends and the target semantic segmentation model is obtained. For example, during training, the parameters of the encoder and decoder of the initial semantic segmentation model, as well as the generator, first discriminator, and second discriminator of the generative adversarial network model, are initialized. First, the parameters of the first and second discriminators are fixed, and the parameters of the encoder, decoder, and generator are trained. After training for a set number of rounds, the parameters of the encoder, decoder, and generator are fixed again, and the parameters of the first and second discriminators are trained. After training for a set number of rounds, the parameters of the first and second discriminators are fixed again, and the parameters of the encoder, decoder, and generator are trained. This process is repeated multiple times until a set number of rounds or a certain preset condition is reached, at which point training stops and the target semantic segmentation model is obtained.

[0050] It should be noted that, with the target semantic segmentation model obtained through this embodiment, when performing semantic prediction on the training samples, it is only necessary to input the training samples into the model to obtain the final second semantic segmentation result. Therefore, the target semantic segmentation model provided in this embodiment does not increase the time performance of the model in semantic prediction.

[0051] Using the above method, a generative adversarial network model is added to the initial semantic segmentation model after training. The initial semantic segmentation model is then trained again based on training samples and generated images to obtain the first semantic segmentation result and the semantic segmentation result of the generated image. Without adding manual annotation to the training samples, the initial semantic segmentation model is retrained only based on the first semantic segmentation result and the semantic segmentation result of the generated image. This allows the obtained target semantic segmentation model to extract semantic segmentation results with more semantic information, thereby improving the model's segmentation accuracy without increasing the model's time performance during inference.

[0052] Figure 3 A flowchart illustrating another training method for a semantic segmentation model provided in this application embodiment.

[0053] based on Figure 2 ,like Figure 3 As shown, S210 above includes:

[0054] S310. Input the training samples into the encoder of the initial semantic segmentation model to obtain the first feature image.

[0055] Here, feature extraction is performed on the training samples to obtain the first feature image of the training samples. Feature extraction can be achieved using the encoder of the initial semantic segmentation model. The feature extraction network model in the encoder can adopt the network structure of the VGG network. The specific training process will not be described in detail here.

[0056] S320. Input the first feature image into the decoder of the initial semantic segmentation model to obtain the first semantic segmentation result.

[0057] In this embodiment, after the training sample enters the encoder of the initial semantic segmentation model, the first feature image is extracted, and then the first feature image is input into the decoder of the initial semantic segmentation model to obtain the predicted first semantic segmentation result corresponding to the training sample.

[0058] For example, taking a "road image" captured by an onboard camera during autonomous driving as the training sample, the road image displays lane lines, arrows, stop lines, road signs, traffic restriction signs, and traffic lights. The road image is input into the encoder of the initial semantic segmentation model to obtain first feature image 1, first feature image 2, first feature image 3, first feature image 4, first feature image 5, and first feature image 6. Then, all or part of the first feature images are input into the decoder of the initial semantic segmentation model to obtain the predicted first semantic segmentation result corresponding to the road image. Here, the first semantic segmentation result includes annotations with predicted identifiers corresponding to lane lines, arrows, stop lines, road signs, traffic restriction signs, and traffic lights, respectively. This is only an example and is not a limitation of this disclosure.

[0059] S330. Calculate the segmentation cross-entropy loss between the first semantic segmentation result and the real labeled sample.

[0060] Since the first semantic segmentation result obtained by the initial semantic segmentation model is the prediction result corresponding to the training sample, there may be a difference between the first semantic segmentation result and the real labeled sample corresponding to the training sample. Therefore, in order to obtain a first semantic segmentation result that is close to the real labeled sample, the difference between the first semantic segmentation result and the real labeled sample can be quickly determined by calculating the segmentation cross-entropy loss between the first semantic segmentation result and the real labeled sample. At the same time, it can also facilitate the adjustment of the parameters of the initial semantic segmentation model in step S340.

[0061] S340. Adjust the parameters of the initial semantic segmentation model based on the segmentation cross-entropy loss to obtain the trained initial semantic segmentation model.

[0062] In this step, the difference information between the first semantic segmentation result and the real labeled sample is determined based on the segmentation cross-entropy loss, so as to adjust the parameters of the initial semantic segmentation model until the initial semantic segmentation model converges. That is, the segmentation cross-entropy loss between the first semantic segmentation result obtained by the trained initial semantic segmentation model and the real labeled sample is within a preset threshold range, so as to minimize the difference between the first semantic segmentation result obtained by the trained initial semantic segmentation model in this embodiment and the real labeled sample.

[0063] It is understood that the segmentation cross-entropy loss can be calculated using any cross-entropy loss function in existing technologies. Using the above implementation method, by calculating the segmentation cross-entropy loss between the first semantic segmentation result and the ground truth labeled sample, the difference between the first semantic segmentation result and the ground truth labeled sample can be determined. By continuously training the initial semantic segmentation model to minimize this segmentation cross-entropy loss, a trained initial semantic segmentation model is obtained. This ensures that the trained initial semantic segmentation model minimizes the difference between the first semantic segmentation result predicted by the training samples and the ground truth labeled sample, thereby improving the segmentation accuracy of the trained initial semantic segmentation model.

[0064] Accordingly, the generated images are trained in the following way:

[0065] S410. Input the training samples into the encoder of the initial semantic segmentation model to obtain the first feature image.

[0066] In this step, in order for the generated image in step S210 to correspond to the content displayed by the training sample, the first feature image is obtained by inputting the training sample into the encoder of the initial semantic segmentation model, so that the generated image obtained in step S220 is a pseudo-original image of the training sample.

[0067] S420. Input the first feature image into the generator of the adversarial generative network model to obtain the generated image.

[0068] Here, by inputting the first feature image into the generator of the adversarial generative network model, a generated image corresponding to the training sample can be obtained. This generated image can not only serve as a training sample for the initial semantic segmentation model, but also determine the accuracy of the initial semantic segmentation model by comparing the generated semantic segmentation result corresponding to the generated image with the first semantic segmentation result corresponding to the training sample.

[0069] By adopting the above implementation method, the first feature image corresponding to the training sample is input into the generator of the adversarial generative network model to obtain the generated image corresponding to the training sample. In subsequent steps, the parameters of the adversarial generative network model can be adjusted by the difference between the training sample and the generated image. At the same time, the generated image can also be used as the training sample of the initial semantic segmentation model to expand the training sample set of the initial semantic segmentation model. By distinguishing the generated semantic segmentation result corresponding to the generated image and the first semantic segmentation result corresponding to the training sample, the parameters of the initial semantic segmentation model can be adjusted to further improve the prediction accuracy of the initial semantic segmentation model.

[0070] This embodiment trains an initial semantic segmentation model based on the generated image to obtain the semantic segmentation result of the generated image, which may include:

[0071] S510. Input the generated image into the encoder of the initial semantic segmentation model to obtain the second feature image.

[0072] S520. Input the second feature image into the decoder of the initial semantic segmentation model to obtain the semantic segmentation result of the generated image.

[0073] Here, the generated image is input into the encoder of the initial semantic segmentation model to obtain the second feature image, and then the second feature image is input into the decoder of the initial semantic segmentation model to enable the initial semantic segmentation model to predict the generated image. This allows the generated image to be used as a training sample for the initial semantic segmentation model, thereby expanding the training sample set of the initial semantic segmentation model. At the same time, by using the initial semantic segmentation model predicted by the training samples to predict the generated image, the comparability between the generated semantic segmentation result corresponding to the generated image and the first semantic segmentation result corresponding to the training samples can be achieved. Furthermore, by discriminating between the generated semantic segmentation result and the first semantic segmentation result, the parameters of the initial semantic segmentation model can be adjusted to further improve the prediction accuracy of the initial semantic segmentation model.

[0074] like Figure 4 As shown, the anti-generative network model in this embodiment can be trained according to the following steps:

[0075] S610. Calculate the L2 distance loss between the generated image and the training samples.

[0076] In this step, the generated image is obtained by adversarially training the first feature image of the training sample in the adversarial generative network model. Here, the difference between the generated image and the training sample can be determined by calculating the L2 distance loss (mean squared error loss function, MSE) between the generated image and the training sample.

[0077] S620. The parameters of the generator of the adversarial generative network model are adjusted based on L2 distance loss to obtain the generator of the trained adversarial generative network model.

[0078] In this step, the difference information between the generated image and the training samples is determined based on the L2 distance loss, so as to adjust the parameters of the generator of the adversarial generative network model. That is, by continuously training the generator to minimize the L2 distance loss, a trained generator can be obtained. This continues until the generator of the adversarial generative network model converges, that is, the L2 distance loss between the generated image obtained by the generator obtained by the trained adversarial generative network model and the training samples is within a preset threshold range, so as to minimize the difference between the generated image obtained by the generator obtained by the trained adversarial generative network model in this embodiment and the training samples.

[0079] S630. Input the first feature image and the second feature image into the first discriminator of the adversarial generative network model. The first discriminator is used to determine whether the second feature image comes from the training sample or the generated image.

[0080] Here, by inputting the first feature image and the second feature image into the first discriminator of the Generative Adversarial Network (GAN) model, it is determined whether the first and second feature images originate from training samples or generated images. If the first discriminator cannot distinguish whether the second feature image originates from training samples or generated images, it is determined that the difference between the second feature image and the first feature image is small. Since the encoder of the same initial semantic segmentation model is used to extract features from the generated images and training samples in this embodiment, it can be determined that the difference between the generated images output by the GAN model and the training samples is small. If the first discriminator can distinguish that the second feature image originates from the generated image, it is determined that the difference between the second feature image and the first feature image is large. Therefore, it can be determined that the difference between the generated images output by the GAN model and the training samples is large, and the parameters of the GAN model are adjusted according to the discrimination result of the first discriminator.

[0081] S640. Calculate the first classification loss between the first feature image and the second feature image.

[0082] Here, since the difference between the generated image obtained by the generator of the generative adversarial network model trained in step S620 and the training sample is very small, the first feature image and the second feature image extracted by the encoder of the same initial semantic segmentation model from the training sample and the generated image should have very small differences. Therefore, by calculating the first classification loss of the first feature image and the second feature image, the difference information between the first feature image and the second feature image in the first discriminator can be determined by the calculated first classification loss, so as to adjust the parameters of the first discriminator of the generative adversarial network model according to the first classification loss.

[0083] Understandably, the first classification loss can be any loss function used in existing technologies to evaluate the degree of difference between the model's predicted values ​​and the true values, such as the absolute value loss function or the exponential loss function.

[0084] S650. Adjust the parameters of the first discriminator of the adversarial generative network model based on the first classification loss to obtain the first discriminator of the trained adversarial generative network model.

[0085] In this step, the parameters of the first discriminator of the adversarial generative network model are adjusted based on the first classification loss until the first discriminator of the adversarial generative network model converges. This means that the discrimination result obtained by the first discriminator of the trained adversarial generative network model is within the preset discrimination range, so as to make the discrimination result obtained by the first discriminator of the trained adversarial generative network model in this embodiment more accurate.

[0086] Using the above implementation method, the generator of the adversarial generative network model is obtained by adjusting the parameters of the generator based on L2 distance loss. Then, the first discriminator of the adversarial generative network model is obtained by adjusting the parameters of the first discriminator based on the first classification loss. This results in an adversarial generative network model that can generate samples with relatively small differences from the training samples.

[0087] Furthermore, after retraining the initial semantic segmentation model and the adversarial generative network model based on the first semantic segmentation result and the generated image semantic segmentation result to obtain the target semantic segmentation model, this embodiment may further include:

[0088] S710. Use the second discriminator to distinguish between the first semantic segmentation result and the generated image semantic segmentation result.

[0089] In this step, the second discriminator is used to determine whether the semantic segmentation result of the generated image comes from the training sample or the generated image.

[0090] It should be noted that by using a second discriminator to distinguish between the first semantic segmentation result and the generated image semantic segmentation result, it can be determined whether the generated image semantic segmentation result originates from the training sample or the generated image. If the second discriminator cannot distinguish whether the first semantic segmentation result and the generated image semantic segmentation result originate from the training sample or the generated image, it is determined that the difference between the first semantic segmentation result and the generated image semantic segmentation result is small. Furthermore, since the same initial semantic segmentation model is used to perform semantic segmentation on the generated image and the training sample in this embodiment, it can be determined that the difference between the output results of the initial semantic segmentation model is small. If the second discriminator can distinguish whether the first semantic segmentation result originates from the training sample and whether the generated image semantic segmentation result originates from the generated image, it is determined that the difference between the first semantic segmentation result and the generated image semantic segmentation result is large. Therefore, it can be determined that the difference between the output results of the initial semantic segmentation model is large, and the parameters of the initial semantic segmentation model are adjusted based on the discrimination result of the second discriminator.

[0091] S720. Calculate the second classification loss of the first semantic segmentation result and the generated image semantic segmentation result.

[0092] Here, if the difference between the generated image obtained by the generator of the generative adversarial network model trained in step S620 and the training sample is small, then the semantic segmentation results of the training sample and the generated image by the same initial semantic segmentation model should have a small difference. Therefore, by calculating the second classification loss of the first semantic segmentation result and the semantic segmentation result of the generated image, the difference information between the first semantic segmentation result and the semantic segmentation result of the generated image in the second discriminator can be determined by the calculated second classification loss, so as to adjust the parameters of the second discriminator of the adversarial generative network model according to the second classification loss.

[0093] Understandably, the second classification loss can employ loss functions already available in the art for evaluating the degree of difference between the model's predicted and true values, such as the absolute value loss function or the exponential loss function.

[0094] S730. Adjust the parameters of the initial semantic segmentation model according to the second classification loss to obtain the trained target semantic segmentation model.

[0095] In this step, the parameters of the second discriminator can be adjusted based on the second classification loss until the second discriminator converges, meaning the discrimination result obtained by the trained second discriminator is within a preset discrimination range, thus making the discrimination result obtained by the second discriminator in this embodiment more accurate. The second discriminator, with its adjusted parameters, then discriminates between the first semantic segmentation result and the generated image semantic segmentation result to determine whether the generated image semantic segmentation result originates from the training samples or the generated image. If the second discriminator can still distinguish that the first semantic segmentation result originates from the training samples and the generated image semantic segmentation result originates from the generated image, it can be determined that the difference between the output results of the initial semantic segmentation model is significant. Therefore, based on the discrimination result of the second discriminator, the parameters of the initial semantic segmentation model are adjusted to further improve the prediction accuracy of the initial semantic segmentation model.

[0096] Using the above implementation method, the first semantic segmentation result and the generated image semantic segmentation result are distinguished by the second discriminator, and then the second classification loss of the first semantic segmentation result and the generated image semantic segmentation result is calculated. Finally, the parameters of the initial semantic segmentation model are adjusted based on the second classification loss to obtain the generator of the trained adversarial generative network model. Then, the parameters of the first discriminator of the adversarial generative network model are adjusted based on the first classification loss to obtain the trained target semantic segmentation model with higher segmentation accuracy.

[0097] Figure 5 A flowchart of an active sample screening method according to an embodiment of the present invention is provided. This method can be performed by, for example... Figure 1 The electronic device 110 shown performs, but is not limited to, its functions. For example... Figure 5 As shown, it may include the following steps:

[0098] S810: Input the samples to be screened into the target semantic segmentation model;

[0099] S820: Based on the discrimination results of the first discriminator and / or the second discriminator, determine whether to use the sample to be screened as an active sample.

[0100] For example, the sample pool to be screened typically stores numerous sample images. Training a semantic segmentation model requires labeling each sample image pixel-by-pixel with semantic labels, and then using the labeled sample images as training samples, resulting in high labeling costs. In this step, the samples to be screened are unlabeled sample images. The first semantic segmentation result is obtained by inputting the sample to be screened into the initial semantic segmentation model, and the generated semantic segmentation result is obtained by inputting the generated image into the initial semantic segmentation model. Here, the generated image is obtained by inputting the first feature image of the sample to be screened into the generative adversarial network model.

[0101] Furthermore, a second discriminator can be used to distinguish between the first semantic segmentation result and the generated semantic segmentation result to determine the source information of the first semantic segmentation result and the generated semantic segmentation result. That is, to determine whether the first semantic segmentation result comes from the training sample or the generated image, or to determine whether there is a difference between the first semantic segmentation result and the generated semantic segmentation result. If it cannot be determined whether the first semantic segmentation result comes from the training sample or the generated image, or if it is determined that there is no difference between the first semantic segmentation result and the generated semantic segmentation result, then it is determined that the training sample is not a new training sample in the training sample set of the model. If it can be determined that the first semantic segmentation result comes from the training sample, or the generated semantic segmentation result generates an image, or if it is determined that there is a difference between the first semantic segmentation result and the generated semantic segmentation result, then it is determined that the training sample is a new training sample in the training sample set of the model.

[0102] If the second discriminator makes a correct judgment, the sample to be screened will be marked as an active sample and the sample set used for model training will be updated.

[0103] Here, after determining that the sample to be screened is a new training sample in the model training sample set, the training sample is labeled as an active sample and the training sample set is updated. This enables the second discriminator in the target semantic segmentation model to actively screen the training sample. For unlabeled training samples, if the second discriminator can distinguish whether the output is the output result corresponding to the training sample or the output result corresponding to the generated image, then the sample image is helpful to improve the model performance and should be selected as an active sample.

[0104] By employing the above implementation method, the first semantic segmentation result obtained by inputting the sample to be screened into the above-mentioned target semantic segmentation model and the generated semantic segmentation result are judged to determine whether the sample is a new training sample in the model training sample set. If so, the sample to be screened is marked as an active sample, and the model training sample set is updated through the training sample, so as to realize the active learning of training samples by the target semantic segmentation model provided in this application. The active learning method of training samples provided in this implementation method does not require adjusting the percentage threshold parameter of top-k, and can realize the direct judgment of the training sample to be learned, simplifying the screening process of training samples, and screening out new training samples in the model training sample set that are more in line with the needs of the model.

[0105] Furthermore, the first discriminator of the generative adversarial network can be used to distinguish between the first feature image and the second feature image to determine the source information of the first feature image and the second feature image, that is, to determine whether the second feature image comes from the training sample or the generated image, or to determine whether there is a difference between the first feature image and the second feature image.

[0106] If the first discriminator determines the result to be correct, then the sample to be screened will be determined as a new training sample in the model training sample set.

[0107] If the discrimination result is correct, it indicates that the second feature image can be determined to be from the generated image, or the first feature image can be determined to be from the training sample. This proves that there are differences between the first feature image and the second feature image, as well as between the training sample and the generated image. Therefore, the training sample cannot be well parsed and segmented by the target semantic segmentation model provided in this application. Therefore, the training sample should be selected as a new training sample in the model training sample set.

[0108] If the judgment result is incorrect, the sample to be screened will not be determined as a new training sample in the model training sample set.

[0109] If the discrimination result is incorrect, it means that it cannot be determined that the second feature image comes from the generated image, or that the first feature image comes from the training sample. This proves that there is no difference between the first feature image and the second feature image, or between the training sample and the generated image. Therefore, the training sample can still be well parsed and segmented by the target semantic segmentation model provided in this application. Therefore, the training sample should not be selected as a new training sample in the sample set for model training.

[0110] By adopting the above implementation method, the generated image is input into the encoder of the initial semantic segmentation model to obtain the second feature image. Then, the first discriminator of the generative adversarial network is used to distinguish between the first feature image and the second feature image to determine whether the training sample is a new training sample in the model training sample set, so as to realize the active learning of the training sample by the first discriminator of the target semantic segmentation model provided in this application.

[0111] Figure 6 This is a schematic diagram of a training device 600 for a semantic segmentation model according to an embodiment of the present invention. Figure 6 As shown, the device 600 includes:

[0112] The first training module 61 is used to train the initial semantic segmentation model based on training samples and generated images to obtain the first semantic segmentation result and the semantic segmentation result of the generated image; the generated image is obtained by training an adversarial generative network model.

[0113] The second training module 62 is used to retrain the initial semantic segmentation model based on the first semantic segmentation result and the generated image semantic segmentation result to obtain the target semantic segmentation model.

[0114] In some possible implementations, the first training module 61 includes:

[0115] The first feature extraction unit is used to input training samples into the encoder of the initial semantic segmentation model to obtain the first feature image;

[0116] The first semantic segmentation result acquisition unit is used to input the first feature image into the decoder of the initial semantic segmentation model to obtain the first semantic segmentation result;

[0117] The segmentation cross-entropy loss calculation unit is used to calculate the segmentation cross-entropy loss between the first semantic segmentation result and the ground truth labeled sample;

[0118] The initial semantic segmentation model training unit is used to adjust the parameters of the initial semantic segmentation model based on the segmentation cross-entropy loss to obtain the trained initial semantic segmentation model. The initial semantic segmentation model is used to obtain the first semantic segmentation result based on the input training samples.

[0119] In some possible implementations, the first training module 61 includes:

[0120] The first feature image extraction unit is used to input training samples into the encoder of the initial semantic segmentation model to obtain the first feature image;

[0121] An image generation unit is used to input the first feature image into the generator of the generative adversarial network model to obtain a generated image;

[0122] The second feature image extraction unit is used to input the generated image into the encoder of the initial semantic segmentation model to obtain the second feature image;

[0123] The image semantic segmentation result acquisition unit is used to input the second feature image into the decoder of the initial semantic segmentation model to obtain the generated image semantic segmentation result.

[0124] In some implementations, device 600 further includes a third training module, which includes:

[0125] The L2 distance loss calculation unit is used to calculate the L2 distance loss between the generated image and the training samples;

[0126] The generator training unit is used to adjust the parameters of the generator of the adversarial generative network model based on L2 distance loss to obtain the generator of the trained adversarial generative network model.

[0127] The first discriminator discrimination unit is used to input the first feature image and the second feature image into the first discriminator of the adversarial generative network model. The first discriminator is used to determine whether the first feature image and the second feature image are derived from training samples or generated images.

[0128] The first classification loss calculation unit is used to calculate the first classification loss between the first feature image and the second feature image;

[0129] The first discriminator training unit is used to adjust the parameters of the first discriminator of the adversarial generative network model based on the first classification loss, so as to obtain the first discriminator of the trained adversarial generative network model.

[0130] In some implementations, device 600 further includes:

[0131] The semantic segmentation result discrimination module is used to use a second discriminator to discriminate between the first semantic segmentation result and the generated image semantic segmentation result; the second discriminator is used to determine whether the generated image semantic segmentation result comes from the training sample or the generated image;

[0132] The second classification loss calculation module is used to calculate the second classification loss of the first semantic segmentation result and the generated image semantic segmentation result;

[0133] The target semantic segmentation model training module is used to adjust the parameters of the initial semantic segmentation model based on the second classification loss to obtain the trained target semantic segmentation model.

[0134] Figure 7 This is a schematic diagram of an active sample screening device 700 according to an embodiment of the present invention. Figure 7 As shown, the active sample screening device 700 includes:

[0135] The sample input module 71 is used to input the sample to be screened into the target semantic segmentation model;

[0136] The determination module 72 is used to determine whether to use the sample to be screened as an active sample based on the discrimination results of the first discriminator and / or the second discriminator.

[0137] In some possible implementations, the active sample screening device 700 also includes:

[0138] The sample set update module is used to label the sample as an active sample and update the sample set used for model training if the first discriminator and / or the second discriminator correctly identify it.

[0139] In some possible implementations, the active sample screening device 700 also includes:

[0140] The second feature image acquisition unit is used to input the generated image into the encoder of the initial semantic segmentation model to obtain the second feature image;

[0141] The feature image discrimination unit is used to discriminate between the first feature image and the second feature image using the first discriminator of the generative adversarial network model;

[0142] If the judgment result is correct, the removed unit is used to determine the training sample as a new training sample in the model training sample set;

[0143] If the determination of a unit is incorrect, it is used to change the training sample to a new training sample in the model training sample set.

[0144] It should be understood that the embodiments of the training apparatus for the semantic segmentation model and the embodiments of the training method for the semantic segmentation model can correspond to each other, and similar descriptions can be found in the embodiments of the training method for the semantic segmentation model. To avoid repetition, further details are omitted here. Specifically, Figure 6 The device 600 shown, and Figure 7 The apparatus 700 shown can all execute the above-described semantic segmentation model training method embodiments, and the aforementioned and other operations and / or functions of each module in apparatus 600 and 700 are respectively for implementing the corresponding process in the above-described semantic segmentation model training method. For the sake of brevity, they will not be described in detail here.

[0145] The apparatuses 600 and 700 of the embodiments of the present invention have been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that these functional modules can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the semantic segmentation model training method embodiments of the present invention can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the mining data prediction method for the fully mechanized mining face disclosed in the embodiments of the present invention can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software modules can be located in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the semantic segmentation model training method embodiments described above.

[0146] Figure 8 This is a schematic block diagram of an electronic device 800 according to an embodiment of the present invention.

[0147] like Figure 8 As shown, the electronic device 800 may include:

[0148] The system includes a memory 810 and a processor 820. The memory 810 stores computer programs and transfers the program code to the processor 820. In other words, the processor 820 can retrieve and run the computer programs from the memory 810 to implement the methods described in the embodiments of the present invention.

[0149] For example, the processor 820 can be used to execute the above-described method embodiments according to instructions in the computer program.

[0150] In some embodiments of the present invention, the electronic device 820 may include, but is not limited to:

[0151] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0152] In some embodiments of the present invention, the memory 810 includes, but is not limited to:

[0153] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0154] In some embodiments of the present invention, the computer program may be divided into one or more modules, which are stored in the memory 810 and executed by the processor 820 to perform the method provided by the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the controller.

[0155] like Figure 8 As shown, the electronic device 800 may further include:

[0156] Transceiver 830, which can be connected to processor 820 or memory 810.

[0157] The processor 820 can control the transceiver 830 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 830 may include a transmitter and a receiver. The transceiver 830 may further include antennas, and the number of antennas may be one or more.

[0158] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0159] The present invention also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, one embodiment of the present invention also provides a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.

[0160] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Video Disc (DVD)), or a semiconductor medium (e.g., Solid State Disk (SSD)).

[0161] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0162] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.

[0163] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0164] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A training method for a semantic segmentation model, characterized in that, include: The initial semantic segmentation model is trained based on the training samples and the generated images to obtain the first semantic segmentation result and the semantic segmentation result of the generated images; The generated image is obtained by training an adversarial generative network model; The initial semantic segmentation model is retrained based on the first semantic segmentation result and the generated image semantic segmentation result to obtain a target semantic segmentation model, which includes the initial semantic segmentation model and an adversarial generative network model.

2. The method according to claim 1, characterized in that, The step of training the initial semantic segmentation model based on training samples and generated images to obtain the first semantic segmentation result includes: The training samples are input into the encoder of the initial semantic segmentation model to obtain the first feature image; The first feature image is input into the decoder of the initial semantic segmentation model to obtain the first semantic segmentation result; Calculate the segmentation cross-entropy loss between the first semantic segmentation result and the training sample; The parameters of the initial semantic segmentation model are adjusted based on the segmentation cross-entropy loss to obtain the trained initial semantic segmentation model, which is used to output a first semantic segmentation result based on the input training samples.

3. The method according to claim 1, characterized in that, The step of training the initial semantic segmentation model based on training samples and generated images to obtain semantic segmentation results for the generated images includes: The training samples are input into the encoder of the initial semantic segmentation model to obtain the first feature image; The first feature image is input into the generator of the adversarial generative network model to obtain the generated image; The generated image is input into the encoder of the initial semantic segmentation model to obtain the second feature image; The second feature image is input into the decoder of the initial semantic segmentation model to obtain the semantic segmentation result of the generated image.

4. The method according to claim 3, characterized in that, The adversarial generative network model is trained in the following manner: Calculate the L2 distance loss between the generated image and the training samples; The parameters of the generator of the adversarial generative network model are adjusted based on the L2 distance loss to obtain the generator of the trained adversarial generative network model. The first feature image and the second feature image are input into the first discriminator of the adversarial generative network model. The first discriminator is used to determine whether the first feature image and the second feature image originate from the training samples or the generated images. Calculate the first classification loss between the first feature image and the second feature image; The parameters of the first discriminator of the adversarial generative network model are adjusted based on the first classification loss to obtain the first discriminator of the trained adversarial generative network model.

5. The method according to claim 1, characterized in that, The method further includes: A second discriminator is used to distinguish between the first semantic segmentation result and the generated image semantic segmentation result; the second discriminator is used to determine whether the generated image semantic segmentation result originates from the training sample or the generated image; Calculate the second classification loss of the first semantic segmentation result and the generated image semantic segmentation result; The parameters of the initial semantic segmentation model are adjusted based on the second classification loss to obtain the trained target semantic segmentation model.

6. An active sample screening method, characterized in that, include: The samples to be screened are input into the target semantic segmentation model as described in any one of claims 1 to 5; Based on the discrimination results of the first discriminator and / or the second discriminator, determine whether to use the sample to be screened as an active sample.

7. A training device for a semantic segmentation model, characterized in that, include: The first training module is used to train the initial semantic segmentation model based on training samples and generated images to obtain the first semantic segmentation result and the semantic segmentation result of the generated image. The generated image is obtained by training an adversarial generative network model; The second training module is used to retrain the initial semantic segmentation model based on the first semantic segmentation result and the generated image semantic segmentation result to obtain the target semantic segmentation model.

8. An active sample screening device, characterized in that, include: A sample input module is used to input the samples to be screened into the target semantic segmentation model according to any one of claims 1 to 5; The determining module is used to determine whether to include the sample to be screened as an active sample based on the discrimination results of the first discriminator and / or the second discriminator.

9. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1-6.