A road scene semantic segmentation method and system with a self-evaluation mechanism

By constructing a self-evaluation model in the road scene semantic segmentation method, the problem that self-evaluation cannot be performed in the existing technology is solved, objective evaluation and unsupervised tuning of semantic segmentation results are achieved, and decision-making accuracy and algorithm optimization capabilities of the intelligent driving system are improved.

CN114332797BActive Publication Date: 2025-06-10RATE TECH (CHONGQING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111614152.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-06-10
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The existing road scene semantic segmentation algorithm cannot perform self-evaluation during actual use, resulting in the inability to provide an effective decision-making basis for intelligent driving systems and unsupervised tuning.

Method used

A road scene semantic segmentation method with a self-evaluation mechanism is adopted. By building an evaluation network and a self-evaluation model, the semantic segmentation results can be scored during actual use, thereby providing a decision-making basis for the intelligent driving system and supporting unsupervised tuning.

Benefits of technology

It realizes objective evaluation of semantic segmentation results during actual use, provides an effective decision-making basis, and supports unsupervised tuning, which improves the safety of intelligent driving systems and dynamic optimization capabilities of algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332797B_ABST
    Figure CN114332797B_ABST
Patent Text Reader

Abstract

The present invention discloses a road scene semantic segmentation method and system with a self-evaluation mechanism. The method includes: acquiring a road scene video in the vehicle driving environment, and performing semantic segmentation on a preset number of frames of street view original images to predict a segmentation mask image and motion optical flow information; constructing an evaluation network, training and optimizing the evaluation network using a video object segmentation data set to obtain a self-evaluation model; inputting the street view original image, the segmentation mask image and the motion optical flow information into the self-evaluation model to obtain a score of the semantic segmentation result. By pre-constructing the self-evaluation model, after obtaining the semantic segmentation result, using the result to input into the self-evaluation model, and then obtaining a relatively objective score of the semantic segmentation result, this score can objectively evaluate the confidence of the semantic segmentation result, and thus provide an effective decision-making basis for the intelligent driving system and also provide data support for unsupervised optimization in the application process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically to a road scene semantic segmentation method and system with a self-evaluation mechanism. Background Art

[0002] At present, with the continuous development of vehicle intelligent driving technology, road scene semantic segmentation technology, as a core technology in intelligent driving systems, has become the focus of research in this field. Road scene semantic segmentation technology can assist vehicles in densely perceiving the driving environment at the pixel level. During driving, the vehicle camera is used to obtain driving images and input them into the semantic segmentation algorithm. The algorithm automatically segments and classifies the images, and divides the perceived entire image into specific semantic information such as lanes, pedestrians, and vehicles, and then transmits it to the decision-making module that determines the vehicle's driving, so as to perform obstacle avoidance and environmental analysis modules, thereby ensuring safe driving.

[0003] However, the evaluation of existing road scene semantic segmentation algorithms needs to be performed on labeled data sets, and the segmentation accuracy of the algorithm is calculated based on the difference between the algorithm's segmentation results and the manual annotations. However, the performance of various algorithms based on deep learning is different in different data sets. When the algorithm is applied in practice, the specific measured performance can only be judged subjectively by the user's visual inspection. In this way, the practical performance of the algorithm cannot be objectively evaluated, and the accuracy of the information delivered to the decision-making end cannot be guaranteed. In addition, developers cannot obtain accurate feedback on the practical effects of the algorithm, and cannot debug and optimize the algorithm for actual usage scenarios without additional data annotation.

[0004] Therefore, how to provide a road scene semantic segmentation method with self-evaluation function is a problem that technicians in this field need to solve urgently. Summary of the invention

[0005] In view of this, the present invention provides a road scene semantic segmentation method and system with a self-evaluation mechanism. The method comes with a segmentation evaluation mechanism, which can give a segmentation result and a score (confidence) of the segmentation result during actual use, thereby solving the problem that the existing semantic segmentation algorithm cannot realize the self-evaluation function, and thus cannot provide an effective decision-making basis for the intelligent driving system, and cannot provide data support for unsupervised tuning during the application process.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] On the one hand, the present invention provides a road scene semantic segmentation method with a self-evaluation mechanism, the method comprising the following steps:

[0008] Semantic segmentation: Obtain the road scene video in the vehicle driving environment, and perform semantic segmentation on the original street view images of a preset number of frames to predict the segmentation mask image and the motion optical flow information;

[0009] Model construction: Construct an evaluation network, and use the video object segmentation dataset to train and optimize the evaluation network to obtain a self-evaluation model;

[0010] Evaluation result: Input the original street view image, the segmentation mask image, and the motion optical flow information into the self-evaluation model to obtain the score of the semantic segmentation result.

[0011] The beneficial effect of the present invention is that: by pre-constructing a self-evaluation model, after obtaining the semantic segmentation result, using this result to input the constructed self-evaluation model, and then obtaining a relatively objective score of the semantic segmentation result. This score can objectively evaluate the confidence of the semantic segmentation result, and thus provide an effective decision-making basis for the intelligent driving system, and can also provide data support for unsupervised optimization during the application process.

[0012] Further, the steps of the above model construction specifically include:

[0013] Step 1: Construct an evaluation network with a convolutional neural network as the main body;

[0014] Step 2: Based on the video object segmentation dataset, pre-train the evaluation network;

[0015] Step 3: Select the mask images and optical flow information predicted by multiple known algorithms in the video object segmentation dataset, and calculate the segmentation result score by annotating the selected data in the video object segmentation dataset;

[0016] Step 4: Use the mask image, the optical flow information, and the segmentation result score as training data to train the evaluation network;

[0017] Step 5: Use the segmentation mask image and the motion optical flow information predicted in the semantic segmentation step as tuning data to optimize the trained evaluation network to obtain a self-evaluation model.

[0018] Further, when the segmentation mask image is of one category, the evaluation result step specifically includes:

[0019] Obtain the RGB segmentation result map in the form of multiplying the segmentation mask image by the original street view image with a binary mask;

[0020] Input the RGB segmentation result map and the motion optical flow information into the evaluation model to obtain the score of the semantic segmentation result.

[0021] Further, when the segmentation mask image has multiple categories, the evaluation result step specifically includes:

[0022] Split the segmentation mask image according to categories to obtain multiple single-category mask images;

[0023] Respectively, multiply the single-category mask images with the street view original image in the form of binary mask to obtain multiple RGB segmentation result images;

[0024] Input each of the RGB segmentation result images and the corresponding motion optical flow information into the self-evaluation model to obtain the segmentation result scores for each category;

[0025] Average the segmentation result scores for each category to obtain the score of the final semantic segmentation result.

[0026] Further, the score of the semantic segmentation result is any value within 0 to 1. The higher the score value, the better the performance of the semantic segmentation result.

[0027] Further, the above road scene semantic segmentation method with a self-evaluation mechanism further includes:

[0028] Unsupervised tuning: Construct a loss function based on the score of the semantic segmentation result, and use the loss function for fine-tuning optimization. Through this process, online optimization without supervision (no additional manual annotation is required) can be directly performed according to the actual application scenario.

[0029] Further, the loss function is:

[0030] Loss = 1 - s

[0031] s = C 2 (It,M,F)

[0032] where Loss is the loss function, s is the score of the semantic segmentation result, C 2 is the evaluation model, It is the street view original image, M is the semantic segmentation mask image, and F is the motion optical flow information.

[0033] On the other hand, the present invention also provides a road scene semantic segmentation system with a self-evaluation mechanism, and the system includes:

[0034] A scene segmentation module, configured to obtain a road scene video in the vehicle driving environment, and perform semantic segmentation on a preset number of frames of street view original images to predict and obtain a segmentation mask image and motion optical flow information;

[0035] A model construction module, configured to construct an evaluation network, train and tune the evaluation network using a video object segmentation data set to obtain a self-evaluation model; and

[0036] The self-evaluation module is used to input the original street view image, the segmentation mask image, and the motion optical flow information into the self-evaluation model to obtain the score of the semantic segmentation result.

[0037] Furthermore, the above road scene semantic segmentation system with a self-evaluation mechanism further includes an unsupervised tuning module. The unsupervised tuning module is used to construct a loss function based on the score of the semantic segmentation result and use the loss function for fine-tuning optimization.

[0038] The main parts of the above system are the scene segmentation module and the self-evaluation module. The scene segmentation module uses a multi-scale fully convolutional neural network to perform semantic segmentation on road scene images. The self-evaluation module then performs unsupervised autonomous evaluation on the segmentation result and gives the performance score of the segmentation result. The core of this system in the present invention lies in the self-evaluation module. The inputs of this part are the segmentation result of the semantic segmentation algorithm, the original image of the current frame, and the motion optical flow information (i.e., the optical flow amplitude map). Using the original image and the optical flow amplitude map as the reference items of the segmentation result in space and time respectively to automatically score the segmentation result can give an objective evaluation of the algorithm result in the actual application scenario without manual annotation, and thus can assist the intelligent driving system to make more accurate decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0040] Figure 1 It is a schematic flow chart of a road scene semantic segmentation method with a self-evaluation mechanism provided by the present invention;

[0041] Figure 2 It is a schematic structural diagram of a road scene semantic segmentation system with a self-evaluation mechanism in an embodiment of the present invention;

[0042] Figure 3 It is a schematic network structure diagram of the scene segmentation module in an embodiment of the present invention;

[0043] Figure 4 It is a schematic network structure diagram of the self-evaluation module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] On the one hand, referring to the attached Figure 1 , an embodiment of the present invention discloses a road scene semantic segmentation method with a self-evaluation mechanism, and the method includes the following steps:

[0046] S1: Semantic segmentation: Obtain the road scene video in the vehicle driving environment, and perform semantic segmentation on the original street view images of a preset number of frames to predict the segmentation mask image and the motion optical flow information;

[0047] S2: Build a model: Build an evaluation network, and use the video object segmentation data set to train and optimize the evaluation network to obtain a self-evaluation model;

[0048] The above process of building the model specifically includes:

[0049] Step 1: Build an evaluation network with a convolutional neural network as the main body;

[0050] Step 2: Based on the video object segmentation data set, pre-train the evaluation network;

[0051] Step 3: Select the mask images and optical flow information predicted by various known algorithms in the video object segmentation data set, and calculate the segmentation result scores for the selected data in the video object segmentation data set;

[0052] Step 4: Use the mask images, optical flow information, and segmentation result scores as training data to train the evaluation network;

[0053] Step 5: Use the segmentation mask image and motion optical flow information predicted in the semantic segmentation step as tuning data to optimize the trained evaluation network to obtain a self-evaluation model.

[0054] S3: Evaluation result: Input the original street view image, segmentation mask image, and motion optical flow information into the self-evaluation model to obtain the score of the semantic segmentation result.

[0055] In actual evaluation, since the segmentation mask image may be an image of one type or an image after a combination of multiple types, in order to obtain a more accurate score, it is necessary to reasonably score according to the number of image types included in the segmentation mask image. Specifically, it is divided into two cases for processing:

[0056] (1) When the segmented mask image is of one category, the process of the above evaluation result specifically includes the following steps:

[0057] Step 1: Obtain the RGB segmentation result map in the form of multiplying the segmented mask image by the street view original image with a binary mask;

[0058] Step 2: Input both the RGB segmentation result map and the motion optical flow information into the evaluation model to obtain the score of the semantic segmentation result.

[0059] (2) When the segmented mask image is of multiple categories, the process of the above evaluation result specifically includes the following steps:

[0060] Step 1: Split the segmented mask image according to categories to obtain multiple single-category mask images;

[0061] Step 2: Respectively obtain multiple RGB segmentation result maps in the form of multiplying each single-category mask image by the street view original image with a binary mask;

[0062] Step 3: Input each RGB segmentation result map and the corresponding motion optical flow information into the self-evaluation model respectively to obtain the segmentation result scores of each category;

[0063] Step 4: Calculate the average of the segmentation result scores of each category to obtain the final score of the semantic segmentation result.

[0064] In this embodiment, the score of the semantic segmentation result is any value within 0 to 1. The higher the score value, the better the performance of the semantic segmentation result.

[0065] Preferably, the above road scene semantic segmentation method with a self-evaluation mechanism further includes:

[0066] Unsupervised tuning: Construct a loss function based on the score of the semantic segmentation result and use the loss function for fine-tuning optimization. Through this process, online optimization without supervision (no additional manual annotation required) can be directly performed according to the actual application scenario.

[0067] The above loss function is specifically:

[0068] Loss = 1 - s

[0069] s = C 2 (It, M, F)

[0070] where Loss is the loss function, s is the score of the semantic segmentation result, C 2 is the evaluation model, It is the street view original image, M is the semantic segmentation mask image, and F is the motion optical flow information.

[0071] On the other hand, see the appendix Figure 2, an embodiment of the present invention also discloses a road scene semantic segmentation system with a self-evaluation mechanism, which includes:

[0072] A scene segmentation module 1, configured to obtain a road scene video in the vehicle driving environment, perform semantic segmentation on a preset number of frames of street view original images, and predict a segmentation mask image and motion optical flow information;

[0073] A model construction module 2, configured to construct an evaluation network, train and optimize the evaluation network using a video object segmentation data set to obtain a self-evaluation model; and

[0074] A self-evaluation module 3, configured to input the street view original image, the segmentation mask image, and the motion optical flow information into the self-evaluation model to obtain a score of the semantic segmentation result.

[0075] In this embodiment, refer to Appendix Figure 2 and Appendix Figure 3 , the main body of the above scene segmentation module is a fully convolutional neural network (Fully Convolutional Networks, abbreviated as FCN). It uses multi-level and multi-scale fusion features to enhance the description ability of network features, and after fusing multiple paths of features, it divides into two paths. One path predicts the scene segmentation mask, and the other path predicts the motion information. The input of the network is the street view original images (assumed to be the t-th and t+1-th frames) in the road scene video captured by the in-vehicle camera for two consecutive frames, and the output is the semantic segmentation result of the current frame (t), that is, an N+1-layer probability map (segmentation mask image). Each layer of probability corresponds to a segmentation category (such as lane, pedestrian, vehicle), and the total number of categories N is determined by the annotation of the training data set. The extra layer corresponds to the background category; the output also includes the motion optical flow information between the current frame (t) and the next frame (t+1), that is, a two-dimensional optical flow prediction result, and the two dimensions respectively correspond to the horizontal displacement and vertical displacement of the pixel points. The specific parameters of this network are obtained through alternating training on a large-scale street view data set with annotations (such as the Cityscapes data set) and an optical flow data set with annotations (such as the SINTEL data set).

[0076] The scene segmentation module is mainly responsible for performing pixel-level segmentation and optical flow prediction on the road scene images obtained by the in-vehicle camera during driving. The obtained optical flow result is input into the self-evaluation part as a reference for scoring the segmentation result, and the obtained segmentation mask image will be further output to the decision-making algorithm in the intelligent driving system for the next step of driving planning, collision detection, and other decision-making generation processes.

[0077] In this embodiment, refer to Appendix Figure 2 and Figure 4, the main body of the self-evaluation module is a Convolutional Neural Network (CNN for short). Compared with the multi-layer fusion structure designed for rich features in the scene segmentation module, the network structure of the self-evaluation module is relatively simple in the feature extraction stage, and the focus is on the multi-dimensional comparison and evaluation process of the segmentation results.

[0078] The self-evaluation module introduces two types of reference information: spatial domain and time domain. The input of this part is the original street scene image (spatial domain reference), the motion optical flow information (time domain reference), and the segmentation mask image to be evaluated. The segmentation mask image to be evaluated is in the form of multiplying the binary mask with the original image to obtain the RGB segmentation result map and input it into the self-evaluation model.

[0079] The output of the self-evaluation model is a mask score from 0 to 1, which can directly give the quality score of the mask unsupervised according to the comparison between the mask and its reference information.

[0080] To achieve this goal, the self-evaluation module is pre-trained on the large-scale video object segmentation dataset DAVIS. Dozens of existing algorithms' predicted masks and their scores (i.e., the Jaccard scores calculated from the dataset annotations) on this dataset are selected as training data, and the L2 loss function is used to train the network, enabling the network to gradually have the ability to autonomously evaluate any mask.

[0081] After the pre-training is completed, the output of the scene segmentation module is used as the input to further optimize the network (still using the L2 loss function), and a small amount of annotation is required during the debugging process. After the debugging is completed, the two modules form an integrated network structure, which can simultaneously predict the segmentation mask and unsupervised evaluate the segmentation mask. In actual evaluation, for multi-category segmentation masks, the predicted masks of different categories are split and evaluated one by one (for example, for N segmentation categories, the scores of the masks of the 1st to Nth categories are evaluated respectively), and then the average value is taken as the score of the segmentation result.

[0082] The self-evaluation module autonomously evaluates the results of the segmentation part, gives a judgment on the quality of the segmentation results, and sends it to the algorithm in the decision-making part as a reference. The decision-making algorithm can weight the credibility of the segmentation information according to the self-evaluation score. For example, when the score of the current frame is relatively low, the segmentation results of the previous frames with higher scores are introduced to assist in the analysis, thereby improving the accuracy and safety of the overall system. This score can also be fed back to the algorithm itself for dynamic optimization of the algorithm in the actual scene, etc.

[0083] Preferably, the above road scene semantic segmentation system with a self-evaluation mechanism further includes an unsupervised optimization module, which is used to construct a loss function based on the score of the semantic segmentation result and perform fine-tuning optimization using the loss function.

[0084] The working process of the above road scene semantic segmentation system with a self-evaluation mechanism in an intelligent driving system will be specifically described as follows:

[0085] 1) During the driving process, the in-vehicle camera captures a road scene video and inputs every two consecutive street view images into the scene segmentation module.

[0086] 2) The scene segmentation module predicts a scene segmentation mask image and motion optical flow information based on the images.

[0087] 3) The self-evaluation module reads the mask and optical flow predicted by the scene segmentation module and returns a score.

[0088] 4) The segmentation mask and the score are simultaneously sent to the decision-making algorithm at the back end of the intelligent driving system. The decision-making algorithm determines the confidence of the segmentation information of the current frame in the decision-making based on the score: if the score of the segmentation result is too low, it will not be adopted, and the previous high-score segmentation result or information from signal sources such as radar in the system will be used for supplementation.

[0089] The process of unsupervised tuning of the above road scene semantic segmentation system with a self-evaluation mechanism in a practical scenario will be described in detail as follows:

[0090] 1) During the driving process, the algorithm simultaneously gives the segmentation result and its evaluation score.

[0091] 2) The network directly uses 1 minus the evaluation score as the loss function for fine-tuning training, and can directly perform unsupervised (no additional manual annotation required) online optimization.

[0092] Let the networks of the segmentation part and the self-evaluation part be C 1 and C 2 , the input of the segmentation part is the image I t 、I t+1 , and the output is the segmentation mask image M and the motion optical flow information F. Then we have:

[0093] C 1 (I t , I t+1 ) = M, F

[0094] The calculation process of the score s of the self-evaluation part is as follows:

[0095] s = C 2 (I t , M, F)

[0096] Given C 2What is learned during training is the Jaccard score, which ranges from 0 to 1. The higher the score, the better the mask performance. Therefore, during actual operation, the score s can be equivalent to the mask performance, and the overall system C can be optimized online by maximizing the score s (that is, minimizing 1 - s). 1 +C 2 The tuning loss function can be expressed as follows:

[0097] Loss = 1 - s = 1 - C 2 (I t , M, F) = 1 - C 2 (I t , C 1 (I t , I t+1 ))

[0098] It is not difficult to find that the road scene image semantic segmentation method in the present invention has a built - in evaluation function, which can simultaneously give the segmentation result and an objective score for this result in an unlabeled actual driving scene, providing more detailed reference information for the decision - making algorithm of the intelligent driving system, improving the safety of the overall system, and at the same time providing the possibility for unsupervised dynamic optimization of the algorithm.

[0099] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0100] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A road scene semantic segmentation method with a self-evaluation mechanism, characterized in that, it includes: Semantic segmentation: Obtain the road scene video in the vehicle driving environment, and perform semantic segmentation on the original street view images of a preset number of frames to predict the segmentation mask image and the motion optical flow information; Build a model: Build an evaluation network, and use the video object segmentation dataset to train and optimize the evaluation network to obtain a self-evaluation model; Evaluation result: Input the original street view image, the segmentation mask image, and the motion optical flow information into the self-evaluation model to obtain the score of the semantic segmentation result; Among them, building the model specifically includes: Taking a convolutional neural network as the main body, build an evaluation network; Based on the video object segmentation dataset, pre-train the evaluation network; Select the mask images and optical flow information predicted by multiple known algorithms in the video object segmentation dataset, and calculate the segmentation result scores obtained by annotating the selected data in the video object segmentation dataset; Use the mask image, the optical flow information, and the segmentation result score as training data to train the evaluation network; Use the segmentation mask image and the motion optical flow information predicted in the semantic segmentation step as tuning data to optimize the trained evaluation network to obtain a self-evaluation model.

2. A road scene semantic segmentation method with a self-evaluation mechanism according to claim 1, characterized in that, when the segmentation mask image is of one category, the evaluation result step specifically includes: Multiply the segmentation mask image with the original street view image in the form of a binary mask to obtain an RGB segmentation result map; Input the RGB segmentation result map and the motion optical flow information into the self-evaluation model to obtain the score of the semantic segmentation result.

3. A road scene semantic segmentation method with a self-evaluation mechanism according to claim 1, characterized in that, when the segmentation mask image is of multiple categories, the evaluation result step specifically includes: Split the segmentation mask image according to categories to obtain multiple single-category mask images; Multiply each single-category mask image with the original street view image in the form of a binary mask to obtain multiple RGB segmentation result maps; Input each RGB segmentation result map and the corresponding motion optical flow information into the self-evaluation model to obtain the segmentation result scores of each category; Take the average of the segmentation result scores of each category to obtain the score of the final semantic segmentation result.

4. A road scene semantic segmentation method with a self-evaluation mechanism according to claim 1, characterized in that, the score of the semantic segmentation result is any value within 0 to 1.

5. A road scene semantic segmentation method with a self-evaluation mechanism according to claim 1, characterized in that, it further includes: Unsupervised tuning: Build a loss function based on the score of the semantic segmentation result, and use the loss function for fine-tuning and optimization.

6. A road scene semantic segmentation method with a self-evaluation mechanism according to claim 5, characterized in that, the loss function is: Loss = 1 - s s = C 2 (I t , M, F) Among them, Loss is the loss function, s is the score of the semantic segmentation result, C 2 is the evaluation model, I t is the original street view image, M is the semantic segmentation mask image, and F is the motion optical flow information.

7. A road scene semantic segmentation system with a self-evaluation mechanism, characterized in that, it includes: A scene segmentation module, which is used to obtain the road scene video in the vehicle driving environment, and perform semantic segmentation on the original street view images of a preset number of frames, and predict the segmentation mask image and the motion optical flow information; A model construction module, which is used to construct an evaluation network, train and optimize the evaluation network using a video object segmentation dataset, and obtain a self-evaluation model; A self-evaluation module, which is used to input the original street view image, the segmentation mask image, and the motion optical flow information into the self-evaluation model to obtain a score of the semantic segmentation result; Among them, the specific steps of constructing the model include: Taking a convolutional neural network as the main body, construct an evaluation network; Based on the video object segmentation dataset, pre-train the evaluation network; Select the mask images and optical flow information predicted by a variety of known algorithms in the video object segmentation dataset, and calculate the segmentation result scores obtained by annotating the selected data in the video object segmentation dataset; Use the mask image, the optical flow information, and the segmentation result score as training data to train the evaluation network; Use the segmentation mask image and the motion optical flow information predicted in the semantic segmentation step as tuning data to tune the trained evaluation network to obtain a self-evaluation model.

8. The road scene semantic segmentation system with a self-evaluation mechanism according to claim 7, characterized in that, it further includes an unsupervised tuning module, and the unsupervised tuning module is used to construct a loss function based on the score of the semantic segmentation result, and perform fine-tuning optimization using the loss function.

Citation Information

Patent Citations

  • Adaptive adversarial learning-based urban traffic scene semantic segmentation method and system

    CN110111335A

  • Map element semantic segmentation robustness enhancement method and system

    CN112862839A