Model training method and device based on multi-point supervision and vehicle control method and device
By setting up intermediate supervision nodes and auxiliary dimension adapters in the feature extraction layer of the deep semantic segmentation model, explicitly supervising the intermediate layer and enhancing gradient propagation, the problems of gradient vanishing and insufficient optimization of the intermediate layer are solved, and the semantic segmentation performance of the model and the accuracy of street scene semantic understanding are improved.
Patent Information
- Application Number
- CN202510542891.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
During the training process, deep semantic segmentation models are prone to encounter problems such as gradient disappearance and insufficient intermediate layer optimization, which leads to a decline in model performance and affects the accuracy of street scene semantic understanding.
A model training method based on multi-point supervision is proposed. By setting up multiple intermediate supervision nodes in the feature extraction layer, explicitly supervising the intermediate layer of the model, introducing auxiliary dimension adapters to enhance gradient propagation, and optimizing model parameters through target loss value.
By explicitly supervising the intermediate layer and enhancing gradient propagation, the problems of gradient vanishing and insufficient optimization of intermediate features are alleviated, and the semantic segmentation performance of deep networks and the accuracy of street scene semantic understanding is improved.
Smart Images

Figure CN120070900A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, a vehicle control method, and a device based on multi-point supervision. Background Art
[0002] In the field of autonomous driving, in order to accurately perceive the driving environment, a deep semantic segmentation model is usually used to capture hierarchical features in the driving environment to identify road objects, lanes, drivable areas, etc. However, as the network depth of the semantic segmentation model increases, limitations such as gradient disappearance and insufficient optimization of intermediate layers will be faced during the model training process. These limitations reduce the potential of deep network architectures and affect the performance of model semantic segmentation. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a model training method, a vehicle control method, and a device based on multi-point supervision, aiming to improve the accuracy of street scene semantic understanding.
[0004] To achieve the above object, in the first aspect of the embodiments of this application, a model training method based on multi-point supervision is proposed. The method includes: Obtain a street scene image, an image semantic label of the street scene image, and an initial semantic segmentation model. The initial semantic segmentation model includes a feature extraction layer and an output layer, and a plurality of intermediate supervision nodes are provided in the feature extraction layer; Extract the target street scene semantic features of the street scene image and the intermediate street scene semantic features of the street scene image at each of the intermediate supervision nodes through the feature extraction layer, and predict corresponding intermediate street scene semantic labels based on each of the intermediate street scene semantic features through the output layer, and predict a target street scene semantic label based on the target street scene semantic features through the output layer; For each intermediate supervision node, predict an auxiliary street scene semantic label of the corresponding intermediate street scene semantic feature through the intermediate supervision node; Calculate the target loss value according to each of the intermediate street scene semantic labels, each of the auxiliary street scene semantic labels, the image semantic label, and the target street scene semantic label; Update the model parameters of the initial semantic segmentation model according to the target loss value to obtain a target semantic segmentation model.
[0005] In some embodiments, the calculating the target loss value according to each of the intermediate street scene semantic labels, each of the auxiliary street scene semantic labels, the image semantic label, and the target street scene semantic label includes: Calculate the negative entropy loss of each intermediate supervision node according to each of the intermediate street scene semantic labels; Calculate the cross-entropy loss according to the image semantic label and the target street view semantic label; Calculate the mutual information loss of each intermediate supervision node according to each auxiliary street view semantic label and the image semantic label; Perform a weighted sum of the negative entropy loss of each intermediate supervision node, the cross-entropy loss, and the mutual information loss of each intermediate supervision node to obtain the target loss value.
[0006] In some embodiments, the performing a weighted sum of the negative entropy loss of each intermediate supervision node, the cross-entropy loss, and the mutual information loss of each intermediate supervision node to obtain the target loss value includes: Determine the first weight and the second weight of each intermediate supervision node; Perform a weighted sum of the negative entropy loss of each intermediate supervision node and the corresponding first weight to obtain a first loss value; Perform a weighted sum of the mutual information loss of each intermediate supervision node and the corresponding second weight to obtain a second loss value; Perform a loss sum of the cross-entropy loss, the first loss value, and the second loss value to obtain the target loss value.
[0007] In some embodiments, each intermediate supervision node is connected to an auxiliary dimension adapter. After updating the model parameters of the initial semantic segmentation model according to the target loss value to obtain a target semantic segmentation model, the model training method based on multi-point supervision further includes: For each auxiliary dimension adapter, calculate the parameter gradient of the target loss value with respect to the model parameters of the auxiliary dimension adapter; Determine the parameter update step size; Update the model parameters of the auxiliary dimension adapter according to the parameter update step size and the parameter gradient.
[0008] In some embodiments, each intermediate supervision node is connected to an auxiliary dimension adapter. For each intermediate supervision node, predicting the auxiliary street view semantic label of the corresponding intermediate street view semantic feature through the intermediate supervision node includes: Obtain the street view semantic categories included in the target street view semantic label and the number of categories of the street view semantic categories; Call a target generator according to the street view semantic categories and the number of categories. Through the target generator, generate a synthetic street view image based on the intermediate street view semantic feature and the target street view semantic feature of each intermediate supervision node; Extract the intermediate synthetic street view semantic feature of the synthetic street view image at each intermediate supervision node through the feature extraction layer; For each intermediate supervision node, perform feature fusion on the intermediate street view semantic feature and the intermediate synthesized street view semantic feature of the intermediate supervision node to obtain a fused semantic feature of the intermediate supervision node; Predict an auxiliary street view semantic label of the corresponding fused semantic feature through the auxiliary dimension adapter connected by each intermediate supervision node.
[0009] In some embodiments, the generating, by the target generator, a synthesized street view image based on the intermediate street view semantic feature and the target street view semantic feature of each intermediate supervision node includes: Perform feature splicing on the intermediate street view semantic feature and the target street view semantic feature of each intermediate supervision node to obtain a first reference semantic feature; Perform multi-scale feature extraction on the reference semantic feature to obtain sub-semantic features of multiple scales; Obtain random noise, and perform feature fusion on the random noise and the sub-semantic features of multiple scales to obtain a second reference semantic feature; Generate the synthesized street view image by the target generator based on the second reference semantic feature.
[0010] To achieve the above object, a second aspect of the embodiments of the present application proposes a vehicle control method, and the method includes: Obtain a target street view image collected by a target vehicle; Perform street view semantic segmentation on the target street view image through a target semantic segmentation model to obtain a predicted street view semantic label; wherein, the target semantic segmentation model is trained according to the above-mentioned model training method based on multi-point supervision; Control the driving of the target vehicle according to the predicted street view semantic label.
[0011] To achieve the above object, a third aspect of the embodiments of the present application proposes a vehicle control device, and the device includes: An image acquisition module, configured to obtain a target street view image collected by a target vehicle; A semantic segmentation module, configured to perform street view semantic segmentation on the target street view image through a target semantic segmentation model to obtain a predicted street view semantic label; wherein, the target semantic segmentation model is trained according to the above-mentioned model training method based on multi-point supervision; A control module, configured to control the driving of the target vehicle according to the predicted street view semantic label.
[0012] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the model training method based on multi-point supervision in the first aspect or the vehicle control method in the second aspect.
[0013] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the model training method based on multi-point supervision in the first aspect or the vehicle control method in the second aspect.
[0014] The model training method based on multi-point supervision, vehicle control method, vehicle control device, electronic device and computer-readable storage medium in the embodiments of the present application obtain a street view image, an image semantic label of the street view image and an initial semantic segmentation model, and train the initial semantic segmentation model based on the street view image and the image semantic label. The initial semantic segmentation model includes a feature extraction layer and an output layer. In order to alleviate problems such as gradient disappearance and insufficient optimization of the intermediate layer that occur during the training of the initial semantic segmentation model, multiple intermediate supervision nodes are set in the feature extraction layer to explicitly supervise the intermediate layer of the model through the intermediate supervision nodes, so that the model can learn intermediate features related to the street view semantic understanding task. In order to capture semantic features that play a key role in street view semantic understanding in the driving environment, the target street view semantic features of the street view image and the intermediate street view semantic features of the street view image at each intermediate supervision node are extracted through the feature extraction layer. The corresponding intermediate street view semantic label is predicted based on each intermediate street view semantic feature through the output layer, and the target street view semantic label is predicted based on the target street view semantic feature through the output layer, so as to optimize the semantic segmentation performance of the model based on the intermediate street view semantic label and the target street view semantic label of each intermediate supervision node. For each intermediate supervision node, an auxiliary street view semantic label of the corresponding intermediate street view semantic feature is predicted through the intermediate supervision node, so as to assist in adjusting the parameters of the intermediate layer of the model based on the auxiliary street view semantic label, thereby improving the accuracy of intermediate feature extraction. According to each intermediate street view semantic label, each auxiliary street view semantic label, the image semantic label and the target street view semantic label, a target loss value is calculated to guide the training process of the initial semantic segmentation model based on the target loss value. The model parameters of the initial semantic segmentation model are updated according to the target loss value to obtain a target semantic segmentation model. By setting intermediate supervision nodes to perform explicit supervision on the feature learning process of the intermediate layer of the model, the parameters of the intermediate layer are optimized, and an additional gradient path is introduced through the intermediate supervision nodes, alleviating the gradient disappearance problem, thereby improving the accuracy of semantic segmentation. Description of the Drawings
[0015] Figure 1It is a flowchart of the model training method based on multi-point supervision provided by an embodiment of the present application; Figure 2 It is Figure 1 a flowchart of step S130 in Figure 3 It is Figure 2 a flowchart of step S220 in Figure 4 It is Figure 1 a flowchart of step S140 in Figure 5 It is Figure 4 a flowchart of step S440 in Figure 6 It is another flowchart of the model training method based on multi-point supervision provided by an embodiment of the present application; Figure 7 It is a flowchart of the vehicle control method provided by an embodiment of the present application; Figure 8 It is a schematic structural diagram of the vehicle control device provided by an embodiment of the present application; Figure 9 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0016] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0019] In the field of autonomous driving, in order to achieve accurate perception of the driving environment, a deep semantic segmentation model is usually adopted to capture hierarchical features in the driving environment to identify road objects, lanes, drivable areas, etc. However, as the network depth of the semantic segmentation model increases, limitations such as gradient disappearance and insufficient optimization of intermediate layers will be faced during the model training process. These limitations reduce the potential of deep network architectures and affect the performance of model semantic segmentation.
[0020] Based on this, the embodiments of the present application provide a model training method, a vehicle control method, a vehicle control device, an electronic device, and a computer-readable storage medium based on multi-point supervision, aiming to improve the accuracy of street scene semantic understanding.
[0021] The model training method, vehicle control method, vehicle control device, electronic device, and computer-readable storage medium based on multi-point supervision provided by the embodiments of the present application are specifically described through the following embodiments. First, the model training method based on multi-point supervision in the embodiments of the present application is described.
[0022] The model training method based on multi-point supervision provided by the embodiments of the present application relates to the field of artificial intelligence technology. The model training method based on multi-point supervision provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the model training method based on multi-point supervision, etc., but is not limited to the above forms.
[0023] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0024] Figure 1 is an alternative flowchart of a model training method based on multi-point supervision provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S110 to S150.
[0025] Step S110: Obtain a street view image, an image semantic label of the street view image, and an initial semantic segmentation model. The initial semantic segmentation model includes a feature extraction layer and an output layer, and multiple intermediate supervision nodes are set in the feature extraction layer. Step S120: Extract the target street view semantic features of the street view image and the intermediate street view semantic features of the street view image at each intermediate supervision node through the feature extraction layer, and predict the corresponding intermediate street view semantic labels based on each intermediate street view semantic feature through the output layer, and predict the target street view semantic label based on the target street view semantic feature through the output layer. Step S130: For each intermediate supervision node, predict the auxiliary street view semantic label of the corresponding intermediate street view semantic feature through the intermediate supervision node. Step S140: Calculate the target loss value according to each intermediate street view semantic label, each auxiliary street view semantic label, the image semantic label, and the target street view semantic label. Step S150: Update the model parameters of the initial semantic segmentation model according to the target loss value to obtain the target semantic segmentation model.
[0026] In step S110 of some embodiments, a street view image of the surrounding environment is collected according to the shooting device carried by the vehicle, and the street view image is semantically annotated to obtain an image semantic label. The image semantic label is used to identify and classify the semantic categories of each element in the street view image, such as lanes, buildings, pedestrians, drivable areas, and traffic signs, etc. The image semantic label is a pixel-level semantic label, and each pixel point in the street view image is marked as a specific semantic category. An initial semantic segmentation model is obtained. The initial semantic segmentation model is a semantic segmentation model to be trained. Through the initial semantic segmentation model, each pixel in the street view image can be assigned to a predefined semantic category. The initial semantic segmentation model can be designed according to the actual situation. The initial semantic segmentation model can be a fully convolutional neural network, a U-shaped network, etc. The initial semantic segmentation model includes a feature extraction layer and an output layer. The feature extraction layer is used to extract image feature representations useful for semantic segmentation from the street view image, and the output layer is used to predict the semantic category of the street view image according to the image feature representations extracted by the feature extraction layer.
[0027] As the network depth increases, problems such as gradient disappearance and insufficient optimization of intermediate features will occur in the deep learning network. To improve the semantic segmentation performance of the deep learning network, multiple intermediate supervision nodes are set in the feature extraction layer. Each intermediate supervision node is connected with an auxiliary dimension adapter. The auxiliary dimension adapter is added to the network layer where the intermediate supervision node is located in the form of a pluggable module to explicitly supervise the intermediate features output by the intermediate layer based on the intermediate supervision node and correct the features of the intermediate features, so that the deep layer of the feature extraction layer can accurately encode semantic-rich and discriminative feature representations to achieve street view semantic segmentation in different driving environments such as occlusion and weather changes.
[0028] Determine the number of nodes of the intermediate supervision nodes. The number of nodes is less than the number of network layers in the feature extraction layer. The intermediate supervision nodes can be randomly set in the network layers of the feature extraction layer. Only one intermediate supervision node can be set in one network layer. The feature extraction layer includes multiple network layers, and there are continuous network layers between every two adjacent intermediate supervision nodes.
[0029] In step S120 of some embodiments, the street view image is subjected to image feature extraction through the feature extraction layer to obtain the target street view semantic features output by the feature extraction layer in the last network layer, and the intermediate street view semantic features output by each intermediate supervision node. Each intermediate street view semantic feature is respectively input into the output layer for street view semantic segmentation to obtain the intermediate street view semantic label corresponding to each intermediate street view semantic feature. The intermediate street view semantic label is the semantic category predicted by the output layer according to the intermediate street view semantic feature. The target street view semantic features are input into the output layer for street view semantic segmentation to obtain the target street view semantic label. The target street view semantic label is the semantic category predicted by the output layer according to the target street view semantic features.
[0030] Please refer to Figure 2 , in some embodiments, step S130 may include but is not limited to steps S210 to S250: Step S210, obtaining the street view semantic categories included in the target street view semantic label and the number of categories of the street view semantic categories; Step S220, calling the target generator according to the street view semantic categories and the number of categories, and generating a synthetic street view image through the target generator based on the intermediate street view semantic features and the target street view semantic features of each intermediate supervision node; Step S230, extracting the intermediate synthetic street view semantic features of the synthetic street view image at each intermediate supervision node through the feature extraction layer; Step S240, for each intermediate supervision node, fusing the intermediate street view semantic features and the intermediate synthetic street view semantic features of the intermediate supervision node to obtain the fused semantic features of the intermediate supervision node; Step S250, predicting the auxiliary street view semantic labels of the corresponding fused semantic features through the auxiliary dimension adapters connected to each intermediate supervision node.
[0031] In step S210 of some embodiments, the street view semantic categories included in the target street view semantic label and the number of categories of the street view semantic categories are obtained. The street view semantic categories are used to indicate the semantic categories of the elements in the street view image, such as lanes, traffic signs, etc. The number of categories is the number of street view semantic categories included in the target street view semantic label. For example, if the target street view semantic label includes lanes, traffic signs, and buildings, the number of categories is 3.
[0032] In step S220 of some embodiments, in order to improve the semantic understanding ability of the model for complex street scene environments, the target generator is called according to the street scene semantic categories and the number of categories to expand the sample size. If the street scene semantic categories included in the target street scene semantic label are preset categories, the target generator is called. The preset categories can be set according to the actual situation. For example, in order to improve the accuracy of street scene semantic understanding under bad weather conditions, the preset categories can be bad weather categories such as rain, fog, wind, etc. In order to improve the accuracy of street scene semantic understanding under complex road conditions, the preset categories can be complex road condition categories such as mountain roads, rural roads, etc. If the preset categories are not included in the target street scene semantic label, then the number of categories of the street scene semantic categories is judged. A quantity threshold is preset in advance. If the number of categories is greater than or equal to the quantity threshold, it indicates that the street scene indicated by the street scene image is a complex scene, such as a scene with heavy traffic flow, a scene with a complex road environment, etc., then the target generator is called. If the number of categories is less than the quantity threshold, for each intermediate supervision node, the intermediate street scene semantic features output by the intermediate supervision node are input into the auxiliary dimension adapter connected to the intermediate supervision node for street scene semantic segmentation to obtain the auxiliary street scene semantic label corresponding to the intermediate street scene semantic features. The auxiliary street scene semantic label is the semantic category predicted by the auxiliary dimension adapter according to the intermediate street scene semantic features.
[0033] The target generator is a pre-trained generator and can be obtained through generative adversarial learning. The target generator generates images according to the intermediate street scene semantic features and the target street scene semantic features of each intermediate supervision node to obtain synthetic street scene images. The intermediate street scene semantic features contain the details and local structures of the street scene image, and the target street scene semantic features contain the global street scene semantics of the street scene image. The synthetic street scene images generated by combining the two can be consistent with the street scene image in terms of details and overall structure, improving the quality of image generation.
[0034] In step S230 of some embodiments, the feature extraction layer extracts image features from the synthetic street scene images to obtain the output features of the synthetic street scene images at each intermediate supervision node, and uses the output features as the intermediate synthetic street scene semantic features.
[0035] In step S240 of some embodiments, for each intermediate supervision node, the intermediate street scene semantic features output by the street scene image at the intermediate supervision node and the intermediate synthetic street scene semantic features output by the synthetic street scene image at the same intermediate supervision node are fused in features, so as to accurately express local details of the street scene such as traffic signs and driving road conditions at the same intermediate supervision node, and obtain the fused semantic features of the intermediate supervision node, thereby improving the street scene semantic segmentation performance of the auxiliary dimension adapter based on the fused semantic features.
[0036] In step S250 of some embodiments, for each intermediate supervision node, the fused semantic features output by the intermediate supervision node are input into the auxiliary dimension adapter connected to the intermediate supervision node for street scene semantic segmentation to obtain auxiliary street scene semantic labels. By using the auxiliary dimension adapter to assist in street scene semantic segmentation, the number of gradient propagation paths is increased, and the problem of gradient disappearance is alleviated to a certain extent. Moreover, the auxiliary street scene semantic labels output by the auxiliary dimension adapter can also assist in updating the model parameters of the initial semantic segmentation model, enabling the intermediate layer to output more accurate intermediate features, making up for the deficiency of the deep neural network in optimizing intermediate features, and greatly improving the street scene semantic segmentation performance of the initial semantic segmentation model.
[0037] Through the above steps S210 to S250, more accurate auxiliary street scene semantic labels can be obtained, so as to prompt the model to perform more accurate feature learning on the street scene image based on the auxiliary street scene semantic labels.
[0038] Please refer to Figure 3 , in some embodiments, step S220 may include but is not limited to steps S310 to S340: Step S310, perform feature splicing on the intermediate street scene semantic features and the target street scene semantic features of each intermediate supervision node to obtain the first reference semantic feature; Step S320, perform multi-scale feature extraction on the reference semantic feature to obtain sub-semantic features of multiple scales; Step S330, obtain random noise, and perform feature fusion on the random noise and the sub-semantic features of multiple scales to obtain the second reference semantic feature; Step S340, generate a synthetic street scene image based on the second reference semantic feature through the target generator.
[0039] In step S310 of some embodiments, in order for the target generator to generate an image consistent with the style of the street scene image, a comprehensive feature representation of the street scene image needs to be provided to the target generator. Therefore, according to the node positions of the intermediate supervision nodes in the feature extraction layer, the intermediate street scene semantic features output by each intermediate supervision node are sequentially spliced, and the spliced street scene semantic features are spliced with the target street scene semantic features again to obtain the first reference semantic feature.
[0040] In step S320 of some embodiments, the first reference semantic features include features of different scales such as building roads, roads, pedestrians, vehicles, etc. To enable the target generator to capture these semantic elements of different scales, multi-scale feature extraction is performed on the reference semantic features through a feature pyramid network to obtain sub-semantic features of multiple scales, where the scale refers to the resolution of the sub-semantic features. The feature pyramid network includes a bottom-up path and a top-down path. The bottom-up path includes N first network layers connected in series in sequence, and the top-down path includes N second network layers connected in series in sequence, where N is an integer greater than 1. The first reference semantic features are downsampled by a factor of two through the bottom-up path to obtain the first feature maps output by each first network layer.
[0041] If the first feature map output by the i-th first network layer is denoted as , then the resolution of the first feature map is of the first reference semantic features. The first feature map output by the last first network layer is upsampled by a factor of two through the first second network layer of the top-down path to obtain the second feature map output by the first second network layer. After the second feature map is upsampled by a factor of two through the second second network layer, it is feature-fused (i.e., feature-added) with the first feature map output by the (N - 1)-th first network layer to obtain the second feature map output by the second network layer. After the second feature map output by the previous network layer is upsampled by a factor of two through the i-th second network layer, it is feature-fused with the first feature map output by the (N - i + 1)-th first network layer to obtain the second feature map output by the i-th second network layer, where i is an integer greater than or equal to 2 and less than or equal to N. The above steps are repeated until multiple second feature maps are obtained, and these multiple second feature maps are used as sub-semantic features of multiple scales.
[0042] In step S330 of some embodiments, to generate diverse street view images, random noise is obtained, and the random noise can follow a normal distribution, a uniform distribution, etc. To enable the target generator to accurately recognize semantic elements of different scales, the random noise and the sub-semantic features of each scale are feature-fused to obtain the second reference semantic features.
[0043] In step S340 of some embodiments, the target generator may adopt different network architectures, such as a fully connected network, a convolutional network, etc. The second reference semantic feature is input into the target generator so that the target generator can capture the street view semantics implied in the second reference semantic feature and generate a synthetic street view image.
[0044] Through the above steps S310 to S340, a synthetic street view image can be obtained to expand the sample scale based on the synthetic street view image and achieve an accurate expression of the intermediate street view semantic feature output by the intermediate supervision node.
[0045] Please refer to Figure 4 , in some embodiments, step S140 may include but is not limited to steps S410 to S440: Step S410, calculate the negative entropy loss of each intermediate supervision node according to each intermediate street view semantic label; Step S420, calculate the cross-entropy loss according to the image semantic label and the target street view semantic label; Step S430, calculate the mutual information loss of each intermediate supervision node according to each auxiliary street view semantic label and the image semantic label; Step S440, perform weighted summation on the negative entropy loss, cross-entropy loss of each intermediate supervision node and the mutual information loss of each intermediate supervision node to obtain the target loss value.
[0046] In step S410 of some embodiments, in order to prevent overfitting caused by multi-point supervision, negative entropy is used as a regularization term to punish the intermediate features output by the intermediate supervision node by maximizing the negative entropy, reduce the uncertainty of the model's prediction results, and thus improve the generalization ability of the model to unknown driving environments. The intermediate street view semantic label is used to indicate the probability that each pixel point in the intermediate street view semantic feature belongs to a specific category. The negative entropy is calculated according to the intermediate street view semantic label corresponding to the m-th intermediate supervision node to obtain the negative entropy loss of the m-th intermediate supervision node. The calculation formula of the negative entropy loss is expressed as: , where, is the negative entropy loss of the -th intermediate supervision node; D is the sample data set; |D| is the number of street view images in the sample data set; i represents the i-th street view image in the sample data set; is the pixel set composed of pixel points; j represents the j-th pixel point; k represents the k-th semantic category; K is the number of classification categories; is the j-th pixel point of the intermediate street view semantic feature of the i-th street view image output by the -th intermediate supervision node; is The probability of belonging to the k-th semantic category; are the model parameters of the initial semantic segmentation model.
[0047] In step S420 of some embodiments, cross-entropy calculation is performed based on the image semantic label and the target street view semantic label to calculate the cross-entropy loss. The cross-entropy loss is used to measure the difference between the image semantic label and the target street view semantic label. The calculation formula of the cross-entropy loss is expressed as: , where, is the cross-entropy loss; D is the sample data set; |D| is the number of street view images in the sample data set; i represents the i-th street view image in the sample data set; is the pixel set composed of pixel points; j represents the j-th pixel point; K is the number of classification categories; represents the true probability that the j-th pixel point in the i-th street view image indicated by the image semantic label belongs to the k-th semantic category; represents the j-th pixel point in the i-th street view image; is the predicted probability of belonging to the k-th semantic category; are the model parameters of the initial semantic segmentation model.
[0048] In step S430 of some embodiments, in order to measure the distance between the intermediate features output by the intermediate supervision node and the ground truth, mutual information calculation is performed based on each auxiliary street view semantic label and the image semantic label to obtain the mutual information loss of each intermediate supervision node. By maximizing the mutual information loss, the hidden feature dimension and the ground truth dimension can be aligned, so that the auxiliary street view semantic label gradually approximates the image semantic label, thereby performing explicit supervision on the intermediate features. The calculation formula of the mutual information loss is expressed as: , where, is the mutual information loss of the -th intermediate supervision node; D is the sample data set; |D| is the number of street view images in the sample data set; i represents the i-th street view image in the sample data set; is the pixel set composed of pixel points; j represents the j-th pixel point; K is the number of classification categories; represents the true probability that the j-th pixel point in the i-th street view image indicated by the image semantic label belongs to the k-th semantic category; is the -th pixel point of the intermediate street view semantic feature of the i-th street view image output by the intermediate supervision node; is the The auxiliary street view semantic labels output by the auxiliary dimension adapters connected to the intermediate supervision nodes, where the auxiliary street view semantic labels indicate the predicted probability belonging to the k-th semantic class; is the parameter of the auxiliary dimension adapter connected to the
[0049] In step S440 of some embodiments, the negative entropy loss, cross-entropy loss, and mutual information loss of each intermediate supervision node are weighted and summed to obtain a target loss value, so as to guide the model optimization process based on the target loss value, thereby improving the street view semantic segmentation performance of the model.
[0050] Through the above steps S410 to S440, a target loss value can be obtained to explicitly supervise the intermediate features based on the target loss value, so that the intermediate layer can learn features related to the street view semantic segmentation task, and further enable the deep network layer to encode semantic-rich and discriminative feature representations to cope with diverse driving conditions.
[0051] Please refer to Figure 5 , in some embodiments, step S440 may include but is not limited to steps S510 to S540: Step S510, determining the first weight and the second weight of each intermediate supervision node; Step S520, weighting and summing the negative entropy loss of each intermediate supervision node and the corresponding first weight to obtain a first loss value; Step S530, weighting and summing the mutual information loss of each intermediate supervision node and the corresponding second weight to obtain a second loss value; Step S540, summing the cross-entropy loss, the first loss value, and the second loss value to obtain a target loss value.
[0052] In step S510 of some embodiments, the first weight and the second weight of each intermediate supervision node are obtained, where the first weight is the weight for the negative entropy loss, and the second weight is the weight for the mutual information loss.
[0053] In step S520 of some embodiments, for each intermediate supervision node, the negative entropy loss of the intermediate supervision node is multiplied by the first weight of the intermediate supervision node to obtain a first sub-loss. The first sub-losses of each intermediate supervision node are summed to obtain a first loss value.
[0054] In step S530 of some embodiments, for each intermediate supervision node, the mutual information loss of the intermediate supervision node is multiplied by the second weight of the intermediate supervision node to obtain a second sub-loss. The second sub-losses of each intermediate supervision node are summed to obtain a second loss value.
[0055] In step S540 of some embodiments, the cross-entropy loss, the first loss value, and the second loss value are added together to obtain a target loss value. The calculation formula of the target loss value is expressed as: , where, is the target loss value; is the cross-entropy loss; is the number of intermediate supervision nodes; represents the th intermediate supervision node; and are respectively the first weight and the second weight of the th intermediate supervision node; is the negative entropy loss of the th intermediate supervision node; is the mutual information loss of the th intermediate supervision node.
[0056] Through the above steps S510 to S540, the target loss value can be obtained, so as to prompt the model to learn features related to the street view semantic understanding task based on the target loss value, thereby improving the accuracy of street view semantic understanding.
[0057] Please refer to Figure 6 , in some embodiments, after step S140, the model training method based on multi-point supervision may further include but is not limited to steps S610 to S630: Step S610, for each auxiliary dimension adapter, calculate the parameter gradient of the target loss value with respect to the model parameters of the auxiliary dimension adapter; Step S620, determine the parameter update step size; Step S630, update the model parameters of the auxiliary dimension adapter according to the parameter update step size and the parameter gradient.
[0058] In step S610 of some embodiments, for each auxiliary dimension adapter, calculate the first-order partial derivative of the target loss value with respect to the model parameters of the auxiliary dimension adapter to obtain the parameter gradient.
[0059] In step S620 of some embodiments, a parameter update step size is obtained. The parameter update step size is the learning rate, which is used to indicate the step size of parameter update in each iteration of the model. The parameter update step size can directly affect the convergence speed of the model. To better balance the training speed and stability, it is necessary to dynamically adjust the parameter update step size. The parameter update step size can be adjusted based on the loss value. Specifically, if the difference between the target loss values of consecutive multiple training rounds is less than a preset difference threshold, the decay rate is multiplied by the parameter update step size, and the result of the multiplication is used as the updated parameter update step size. The decay rate is a number greater than 0 and less than 1. The parameter update step size can be adjusted based on the training round (number of iterations). The parameter update step size at the t-th training round can be expressed as: , where, is the parameter update step size at the t-th training round; r is the decay rate, and the value of the decay rate is in the range of (0, 1); t is the number of iterations; e is the base of the exponential operation; represents the multiplication operation.
[0060] In step S630 of some embodiments, the parameter update step size is multiplied by the parameter gradient, and the current model parameter of the auxiliary dimension adapter is subtracted from the result of the multiplication to update the current model parameter. The update process of the model parameter of the auxiliary dimension adapter is expressed as: , where, is the model parameter of the m-th auxiliary dimension adapter; is the parameter update step size; is the target loss value for the model parameter of the parameter gradient.
[0061] In the above steps S610 to S630, the auxiliary dimension adapter connected to each intermediate supervision node is optimized through the target loss value, so that the auxiliary dimension adapter can cooperate with the intermediate supervision node to more accurately supervise the intermediate features.
[0062] In step S150 of some embodiments, the initial semantic segmentation model and the auxiliary dimension adapter can be synchronously iteratively updated according to the same target loss value and the same parameter update step size until a preset iteration number threshold is reached. Specifically, the gradient of the target loss value with respect to the model parameter of the initial semantic segmentation model is calculated, and the model parameter of the initial semantic segmentation model is iteratively updated according to the parameter update step size and the gradient until the preset iteration number threshold is reached to obtain the target semantic segmentation model. The update process of the initial semantic segmentation model is expressed as: , where, are the model parameters of the initial semantic segmentation model; is the parameter update step size; is the target loss value For the model parameters gradient.
[0063] The street scene semantic segmentation model based on the deep learning network tends to deepen, and the traditional training paradigm that relies on single-point supervision at the network output end reduces the street scene semantic understanding performance of the model. To alleviate this problem, the embodiments of the present application improve the performance of the deep learning network by introducing intermediate multi-channel supervision (mutual information) and normalization (negative entropy). Theoretical convergence analysis of the training method of the embodiments of the present application shows that the training method has a convergence rate matching that of the standard stochastic gradient descent optimization method, and multi-point supervision does not damage the asymptotic convergence of the model. If the number of iterations of the model is T, the convergence rate is expressed as . Moreover, the training method of the embodiments of the present application is tested on multiple datasets, and the performance of the training method is better than that of the traditional training method, and the average intersection over union is increased by up to 9.19%.
[0064] Figure 7 is an optional flowchart of a vehicle control method provided by the embodiments of the present application, Figure 7 The method in may include but is not limited to steps S710 to S730.
[0065] Step S710, obtaining a target street scene image collected by a target vehicle; Step S720, performing street scene semantic segmentation on the target street scene image through a target semantic segmentation model to obtain a predicted street scene semantic label; wherein, the target semantic segmentation model is trained according to the above-mentioned model training method based on multi-point supervision; Step S730, controlling the driving of the target vehicle according to the predicted street scene semantic label.
[0066] In step S710 of some embodiments, the vehicle terminal is a control system installed on the target vehicle, and the vehicle terminal controls a photographing device carried on the target vehicle to collect an image of the surrounding environment of the vehicle to obtain a target street scene image.
[0067] In step S720 of some embodiments, the vehicle terminal obtains a target semantic segmentation model, the target semantic segmentation model includes a feature extraction layer and an output layer, extracts street scene semantic features of the target street scene image through the feature extraction layer, and outputs a predicted street scene semantic label based on the street scene semantic features through the output layer. The predicted street scene semantic label is a pixel-level semantic label for indicating the semantic category of each pixel point in the target street scene image.
[0068] In step S730 of some embodiments, the vehicle terminal outputs a vehicle control instruction for the target vehicle according to the predicted street view semantic label, and controls the target vehicle to travel according to the vehicle control instruction. The vehicle control instruction may be a vehicle steering instruction, a vehicle acceleration instruction, etc.
[0069] Through the above steps S710 to S730, the target semantic segmentation model is trained based on the multi-point supervision method, so that the intermediate layer of the target semantic segmentation model can learn semantic features associated with the street view semantic segmentation task, alleviating the problems of gradient disappearance and insufficient intermediate feature optimization caused by the deep learning network, and improving the accuracy of street view semantic understanding.
[0070] The embodiment of the present application also provides a model training device based on multi-point supervision, which can implement the above-mentioned model training method based on multi-point supervision. The model training device based on multi-point supervision includes: An acquisition module, configured to acquire a street view image, an image semantic label of the street view image, and an initial semantic segmentation model. The initial semantic segmentation model includes a feature extraction layer and an output layer, and multiple intermediate supervision nodes are provided in the feature extraction layer; A first street view semantic prediction module, configured to extract the target street view semantic features of the street view image and the intermediate street view semantic features of the street view image at each intermediate supervision node through the feature extraction layer, and predict the corresponding intermediate street view semantic label based on each intermediate street view semantic feature through the output layer, and predict the target street view semantic label based on the target street view semantic features through the output layer; A second street view semantic prediction module, configured to, for each intermediate supervision node, predict an auxiliary street view semantic label of the corresponding intermediate street view semantic feature through the intermediate supervision node; A calculation module, configured to calculate a target loss value according to each intermediate street view semantic label, each auxiliary street view semantic label, the image semantic label, and the target street view semantic label; An update module, configured to update the model parameters of the initial semantic segmentation model according to the target loss value to obtain a target semantic segmentation model.
[0071] In some embodiments, the calculation module is further configured to: Calculate the negative entropy loss of each intermediate supervision node according to each intermediate street view semantic label; calculate the cross-entropy loss according to the image semantic label and the target street view semantic label; calculate the mutual information loss of each intermediate supervision node according to each auxiliary street view semantic label and the image semantic label; perform weighted summation on the negative entropy loss of each intermediate supervision node, the cross-entropy loss, and the mutual information loss of each intermediate supervision node to obtain the target loss value.
[0072] In some embodiments, the calculation module is further configured to: Determine the first weight and the second weight of each intermediate supervision node; perform a weighted sum of the negative entropy loss of each intermediate supervision node and the corresponding first weight to obtain a first loss value; perform a weighted sum of the mutual information loss of each intermediate supervision node and the corresponding second weight to obtain a second loss value; perform a loss sum on the cross-entropy loss, the first loss value, and the second loss value to obtain a target loss value.
[0073] In some embodiments, the update module is further configured to: For each auxiliary dimension adapter, calculate the parameter gradient of the target loss value with respect to the model parameters of the auxiliary dimension adapter; determine the parameter update step size; update the model parameters of the auxiliary dimension adapter according to the parameter update step size and the parameter gradient.
[0074] In some embodiments, the second street view semantic prediction module is further configured to: Obtain the street view semantic categories included in the target street view semantic label and the number of categories of the street view semantic categories; call the target generator according to the street view semantic categories and the number of categories, and generate a synthetic street view image through the target generator based on the intermediate street view semantic features and the target street view semantic features of each intermediate supervision node; extract the intermediate synthetic street view semantic features of the synthetic street view image at each intermediate supervision node through the feature extraction layer; for each intermediate supervision node, fuse the intermediate street view semantic features and the intermediate synthetic street view semantic features of the intermediate supervision node to obtain the fused semantic features of the intermediate supervision node; predict the auxiliary street view semantic labels of the corresponding fused semantic features through the auxiliary dimension adapter connected to each intermediate supervision node.
[0075] In some embodiments, the second street view semantic prediction module is further configured to: Perform feature splicing on the intermediate street view semantic features and the target street view semantic features of each intermediate supervision node to obtain a first reference semantic feature; perform multi-scale feature extraction on the reference semantic feature to obtain sub-semantic features of multiple scales; obtain random noise, and fuse the random noise and the sub-semantic features of multiple scales to obtain a second reference semantic feature; generate a synthetic street view image through the target generator based on the second reference semantic feature.
[0076] Please refer to Figure 8 , the embodiments of the present application further provide a vehicle control device, which can implement the above vehicle control method. The vehicle control device includes: An image acquisition module 810, configured to acquire a target street view image collected by a target vehicle; A semantic segmentation module 820, configured to perform street view semantic segmentation on the target street view image through a target semantic segmentation model to obtain a predicted street view semantic label; wherein, the target semantic segmentation model is trained according to the above model training method based on multi-point supervision; A control module 830 is configured to control the driving of a target vehicle according to predicted street view semantic tags.
[0077] The specific implementation manner of this vehicle control device is basically the same as the specific embodiments of the above vehicle control method, and will not be elaborated here.
[0078] An embodiment of this application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned model training method or vehicle control method based on multi-point supervision. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0079] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes: A processor 910, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application; A memory 920, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 920 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 920, and are called by the processor 910 to execute the model training method or vehicle control method based on multi-point supervision of the embodiments of this application; An input / output interface 930, which is used to implement information input and output; A communication interface 940, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.); A bus 950, which transmits information between various components of the device (such as the processor 910, the memory 920, the input / output interface 930, and the communication interface 940); Among them, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are communicatively connected to each other inside the device through the bus 950.
[0080] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training method or vehicle control method based on multi-point supervision is implemented.
[0081] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0082] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0083] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.
[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0085] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0086] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0087] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0088] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0089] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0090] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0092] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.
Claims
1. A model training method based on multi-point supervision, characterized in that: The method comprises: Acquire a street view image, an image semantic label of the street view image, and an initial semantic segmentation model, wherein the initial semantic segmentation model includes a feature extraction layer and an output layer, and the feature extraction layer is provided with a plurality of intermediate supervision nodes; Extracting target street view semantic features of the street view image and intermediate street view semantic features of the street view image at each intermediate supervisory node through the feature extraction layer, predicting corresponding intermediate street view semantic labels based on each intermediate street view semantic feature through the output layer, and predicting target street view semantic labels based on the target street view semantic features through the output layer; For each intermediate supervisory node, predicting the auxiliary street view semantic label of the corresponding intermediate street view semantic feature through the intermediate supervisory node; Calculating the target loss value according to each of the intermediate street view semantic labels, each of the auxiliary street view semantic labels, the image semantic label and the target street view semantic label; The model parameters of the initial semantic segmentation model are updated according to the target loss value to obtain a target semantic segmentation model.
2. The method according to claim 1, characterized in that: The step of calculating the target loss value according to each of the intermediate street view semantic labels, each of the auxiliary street view semantic labels, the image semantic label and the target street view semantic label comprises: According to each of the intermediate street view semantic labels, calculating the negative entropy loss of each intermediate supervisory node; Calculating a cross entropy loss according to the image semantic label and the target street view semantic label; Calculating the mutual information loss of each intermediate supervisory node according to each of the auxiliary street view semantic labels and the image semantic labels; The negative entropy loss, the cross entropy loss and the mutual information loss of each intermediate supervisory node are weightedly summed to obtain the target loss value.
3. The method according to claim 2, characterized in that The weighted summation of the negative entropy loss, the cross entropy loss and the mutual information loss of each intermediate supervisory node to obtain the target loss value includes: Determining a first weight and a second weight of each intermediate supervisory node; Performing a weighted summation on the negative entropy loss of each intermediate supervisory node and the corresponding first weight to obtain a first loss value; Perform a weighted summation of the mutual information loss of each intermediate supervisory node and the corresponding second weight to obtain a second loss value; The cross entropy loss, the first loss value, and the second loss value are summed to obtain the target loss value.
4. The method according to claim 1, characterized in that: Each intermediate supervision node is connected to an auxiliary dimension adapter. After updating the model parameters of the initial semantic segmentation model according to the target loss value to obtain the target semantic segmentation model, the model training method based on multi-point supervision further includes: For each auxiliary dimension adapter, calculating a parameter gradient of the target loss value with respect to a model parameter of the auxiliary dimension adapter; Determine the parameter update step size; The model parameters of the auxiliary dimension adapter are updated according to the parameter update step size and the parameter gradient.
5. The method according to claim 1, characterized in that Each intermediate supervisory node is connected to an auxiliary dimension adapter, and for each intermediate supervisory node, the auxiliary street view semantic label of the corresponding intermediate street view semantic feature is predicted by the intermediate supervisory node, including: Obtaining the street view semantic category included in the target street view semantic tag and the number of the street view semantic category; Calling a target generator according to the street view semantic category and the number of categories, and generating a synthetic street view image through the target generator based on the intermediate street view semantic features of each intermediate supervisory node and the target street view semantic features; Extracting intermediate synthetic street view semantic features of the synthetic street view image at each intermediate supervisory node through the feature extraction layer; For each intermediate supervisory node, the intermediate street view semantic feature and the intermediate synthetic street view semantic feature of the intermediate supervisory node are subjected to feature fusion to obtain a fused semantic feature of the intermediate supervisory node; The auxiliary dimension adapter connected through each intermediate supervisory node predicts the auxiliary street view semantic label of the corresponding fused semantic feature.
6. The method according to claim 5, characterized in that The method of generating a synthetic street view image by the target generator based on the intermediate street view semantic features of each intermediate supervisory node and the target street view semantic features includes: Performing feature splicing on the intermediate street view semantic features of each intermediate supervisory node and the target street view semantic features to obtain a first reference semantic feature; Performing multi-scale feature extraction on the reference semantic features to obtain sub-semantic features of multiple scales; Acquire random noise, and perform feature fusion on the random noise and sub-semantic features of multiple scales to obtain a second reference semantic feature; The synthetic street view image is generated by the target generator based on the second reference semantic feature.
7. A vehicle control method, characterized in that: The method comprises: Acquire a target street view image collected by a target vehicle; Performing street view semantic segmentation on the target street view image through a target semantic segmentation model to obtain a predicted street view semantic label; wherein the target semantic segmentation model is trained according to the model training method based on multi-point supervision according to any one of claims 1 to 6; The target vehicle is controlled to travel according to the predicted street view semantic label.
8. A vehicle control device, characterized in that: The device comprises: An image acquisition module, used to acquire a target street view image collected by a target vehicle; A semantic segmentation module, configured to perform street view semantic segmentation on the target street view image through a target semantic segmentation model to obtain a predicted street view semantic label; wherein the target semantic segmentation model is trained according to the model training method based on multi-point supervision according to any one of claims 1 to 6; A control module is used to control the target vehicle to travel according to the predicted street view semantic label.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the model training method based on multi-point supervision as described in any one of claims 1 to 6 or the vehicle control method as described in claim 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the model training method based on multi-point supervision described in any one of claims 1 to 6 or the vehicle control method described in claim 7 is implemented.
Citation Information
Patent Citations
Street view understanding model training method based on large vision model assistance
CN118823719A