Image annotation method, device, equipment and medium based on autonomous driving scene

By adopting the image labeling method of multi-layer cascade feature processing in the autonomous driving scenario, the problem of low efficiency and accuracy of autonomous driving image labeling in the prior art is solved, and efficient and accurate image labeling is achieved.

CN119339380BActive Publication Date: 2025-05-06SHENZHEN RES INST OF BIG DATA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411885221.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In the prior art, the labeling efficiency and accuracy of autonomous driving images are low, especially the calculation load of arbitrary segmentation models when segmenting each picture is large, resulting in limited labeling efficiency and accuracy of the model on the autonomous driving images.

Method used

An image labeling method based on autonomous driving scenarios is proposed. By obtaining the image to be marked from the autonomous driving data, inputting it into the target model for potential feature extraction and multi-layer cascade feature processing, the features are enhanced to represent them at each level, and finally annotating the image according to the target highlighting area.

Benefits of technology

Through this method, the efficiency and accuracy of autonomous driving image annotation can be significantly improved, the calculation load can be reduced, the model can capture high-level semantic information of the image, and the generated annotation can be more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339380B_ABST
    Figure CN119339380B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an image annotation method, device, equipment and medium based on an autonomous driving scenario. An image to be annotated is obtained from autonomous driving data, and the image to be annotated is input into a target model to extract potential features to obtain a first image feature; the first image feature is adapted through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; the corresponding second image feature is represented by a layer-by-layer feature enhancement according to each feature category label through the mapping parameters corresponding to the multi-layer cascade feature processing space in the target model to obtain a target label image corresponding to each feature category label, and each target label image contains a target highlight area corresponding to the corresponding feature category label; the image to be annotated is annotated according to the target highlight area in each target label image. In this way, the efficiency and accuracy of annotating autonomous driving images can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image annotation method, device, equipment and medium based on an autonomous driving scenario. Background Art

[0002] Self-driving cars rely on an accurate understanding of their surroundings to make decisions. To help the system better understand and interpret visual information, different elements in an image can be annotated to improve the perception and safety of self-driving cars.

[0003] In related technologies, objects are usually automatically segmented from three primary color (Red, Green, Blue, RGB) images using a Segment Anything Model (SAM) as semantic labels. However, this segmentation method is relatively rough, and SAM involves a significant computational load when segmenting each image, which reduces the efficiency and accuracy of the model in labeling autonomous driving images. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose an image annotation method, device, equipment and medium based on an autonomous driving scenario, which can improve the efficiency and accuracy of annotating autonomous driving images.

[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes an image annotation method based on an autonomous driving scenario, the method comprising:

[0006] Acquire an image to be annotated from the autonomous driving data, and input the image to be annotated into a target model to extract potential features, thereby obtaining a first image feature;

[0007] Adapting the first image feature through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label;

[0008] By using mapping parameters corresponding to the multi-layer cascade feature processing space in the target model, the corresponding second image features are represented by feature enhancement layer by layer according to each feature category label, so as to obtain a target label image corresponding to each feature category label, wherein each target label image includes a target highlighting area corresponding to the corresponding feature category label;

[0009] The feature processing space is composed of a plurality of cascaded feature processing layers, and a mapping parameter corresponding to each feature processing layer is obtained by minimizing the difference between a sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by the feature processing space performing feature enhancement on a corresponding sample image feature based on a corresponding sample feature type label;

[0010] The image to be labeled is labeled according to the target highlighted area in each target label image.

[0011] Accordingly, a second aspect of an embodiment of the present application proposes an image annotation device based on an autonomous driving scenario, the device comprising:

[0012] An acquisition module, used for acquiring an image to be annotated from the autonomous driving data, and inputting the image to be annotated into a target model to extract potential features, thereby obtaining a first image feature;

[0013] An adaptation module, used for adapting the first image feature through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label;

[0014] An enhancement module is used to perform a layer-by-layer feature enhancement representation of the corresponding second image features according to each feature category label through the mapping parameters corresponding to the multi-layer cascade feature processing space in the target model, so as to obtain a target label image corresponding to each feature category label, and each target label image contains a target highlight area corresponding to the corresponding feature category label; wherein the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by minimizing the difference between the sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by the feature processing space performing feature enhancement representation on the corresponding sample image features based on the corresponding sample feature type label;

[0015] The labeling module is used to label the to-be-labeled image according to the target highlighted area in each target label image.

[0016] In some embodiments, the image annotation device based on the autonomous driving scene further includes a training module for:

[0017] Obtaining a sample image to be annotated corresponding to the sample target label image, and inputting the sample image to be annotated into a preset model to extract potential features, thereby obtaining a plurality of sample first image features;

[0018] Adapting the plurality of sample first image features through the plurality of feature category channels in the preset model to obtain sample second image features corresponding to each feature category channel; each feature category channel corresponds to a sample feature category label of the sample target label image;

[0019] By using the sample first mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are represented by feature enhancement layer by layer according to each sample feature category label, so as to obtain a sample prediction label image corresponding to each sample feature category label, wherein each sample prediction label image includes a sample highlight area corresponding to the corresponding sample feature category label;

[0020] Annotate the sample to-be-annotated image according to the sample highlighted area in each sample predicted label image to obtain a sample annotated image;

[0021] Constructing a target loss according to the difference between the sample annotated image and the corresponding area of ​​the sample target label image;

[0022] The parameters of the preset model are adjusted based on the target loss to obtain a target model.

[0023] In some embodiments, the training module is further used to:

[0024] Sequentially through each feature processing layer of the multi-layer cascade of the feature processing space in the preset model, the sample second image feature is processed according to the sample first mapping parameter of each feature processing layer, to obtain a layer sample prediction label image corresponding to each feature processing layer;

[0025] In each feature processing layer, based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image, determine the sample second mapping parameter of the current feature processing layer;

[0026] By using the sample second mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are feature enhanced layer by layer according to each sample feature category label to obtain a sample prediction label image corresponding to each sample feature category label.

[0027] In some embodiments, the training module is further used to:

[0028] In the current feature processing layer, based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image, determine the layer sub-loss of each sample feature category label in the current feature processing layer;

[0029] Constructing a target constraint model based on the layer sub-loss of each sample feature category label at the current feature processing layer, the sample first mapping parameter, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the layer sample prediction label image;

[0030] The target constraint model is solved to obtain the second mapping parameters of the samples of the current feature processing layer.

[0031] In some embodiments, the training module is further used to:

[0032] Constructing an objective function based on multiple layer sub-losses corresponding to the current feature processing layer based on multiple sample feature category labels;

[0033] For each sample feature category label, obtain the sample second image feature processed by the previous feature processing layer of the current feature processing layer, and the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and construct a first constraint function based on the sample second image feature processed by the previous feature processing layer, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the first scaling parameter;

[0034] Constructing a second constraint function based on the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer, the layer sample predicted label image and the first mapping parameter;

[0035] A target constraint model is constructed based on the target function, the first constraint function and the second constraint function.

[0036] In some embodiments, the training module is further used to:

[0037] Generate a target mapping function of the preset model through multiple sample second mapping parameters corresponding to multiple feature processing layers of the preset model; wherein the target mapping function is used to perform a layer-by-layer feature enhancement representation of the corresponding sample second image features according to each sample feature category label;

[0038] The sample second image feature corresponding to each sample category label is calculated according to the target mapping function to obtain a sample predicted label image corresponding to each sample feature category label.

[0039] In some embodiments, the training module is further used to:

[0040] Obtaining a first product according to the product of a plurality of first scaling parameters corresponding to the plurality of feature processing layers;

[0041] Obtaining a first mapping subfunction based on the product of the first product and the sample second image feature of the corresponding sample category label;

[0042] For each current feature processing layer, obtaining a product of a first scaling parameter subsequent to the current feature processing layer in the feature processing space to obtain a second product;

[0043] Determine the probability distribution of the corresponding sample category label at the current feature processing layer according to the second image feature of the sample processed by the previous feature processing layer of the current feature processing layer and the first scaling parameter;

[0044] Obtaining a first difference based on the probability distribution corresponding to the sample category label and the difference between the corresponding area in the sample target label image;

[0045] determining a second mapping sub-function according to the second product, the first learning parameter and the first difference;

[0046] A target mapping function is obtained according to the sum of a plurality of second mapping sub-functions corresponding to a plurality of feature processing layers in the feature processing space and the first mapping sub-function.

[0047] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the image annotation method based on the autonomous driving scene as described in any one of the embodiments of the first aspect of the present application when executing the computer program.

[0048] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the image annotation method based on the autonomous driving scene as described in any one of the embodiments of the first aspect of the present application.

[0049] The embodiment of the present application obtains a to-be-annotated image from autonomous driving data, and inputs the to-be-annotated image into a target model to extract potential features, thereby obtaining a first image feature; the first image feature is adapted through a plurality of feature category channels in the target model, thereby obtaining a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label; the corresponding second image feature is feature-enhanced layer by layer according to each feature category label through a mapping parameter corresponding to a multi-layer cascaded feature processing space in the target model, thereby obtaining a target label image corresponding to each feature category label, wherein each target label image includes a target highlight area corresponding to the corresponding feature category label; wherein the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by minimizing the difference between a sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by feature-enhancing the corresponding sample image feature by the feature processing space based on the corresponding sample feature type label; and the to-be-annotated image is annotated according to the target highlight area in each target label image. In this way, the model can focus on the features of a specific category through the adaptation of the category channel, improve the efficiency of segmenting targets of different categories without heavy load calculation, and optimize the feature map layer by layer through the multi-layer cascade feature processing space to gradually improve the quality of feature representation, so that the target model can efficiently capture the high-level semantic information in the image, and the generated annotated image more accurately reflects the target area. In summary, this application can improve the efficiency and accuracy of annotating autonomous driving images. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic diagram of the architecture of an image annotation system based on an autonomous driving scenario provided in an embodiment of the present application;

[0051] Figure 2 is a flowchart of an image annotation method based on an autonomous driving scenario provided in an embodiment of the present application;

[0052] Figure 3 is a flow chart of model training provided in an embodiment of the present application;

[0053] Figure 4 is a schematic diagram of an image processing process provided by an embodiment of the present application;

[0054] Figure 5 It is a schematic diagram of multi-layer cascade feature processing provided in an embodiment of the present application;

[0055] Figure 6 is an example diagram of the performance of the LAM provided in the embodiment of the present application;

[0056] Figure 7It is a schematic diagram of functional modules of an image annotation device based on an autonomous driving scenario provided in an embodiment of the present application;

[0057] Figure 8 It is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0061] Self-driving cars rely on an accurate understanding of their surroundings to make decisions. To help the system better understand and interpret visual information, different elements in an image can be annotated to improve the perception and safety of self-driving cars.

[0062] In the related art, objects are usually automatically segmented from RGB images using a Segment Anything Model (SAM) as semantic labels. However, this segmentation method is relatively rough, and SAM involves a significant computational load when segmenting each image, which reduces the efficiency and accuracy of the model in labeling autonomous driving images.

[0063] Based on this, the embodiments of the present application provide an image annotation method, device, equipment and medium based on an autonomous driving scenario, which can improve the efficiency and accuracy of annotating autonomous driving images.

[0064] The image annotation method, device, equipment and medium based on the autonomous driving scenario provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the image annotation system based on the autonomous driving scenario in the embodiments of the present application is described.

[0065] Please refer to Figure 1In some implementations, an embodiment of the present application provides an image annotation system based on an autonomous driving scenario, including a terminal 11 and a server 12.

[0066] Exemplarily, the terminal 11 may be a hardware device installed on the autonomous driving vehicle, including but not limited to an onboard computer, an embedded system, a sensor (such as a camera, a laser radar, etc.), and an edge computing device. The terminal 11 may be responsible for collecting environmental information around the autonomous driving vehicle, such as taking an image through a camera. The terminal 11 may also perform preliminary processing on the collected data, such as data compression, pre-labeling, etc., in order to reduce the burden on the server side, and send the collected data (or the data after preliminary processing) to the server side 12 for further processing.

[0067] Furthermore, the server 12 can be a powerful computing resource set located in a data center or cloud platform, responsible for processing a large amount of data from the terminal and performing complex computing tasks. The server 12 can receive data transmitted from the terminal 11, and use the target model to extract, classify and optimize features, generate high-fidelity semantic annotations, and return the generated annotation results to the terminal 11.

[0068] The image annotation method based on the autonomous driving scenario in the embodiment of the present application can be illustrated by the following embodiment.

[0069] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for enabling the normal operation of the embodiment of the present application will be obtained.

[0070] In the embodiment of the present application, the image annotation device based on the autonomous driving scene will be described from the perspective of the image annotation device based on the autonomous driving scene, and the image annotation device based on the autonomous driving scene can be specifically integrated in a computer device. Figure 2 , Figure 2 This is a flowchart of the steps of the image annotation method based on the autonomous driving scene provided in the embodiment of the present application. The embodiment of the present application takes the image annotation device based on the autonomous driving scene as an example, which is specifically integrated on a terminal or a server. When the processor on the terminal or the server executes the program instructions corresponding to the image annotation method based on the autonomous driving scene, the specific process is as follows:

[0071] Step 101: obtain an image to be annotated from the autonomous driving data, and input the image to be annotated into a target model to extract potential features to obtain a first image feature.

[0072] In some embodiments, in order to ensure that the target model can recognize useful information that is consistent with the actual application scenario, the images to be labeled can be obtained from the autonomous driving data, and the images to be labeled can be input into the visual large model in the target model to extract image features, so as to learn the basic patterns of the images to be labeled.

[0073] The autonomous driving data can be various information collected by the system during operation, including but not limited to image data and other environmental perception data obtained through various sensors (such as cameras, radars, lidars, etc.). The autonomous driving data can also be historical autonomous driving data, which is used to analyze the scene. The autonomous driving data is the key to the system's understanding and interpretation of the surrounding environment.

[0074] The unlabeled images can be images that have been captured by the autonomous driving data system but have not yet been analyzed and labeled in detail to identify specific objects or areas in the image. The unlabeled images usually contain a lot of information, but due to the lack of labels, this information has not been systematically utilized.

[0075] Among them, the target model can be a model for processing autonomous driving data. The target model has the ability to extract potential features from the images to be annotated, and can optimize the identified features through a multi-level optimization mechanism, thereby achieving efficient annotation of the images to be annotated. The target model can focus on the features of a specific category through the adaptation of the category channel, and optimize the feature map layer by layer through a multi-layer cascade feature processing space to gradually improve the quality of feature representation.

[0076] The first image feature may be an initial level feature representation extracted from the image to be annotated. These features are extracted from the image to be annotated by the target model and represent the model's basic understanding of the image content.

[0077] In some embodiments, the target model proposed in this application is first introduced. The full name of the target model can be Label Anything Model, referred to as LAM, which can include the following three parts:

[0078] The first is a pre-trained Vision Transformer (ViT), which is used as the backbone to extract the latent features (first image features) of each input image to be annotated. The first image features are high-level representations of the image to be annotated, capturing the local and global information necessary to perform subsequent tasks.

[0079] The second is the Semantic Class Adapter (SCA), which can fuse the hidden features extracted by the visual transformer into the corresponding C feature category channels, where C is the number of semantic categories in the dataset for training the model. The semantic category adapter uses a layer of conv1x1 to adapt the semantic category, involving relatively few adjustable parameters, which can improve the efficiency of model training.

[0080] The third is the optimized unfolding operator (OptOU), which is implemented through a multi-layer cascade feature processing space, aiming to optimize the second image features output by the semantic category adapter so that the target model can generate high-fidelity annotations. Specifically, OptOU contains multiple feature processing layers, in each layer, the output layer prediction label and the real label can be aligned as much as possible, thereby obtaining the target highlight area in each label image.

[0081] Next, the process of inputting the image to be annotated into the target model to extract potential features and obtain the first image feature is introduced.

[0082] Exemplarily, when the image to be annotated is input into the visual big model (ViT) of the target model, the visual big model can transform the original pixel data of the image to be annotated into a higher-level, more abstract feature representation, namely the first image feature, through a series of encoding operations. Specifically, the visual big model will identify different areas in the image to be annotated and convert the visual information of different areas into mathematical feature vectors, namely the first image feature. For example, the visual big model may recognize that there are four-cornered objects (vehicles) in the image and that there are moving human-shaped objects (pedestrians) in certain positions.

[0083] By extracting the first image features, even without direct manual annotation, the target model can effectively capture the local and global information in the image to be annotated, so that the target model can accurately understand and classify the various elements in the image to be annotated, so as to facilitate the subsequent transfer of the first image features to the next component of the target model for further processing and refinement.

[0084] Step 102, adapting the first image feature through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label.

[0085] In some embodiments, in order to enable the target model to focus more on features of a specific category, the first image features can be refined into specific semantic categories through a semantic category adapter in the target model, that is, mapped to channels corresponding to different categories, thereby removing redundant information and retaining important features related to the specific category, improving the accuracy of classification and segmentation, and improving the efficiency of annotating images to be annotated.

[0086] Among them, the feature category channel can be a channel in the semantic category adapter of the target model, which is used to map the first image feature to a specific semantic category. Each feature category channel corresponds to a semantic category label in the data set. The semantic category adapter realizes feature conversion through a layer of convolution (conv1x1) operation.

[0087] The second image feature may be a feature corresponding to each feature category channel obtained after processing in the semantic category adapter. The second image feature is generated from the first image feature after adaptation, and each feature category channel may output a second image feature.

[0088] The feature category label may be a specific semantic category identifier corresponding to each feature category channel. For example, in an autonomous driving scenario, the feature category labels may include "pedestrian", "vehicle", "traffic light", etc.

[0089] For example, there is an image to be annotated in an autonomous driving scene, which contains multiple elements such as pedestrians, vehicles, and trees. The semantic recognition adapter of the target model extracts the first image feature from the image to be annotated, but the first image feature has not yet clearly distinguished these elements. The first image feature is adapted to different feature category channels with pre-trained parameters through the semantic category adapter, and each feature category channel corresponds to a specific semantic category (such as pedestrian channel, vehicle channel, etc.). In this way, the output (second image feature) of the semantic category adapter can better reflect the characteristics of each element in the image. For example, the semantic category adapter can map features related to pedestrians to the "pedestrian" channel and output the second image features corresponding to the "pedestrian" channel, map features related to vehicles to the "vehicle" channel, output the second image features corresponding to the "vehicle" channel, and so on. Each feature category channel carries information about a specific category, so that the target model can more accurately identify and annotate different elements in the image in subsequent segmentation and classification tasks.

[0090] By adapting the semantic category adapter, the target model can focus more on the features of a specific category, improve the accuracy of classification and segmentation, and facilitate the subsequent further enhanced representation of the second image features corresponding to each feature category channel.

[0091] Step 103, through the mapping parameters corresponding to the multi-layer cascaded feature processing space in the target model, the corresponding second image features are feature enhanced layer by layer according to each feature category label, so as to obtain the target label image corresponding to each feature category label, and each target label image contains the target highlighted area corresponding to the corresponding feature category label; wherein the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by minimizing the difference between the sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by the feature processing space performing feature enhancement representation on the corresponding sample image features based on the corresponding sample feature type label.

[0092] In some embodiments, in order to further optimize the feature representation corresponding to each feature category label to generate higher fidelity semantic annotations, the target model can further enhance the representation of the second image features layer by layer through the multi-layer cascade feature processing space corresponding to OptOU, so that the generated target label image can more accurately reflect the target area.

[0093] The multi-layer cascade feature processing space may be a processing space composed of multiple cascaded feature processing layers in the target model. Each feature processing layer in the feature processing space has specific mapping parameters, and the mapping parameters are used to optimize the feature representation layer by layer. The optimization effect can be accumulated through multi-layer cascades.

[0094] The mapping parameters may be parameters used to control feature processing in a multi-layer cascade feature processing space. Each feature processing layer has a corresponding mapping parameter, which is used to convert features into a form that is more conducive to subsequent classification tasks.

[0095] The target label image may be a representation image of the feature corresponding to each feature category channel that is further enhanced after the second image feature corresponding to the feature category channel is processed by a multi-level feature processing space.

[0096] The target highlighted area may be a specific area that is correctly marked and highlighted in the target label image, and each target highlighted area of ​​the target label corresponds to a specific feature category channel, that is, corresponds to a specific feature category label.

[0097] The sample target label image may be an image sample with correct annotations used in the training process, which is used to train the model to learn the correct annotation method.

[0098] Among them, the sample prediction label image can be a label image generated by a preset model based on sample image features, which is used to evaluate the prediction performance of the preset model.

[0099] Among them, the sample feature type label can be a label used to annotate different semantic categories in the sample image to be annotated in the data used to train the preset model. When training the preset model, the features in the sample image to be annotated can be divided according to the sample feature type label.

[0100] Among them, the sample image feature can be a feature representation extracted from the sample image to be annotated, and the sample image feature can be a second image feature after a semantic category adapter, which is used for subsequent feature enhancement and sample feature type label generation.

[0101] In some embodiments, after acquiring the second image feature, in a multi-layer cascaded feature processing space, each second image feature can be optimized layer by layer by each feature processing layer using pre-adjusted mapping parameters, and the features processed by each feature processing layer will be passed to the next feature for further enhancement. The last layer of the feature processing space will output multiple target highlighted areas corresponding to the multiple feature category labels obtained by the final enhancement.

[0102] Exemplarily, after using a semantic category adapter to map the first image feature extracted by the visual large model to a specific semantic category, at least one second image feature corresponding to a different feature category label is obtained, such as a second image feature corresponding to a pedestrian, a second image feature corresponding to a vehicle, etc.

[0103] Furthermore, if the feature processing space includes three feature processing layers, each layer has its own mapping parameters. The first layer uses the corresponding mapping parameters to process each second image feature to obtain the first optimized feature representation, the second layer continues to use the corresponding mapping parameters to optimize the first optimized feature representation output by the first layer to obtain the second optimized feature representation, and the third layer uses a set of mapping parameters to further optimize the second optimized feature representation to obtain the target label image corresponding to each feature category label.

[0104] For example, in the process of training the preset model and finally obtaining the target model, the mapping parameters can be optimized by establishing a constraint model by comparing the difference between the sample target label image (the image with the correct annotation) and the sample predicted label image (the label image generated by the preset model), so that the feature processing layer can identify the label image that is closer to the sample feature type label according to the adjusted mapping parameters. For example, if the pedestrian area predicted by the model deviates from the actual pedestrian area, the mapping parameters are adjusted to reduce the deviation.

[0105] Through multi-layer cascade feature processing space, each feature processing layer will optimize the second image features for specific feature category labels. The optimization of each layer is based on the previous layer, which can accumulate the improvement effect and continuously reduce the gap between the generated label image and the real label, generating a clearer and more accurate target highlighting area.

[0106] Step 104 , annotating the image to be annotated according to the target highlighted area in each target label image.

[0107] In some embodiments, in order to improve the accuracy and consistency of annotation, the images to be annotated may be annotated according to the target highlighted areas in each label image to ensure that the label images generated by the model can accurately reflect the various semantic categories in the input image.

[0108] Exemplarily, the target highlighted region represents the part of the image to be labeled that is most likely to belong to a specific feature category label. According to the position of the target highlighted region in the label image, these regions can be applied to the image to be labeled. Specifically, a region with the same position and range as that in the target label image can be drawn on the image to be labeled and assigned a corresponding feature category label.

[0109] The embodiment of the present application obtains a to-be-annotated image from autonomous driving data, and inputs the to-be-annotated image into a target model to extract potential features, thereby obtaining a first image feature; the first image feature is adapted through a plurality of feature category channels in the target model, thereby obtaining a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label; the corresponding second image feature is feature-enhanced layer by layer according to each feature category label through a mapping parameter corresponding to a multi-layer cascaded feature processing space in the target model, thereby obtaining a target label image corresponding to each feature category label, wherein each target label image includes a target highlight region corresponding to the corresponding feature category label; wherein the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameter corresponding to each feature processing layer is obtained by minimizing the difference between a sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by feature-enhancing the corresponding sample image feature by the feature processing space based on the corresponding sample feature type label; and the to-be-annotated image is annotated according to the target highlight region in each target label image. In this way, the model can focus on the features of a specific category through the adaptation of the category channel, improve the efficiency of segmenting targets of different categories without heavy load calculation, and optimize the feature map layer by layer through the multi-layer cascade feature processing space to gradually improve the quality of feature representation, so that the target model can efficiently capture the high-level semantic information in the image, and the generated annotated image more accurately reflects the target area. In summary, this application can improve the efficiency and accuracy of annotating autonomous driving images.

[0110] Please refer to Figure 3 In some implementations, in order to enable the preset model to learn useful patterns and rules from the data, and then to make accurate predictions or perform tasks on new, unprocessed data, the preset model can be trained to maintain high performance on new data, thereby reducing reliance on manual intervention and improving processing efficiency and reliability of results. For example, the target model can be trained in the following ways:

[0111] Step 201, obtaining a sample image to be annotated corresponding to the sample target label image, and inputting the sample image to be annotated into a preset model to extract potential features, thereby obtaining a plurality of sample first image features;

[0112] Step 202, adapting multiple sample first image features through multiple feature category channels in a preset model to obtain sample second image features corresponding to each feature category channel; each feature category channel corresponds to a sample feature category label of the sample target label image;

[0113] Step 203, by using the sample first mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are represented by feature enhancement layer by layer according to each sample feature category label, and a sample prediction label image corresponding to each sample feature category label is obtained, and each sample prediction label image contains a sample highlight area corresponding to the corresponding sample feature category label;

[0114] Step 204, labeling the sample to-be-labeled image according to the sample highlighted area in each sample predicted label image to obtain a sample labeled image;

[0115] Step 205, constructing a target loss according to the difference between the sample annotated image and the corresponding area of ​​the sample target label image;

[0116] Step 206, adjusting the parameters of the preset model based on the target loss to obtain the target model.

[0117] Among them, the sample target label image can be an image sample with correct annotations used in the training process, which is used to train the preset model to learn the correct annotation method.

[0118] The sample image to be labeled may be an image from which labels are removed from the sample target label image, or may be another separately set training sample image.

[0119] Among them, the preset model can be a model for processing autonomous driving data, and the preset model can be trained at least once to obtain a trained target model.

[0120] Among them, the sample first image feature can be an initial level feature representation extracted from the sample image to be labeled by the visual large model of the preset model, and the sample first image feature represents the preset model's basic understanding of the sample image to be labeled.

[0121] The sample second image feature may be a feature corresponding to each feature category channel obtained after processing by a semantic category adapter of a preset model. The sample second image feature is generated from the sample first image feature after adaptation, and each feature category channel may output a sample second image feature.

[0122] The sample feature category label may be a label used to label different semantic categories in the sample image to be labeled in the data used to train the preset model. When training the preset model, the features in the sample image to be labeled may be divided according to the sample feature type label.

[0123] The sample first mapping parameter may be a parameter used to control feature processing in a multi-layer cascade feature processing space. Each feature processing layer has a corresponding adjustable sample first mapping parameter, which is used to convert the feature into a form closer to the sample feature category label.

[0124] Among them, the sample prediction label image can be a label image generated by a preset model based on sample image features, which is used to evaluate the prediction performance of the preset model.

[0125] The sample highlight region may be a region obtained after the sample second image features are further enhanced according to feature categories by a multi-layer cascade feature processing space of a preset model. Each feature category channel corresponds to a sample category label, and a sample prediction label image with a sample highlight region is output accordingly.

[0126] The sample annotated image may be an image obtained by annotating the sample to-be-annotated image according to the sample highlighted area in each sample predicted label image.

[0127] The target loss may be a loss function constructed according to the difference between the corresponding regions of the sample annotated image and the sample target label image.

[0128] In some embodiments, the structure of the preset model is the same as the structure of the target model described above, both including a large visual model, a semantic category adapter, and an optimization-oriented expansion algorithm, and the structure of the preset model is not described here. In the preset model, the large visual model is pre-trained and can be directly applied to the preset model. In the process of training the preset model, the parameters of the semantic category adapter and the multi-layer cascade feature processing space are mainly adjusted.

[0129] Please refer to Figure 4 For example, a seed image, that is, a sample target label image, can be obtained first, and a sample image to be labeled without annotations can be generated based on the sample target label image. Furthermore, the sample image to be labeled can be input into a preset model, and potential features of the sample image to be labeled can be extracted through the visual macro model of the preset model, such as the outline of pedestrians, the body lines of cars, and the texture of trees.

[0130] Furthermore, a semantic category adapter can be used to process the extracted sample first image features. The semantic category adapter can map the first image features to a specific feature category channel, thereby obtaining second image features corresponding to different sample feature category labels. Each feature category channel corresponds to a specific sample feature category label, and will output a second image feature corresponding to the feature category channel. For example, if the sample feature category label is car, then the feature category channel will output a second image feature related to the car. The semantic category adapter will simultaneously output multiple second image features corresponding to multiple feature category channels to achieve preliminary calibration of the features corresponding to each sample feature category label in the image to be annotated.

[0131] Furthermore, a multi-layer cascade optimization mechanism can be used to optimize the second image features layer by layer. The multi-layer cascade optimization mechanism includes multiple feature processing layers, each of which has a first mapping parameter to be adjusted for enhancing feature representation. After each feature processing layer processes all sample second image features, a target constraint model can be constructed in the feature processing layer. The target constraint model constructs a layer sub-loss based on the gap between the feature area of ​​each sample feature category label identified by the target constraint model in the current feature processing layer and the feature area calibrated by the sample feature category label in the sample target label image, and determines the sample second mapping parameter of the corresponding feature processing layer based on the layer sub-loss.

[0132] Furthermore, the sample second image features can be feature enhanced layer by layer according to multiple sample second mapping parameters corresponding to multiple feature processing layers to obtain a sample prediction label image corresponding to each sample feature category label output in the sample feature space, and each sample prediction label image contains a sample highlight area corresponding to the corresponding sample feature category label.

[0133] Furthermore, after labeling the sample to-be-labeled image according to all sample highlight areas corresponding to all sample feature category labels, a sample labeled image can be obtained. The labeling method can be manual labeling according to each sample highlight area, or automatic labeling through a model, etc. This application does not limit the specific labeling method.

[0134] Exemplarily, a target loss can be constructed based on the difference between the corresponding areas of the sample annotated image and the sample target label image. The target loss can be a cross entropy loss, an L1 loss, etc. The target loss can be used to determine the correctness of the feature annotation of each sample feature category label in the sample annotated image.

[0135] Furthermore, the parameters of the preset model can be adjusted based on the target loss. Since the large visual model has been pre-trained and has good feature extraction capabilities, there is no need to adjust the parameters of the large visual model. Instead, it is only necessary to adjust the parameters of the semantic category adapter and the parameters of the multi-layer cascade feature processing space.

[0136] In some embodiments, for the architecture of the semantic category adapter, a single-layer "1*1" convolutional layer (conv1x1) combined with an activation function (ReLU) can be used. Conv1x1 contains a small number of learnable parameters, which can promote the fast convergence of the preset model and meet the needs of small-scale training data. Specifically, the number of adjustable parameters in the semantic category adapter It can be expressed as:

[0137] = ;

[0138] Where N is the number of channels output by the large visual model, is the number of feature category channels of the semantic category adapter, 1x1 is the product of kernel width and height; +1 is the bias term.

[0139] Furthermore, OptOU contains K layers (the value of K is determined according to the actual situation), and the mapping parameters of each layer contain two learnable parameters, namely the first scaling parameter and the first learning parameter (which will be further introduced later). Therefore, OptOU involves two trainable parameters in total. Therefore, the total number of learnable parameters of the preset model is The preset model in this application contains a small number of learnable parameters, so only less training data is needed to achieve fast convergence of the preset model, which greatly improves the training efficiency of the model and saves computing resources. For example, a single annotated RGB seed can be used as a sample target label image, and the parameters of the semantic category adapter can be trained by back propagation through the sample to-be-annotated image. and OptOU parameters ,Right now ,

[0140] in, is the first scaling parameter, is the first learning parameter, Target loss.

[0141] In some embodiments, when the sample target label image covers all the sample feature category labels, only one sample target label image can be used to train the preset model to obtain the target model. For example, if there are 10 sample feature category labels, and the sample target label image a covers the features corresponding to these 10 sample feature category labels, then only the sample target label image a needs to be selected to train the preset model to obtain the target model. If the sample target label image does not cover all the sample feature category labels, images corresponding to the remaining other sample feature category labels can be specially selected for further training to obtain the target model. For example, if there are 4 sample feature category labels, namely, people, cars, trees and street lights, and the sample target label image b covers the features corresponding to 3 sample feature category labels, such as people, cars and trees, then the image corresponding to the remaining 1 sample feature category label, that is, the image containing the street light, can be selected to train the preset model to obtain the target model, thereby improving the efficiency of training.

[0142] By training the preset model in the above way, the preset model can learn effective feature representations from images and generate interpretable, high-fidelity semantic annotations without relying on handcrafted prompts. In this way, not only the time and economic cost required to manually annotate large amounts of data are significantly reduced, but also the annotation process is accelerated through optimization mechanisms, improving the accuracy and efficiency of annotations.

[0143] In some embodiments, in order to improve the recognition accuracy of the preset model for different semantic categories in the image by optimizing the feature representation layer by layer, the image features can be processed according to a specific first mapping parameter through a multi-layer cascade of feature processing layers, and the first mapping parameter can be further optimized according to the processing result, so that the preset model can gradually enhance the feature representation at each layer, which not only improves the generalization ability of the preset model, but also ensures the robustness and accuracy of the preset model when processing new data. For example, step 203 may include:

[0144] (203.1) sequentially passing through each feature processing layer of the multi-layer cascade of feature processing space in the preset model, processing the sample second image feature according to the sample first mapping parameter of each feature processing layer, and obtaining a layer sample prediction label image corresponding to each feature processing layer;

[0145] (203.2) in each feature processing layer, determining the sample second mapping parameter of the current feature processing layer based on the difference between the layer sample highlighted area of ​​the layer sample predicted label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image;

[0146] (203.3) By using the sample second mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are represented by feature enhancement layer by layer according to each sample feature category label, and the sample prediction label image corresponding to each sample feature category label is obtained.

[0147] Among them, the layer sample prediction label image can be a plurality of layers of prediction label images corresponding to the plurality of sample feature category labels generated by each feature processing layer using the first mapping parameter to process the sample second image features in the multi-layer cascade feature processing space.

[0148] The layer sample highlighted region may be a specific region in the layer sample predicted label image that is annotated and highlighted according to the sample feature category label. The layer sample highlighted region represents the position and range of the sample feature category label in the sample to be annotated image.

[0149] The sample second mapping parameter may be a parameter adjusted according to the feature recognition effect of the sample first mapping parameter in the multi-layer cascade feature processing space. The sample second mapping parameter is more conducive to the execution of subsequent classification tasks than the sample first mapping parameter.

[0150] For example, if multiple sample second image features corresponding to multiple feature category channels have been obtained, these features can be processed by the first feature processing layer in the multi-layer cascade feature processing space using the sample first mapping parameters to obtain the first layer sample predicted label image. The layer sample predicted label image will show preliminary feature enhancement results, such as the layer sample highlighted area corresponding to pedestrians.

[0151] Furthermore, the layer sub-loss of the current feature processing layer can be determined by the difference between the layer sample highlighting area corresponding to each sample feature category label and the corresponding area of ​​the sample feature category label in the sample target label image, and a target constraint model can be constructed based on the layer sub-loss to obtain the second mapping parameter under the condition of minimizing the layer sub-loss. Exemplarily, the sample feature category labels of the sample image to be labeled are tree, car, and person. After being processed by the semantic class adapter, three sample second image features related to tree, car, and person are obtained, which are sample second image feature a1, sample second image feature b1, and sample second image feature c1 corresponding to tree, car, and person, respectively. Exemplarily, if the feature processing space has three feature processing layers, for the first feature processing layer, the three sample second images can be feature enhanced to obtain image feature a2, image feature b2, and image feature c2 corresponding to tree, car, and person, respectively. Then, the difference between the tree-related layer sample image area in the image feature a2 and the corresponding area of ​​the tree in the sample target label image can be calculated respectively to determine the layer sub-loss of the tree in the first layer feature processing layer, the difference between the car-related layer sample image area in the image feature b2 and the corresponding area of ​​the car in the sample target label image can be calculated to determine the layer sub-loss of the car in the first layer feature processing layer, the difference between the person-related layer sample image area in the image feature c2 and the corresponding area of ​​the person in the sample target label image can be calculated to determine the layer sub-loss of the person in the first layer feature processing layer, and the final layer sub-loss of the first layer feature processing layer is constructed based on the layer sub-loss of the tree, the layer sub-loss of the car, and the layer sub-loss of the person.

[0152] Furthermore, a target constraint model can be constructed through the sub-layer loss of the first feature processing layer to obtain the second mapping parameters of the samples of the first feature processing layer under the condition of minimizing the sub-layer loss.

[0153] Furthermore, the image features a2, b2 and c2 output by the first feature processing layer can be input to the second feature processing layer, and so on, to finally obtain the sample second mapping parameters of each feature processing layer.

[0154] Through multi-layer cascade feature processing space, the preset model can optimize the feature representation layer by layer. Each layer will be optimized according to the results of the previous layer, and the mapping parameters will be adjusted by learning the difference between the sample target label image and the layer sample predicted label image, so as to evaluate and improve the feature classification ability of the preset model according to the adjusted second mapping parameters.

[0155] In some implementations, in order to improve the accuracy and reliability of image annotation, the quality of feature representation can be gradually improved by optimizing the mapping parameters of each feature processing layer in a multi-layer cascade feature processing space, and finally a high-quality predicted label image can be generated. For example, (203.2) may include:

[0156] (203.2.1) In the current feature processing layer, determine the layer sub-loss of each sample feature category label in the current feature processing layer based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image;

[0157] (203.2.2) Construct a target constraint model based on the layer sub-loss of each sample feature category label in the current feature processing layer, the sample first mapping parameter, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the layer sample predicted label image;

[0158] (203.2.3) Solve the target constraint model to obtain the second mapping parameters of the samples in the current feature processing layer.

[0159] The layer sub-loss may be a loss determined based on the difference between the layer sample highlighted area of ​​the layer sample predicted label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image in the current feature processing layer. The layer sub-loss is a measure that quantifies the difference between the predicted label image generated by the current feature processing layer and the true label image.

[0160] Among them, the target constraint model can be a mathematical model constructed in the current feature processing layer based on layer sub-loss and parameters related to feature processing between feature processing layers and within feature processing layers.

[0161] Exemplarily, the layer sub-loss can be obtained by calculating the difference between the layer sample highlighting region of a specific sample feature category label (e.g., pedestrian, vehicle, etc.) in the layer sample prediction label image and the corresponding region of the sample feature category label in the sample target label image. The layer sub-loss can be measured by a loss function (e.g., cross entropy loss, etc.).

[0162] Furthermore, continuing with the above example of pedestrians, the construction of the target constraint model aims to find a second mapping parameter that minimizes the difference between the prediction of the pedestrian area and the correct area of ​​the pedestrian in the actual sample target label image, while adding some constraints (such as regularization terms) to prevent overfitting. By solving the target constraint model, a set of optimized second mapping parameters can be obtained. The second mapping parameters make the prediction of the pedestrian area closer to the true label.

[0163] Exemplarily, the target constraint model constructed may be in the form of:

[0164]

[0165]

[0166]

[0167] in, Represents the input of the current feature processing layer. The input of the current feature processing layer can be the second image feature of the sample (if the current feature processing layer is the first layer of the feature processing space), or it can be the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer, which is specifically determined according to the position of the feature processing layer; Represents the layer sample prediction label image output by the current feature processing layer for the corresponding sample feature category label; Represents the mapping from input to output of the current feature processing layer, that is = ( ); Indicates the corresponding area of ​​the sample feature category label in the sample target label image; is the first scaling parameter of the k-th layer input relative to the (k-1)-th layer output. The sample first mapping parameters include , after the previous layer obtains the layer sample prediction label image, the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer can be obtained by multiplying it with the first scaling parameter; ) represents the layer sub-loss of the current feature processing layer, and the layer sub-loss is determined by the layer sample image area corresponding to each sample feature category label in the feature processing layer and the loss of the corresponding sample feature category label in the current feature processing layer; Represents the second image feature of the sample after being processed by the previous feature processing layer.

[0168] For example, when the sample feature category labels are pedestrians, trees, and cars, after the semantic category adapter, the sample second image feature a1 corresponding to pedestrians, the sample second image feature b1 corresponding to trees, and the sample second image feature c1 corresponding to cars can be obtained. Through the first scaling parameter, the first feature processing layer of the feature processing space can receive the sample second image feature a2, the sample second image feature b2, and the sample second image feature c2 scaled by the previous level. The first feature processing layer is based on Mapping from input to output is performed to obtain sample second image features a3, sample second image features b3, and sample second image features c3. Sample second image features a3, sample second image features b3, and sample second image features c3 respectively include the layer sample highlighting area of ​​pedestrians, the layer sample highlighting area of ​​trees, and the layer sample highlighting area of ​​vehicles.

[0169] Furthermore, through the differences between the highlighted areas of the pedestrian layer samples, the highlighted areas of the tree layer samples, and the highlighted areas of the car layer samples and the sample target label images of pedestrians, trees, and cars (for example, calculated by the cross entropy loss function), the losses of pedestrians, trees, and cars in the first feature processing layer can be calculated respectively, and the corresponding losses of pedestrians, trees, and cars can be added to obtain the sub-layer loss.

[0170] Furthermore, the second mapping parameter that minimizes the layer loss can be solved by constructing a target constraint model. Specifically, the second image feature of the sample processed by the previous feature processing layer can be used to obtain the second mapping parameter that minimizes the layer loss. , the second image feature of the sample received by the current feature processing layer after being processed by the previous feature processing layer (indicating that it has been processed by the first scaling parameter between layers) and the first scaling parameter , construct the first constraint function, that is, the first constraint function is .

[0171] Based on the second image feature of the sample received by the current feature processing layer after processing by the previous feature processing layer (indicating that it has been processed by the first scaling parameter between layers) , layer sample prediction label image and , construct the second constraint function, that is, the second constraint function is .

[0172] By constructing the first constraint function and the second constraint function, it is possible to ensure that the mapping parameters remain within a reasonable range during the optimization process, thereby preventing unstable or unreasonable results during the optimization process.

[0173] It should be noted that the processing of the subsequent feature processing layers was the same as that of the first feature processing layer and the steps of constructing the target sleepiness model, which will not be repeated here.

[0174] Through the above process, the preset model can gradually adjust the mapping parameters of each feature processing level and finally generate a higher quality predicted label image.

[0175] In some implementations, in order to ensure that the mapping parameters remain within a reasonable range during the optimization process, the solved second mapping parameters may be constrained to improve the stability of the optimization process and avoid numerical instability. For example, (203.2.2) may include:

[0176] (203.2.2.1) Construct an objective function based on multiple layer sub-losses corresponding to multiple sample feature category labels in the current feature processing layer;

[0177] (203.2.2.2) For each sample feature category label, obtain the sample second image feature processed by the previous feature processing layer of the current feature processing layer, and the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and construct a first constraint function based on the sample second image feature processed by the previous feature processing layer, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the first scaling parameter;

[0178] (203.2.2.3) Construct a second constraint function based on the second image features of the samples processed by the previous feature processing layer, the layer sample prediction label image and the inter-layer mapping parameters received by the current feature processing layer;

[0179] (203.2.2.4) Based on the objective function, the first constraint function and the second constraint function, a target constraint model is constructed.

[0180] Among them, the objective function can be used to measure the difference between the layer sample highlighted area corresponding to the layer sample prediction label image and the corresponding area of ​​the corresponding layer sample prediction label, and serve as the minimization target in the optimization process.

[0181] The second image feature of the sample processed by the previous feature processing layer of the current feature processing layer may be the second image feature of the sample processed by the previous feature processing layer.

[0182] Among them, the second image feature of the sample processed by the previous feature processing layer and received by the current feature processing layer may represent the second image feature of the sample processed by the previous feature processing layer and received by the current feature processing layer, which is obtained by scaling the second image feature of the sample processed by the previous feature processing layer through the first scaling parameter.

[0183] The first scaling parameter may be the first scaling parameter of the input of the feature processing layer of the current feature processing layer (the kth layer) relative to the output of the previous feature processing layer (the k−1th layer), which is used to adjust the importance of contributions between different layers.

[0184] The first constraint function may be to ensure that the input of the current feature processing layer is the output of the previous feature processing layer adjusted by the first scaling parameter.

[0185] Among them, the inter-layer mapping parameters can be used to characterize the mapping relationship from input to output of the current feature processing layer.

[0186] The second constraint function may be to ensure that the output of the current feature processing layer is obtained by transforming the input through the inter-layer mapping parameters.

[0187] Please refer to Figure 4 and Figure 5,Exemplary, the form of the target constraint model can be as follows:

[0188]

[0189]

[0190]

[0191] in, Represents the input of the current feature processing layer. The input of the current feature processing layer can be the second image feature of the sample (if the current feature processing layer is the first layer of the feature processing space), or it can be the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer. Represents the layer sample prediction label image output by the current feature processing layer for the corresponding sample feature category label; Represents the mapping from input to output of the current feature processing layer (inter-layer mapping parameter), that is = ( ); Indicates the corresponding area of ​​the sample feature category label in the sample target label image; is the first scaling parameter of the k-th layer input relative to the (k-1)-th layer output. The sample first mapping parameters include , after the previous layer obtains the layer sample prediction label image, the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer can be obtained by multiplying it with the first scaling parameter; ) represents the layer sub-loss of the current feature processing layer, and the layer sub-loss is determined by the layer sample image area corresponding to each sample feature category label in the feature processing layer and the loss of the corresponding sample feature category label in the current feature processing layer; Represents the second image feature of the sample after being processed by the previous feature processing layer.

[0192] In some implementations, each feature processing layer constructs a target constraint model to optimize the mapping parameters of each layer to obtain updated sample second mapping parameters of each layer.

[0193] By defining an optimization problem for each layer, the preset model can be allowed to focus on solving specific optimization goals in each feature processing layer. In this way, the preset model can gradually learn more complex features, and each layer can focus on improving its own output, thus forming a hierarchical feature learning process in the entire network to speed up the convergence of the model and improve the accuracy of feature processing of the entire model.

[0194] In some implementations, in order to apply the sample second mapping parameters to the optimization of the features by the preset model and verify the optimization capability of the model for the features, the sample second image features may be converted into the output of the multi-layer cascade feature processing space by the sample second mapping parameters to improve the overall performance of the model and the accuracy of prediction. For example, (203.3) may include:

[0195] (203.3.1) Generate a target mapping function of a preset model by using multiple sample second mapping parameters corresponding to multiple feature processing layers of the preset model; wherein the target mapping function is used to perform a layer-by-layer feature enhancement representation of the corresponding sample second image features according to each sample feature category label;

[0196] (203.3.2) Calculate the sample second image features corresponding to each sample category label according to the target mapping function to obtain the sample predicted label image corresponding to each sample feature category label.

[0197] Among them, the target mapping function can be a function used to process the sample second image features by integrating the hierarchical information of all feature processing layers in the feature processing space. The target mapping function processes the sample second image features to generate the output of the entire preset model.

[0198] Exemplarily, global adjustments can be made through the target mapping function. Specifically, the importance of global features can be adjusted by calculating the product of the first scaling parameters of all feature processing layers (first product) to ensure the consistency of features throughout the entire processing process. Further, local refinement can be performed through the target mapping function by calculating the scaling parameter product of all layers after the current feature processing layer (second product), focusing on the specific contribution of each feature processing layer to the final output, and achieving fine adjustment of local features. Furthermore, the difference (first difference) between the probability distribution of the corresponding sample category label of the current layer in the current feature processing layer and the corresponding area in the sample target label image can be timely adjusted to adjust the parameters of the preset model so that the preset model's prediction of the image is closer to the true label.

[0199] In this way, the preset model can gradually optimize the mapping parameters of each feature processing layer, and ultimately generate higher quality predicted label images, thereby improving the prediction accuracy and robustness of the model.

[0200] In some embodiments, in order to obtain more accurate prediction results, the feature representation can be optimized layer by layer through the product of the scaling factor and the mapping subfunction. Specifically, the first product and the first mapping subfunction ensure the consistency adjustment of the global features, while the second product, the probability distribution, the first difference and the second mapping subfunction allow local optimization. Finally, the global and local information are integrated through the target mapping function to obtain accurate prediction results. In this way, the explanatory power and prediction accuracy of the preset model are enhanced, while ensuring the reasonable transfer and optimization of features between layers. For example, (203.3.1) may include:

[0201] (203.3.1.1) Obtain a first product based on the product of multiple first scaling parameters corresponding to multiple feature processing layers;

[0202] (203.3.1.2) obtaining a first mapping subfunction based on the first product and the product of the sample second image feature corresponding to the sample category label;

[0203] (203.3.1.3) For each current feature processing layer, obtain the product of the first scaling parameter subsequent to the current feature processing layer in the feature processing space to obtain a second product;

[0204] (203.3.1.4) Determine the probability distribution of the corresponding sample category label at the current feature processing layer based on the second image feature of the sample processed by the previous feature processing layer of the current feature processing layer and the first scaling parameter;

[0205] (203.3.1.5) Obtain a first difference value based on the probability distribution corresponding to the sample category label and the difference value of the corresponding area in the sample target label image;

[0206] (203.3.1.6) Determine a second mapping subfunction based on the second product, the first learning parameter, and the first difference;

[0207] (203.3.1.7) Obtain a target mapping function based on the sum of multiple second mapping sub-functions corresponding to multiple feature processing layers in the feature processing space and the first mapping sub-function.

[0208] The first product can be the product of the first scaling parameters of all feature processing layers. For example, if the feature processing space has three feature processing layers, then the first product can be the product of the first feature processing layer corresponding to , the second feature processing layer corresponds to , the third feature processing layer corresponds to The product of .

[0209] The first mapping subfunction may be the product of the second image feature of the sample and the first product.

[0210] The second product may be the product of the first scaling parameters of all subsequent feature processing layers of the current feature processing layer. For example, there are three feature processing layers in the feature processing space, the current feature processing layer is the first feature processing layer, and the first scaling parameter corresponding to the first feature processing layer is , the first scaling parameter corresponding to the second feature processing layer is , the first scaling parameter corresponding to the third feature processing layer is , then the second product of the first feature processing layer is and The product of .

[0211] The probability distribution can be the first scaling parameter of the current feature processing layer k after the feature processing layer k-1 of the previous feature processing layer k processes the second image feature of the sample. After scaling, the sample is input into the probability distribution function to obtain the probability distribution of the sample category label of the corresponding sample second image feature at the current feature processing layer.

[0212] The first difference may be a difference between the probability distribution of the current feature processing layer and the corresponding area in the sample target label image.

[0213] The second mapping sub-function may be a function constructed by multiplying the second product, the first learning parameter, and the first difference.

[0214] In some implementations, the target mapping function may be expressed as follows:

[0215] ;

[0216] in, Represents the first product, which is used to scale the second image feature of the sample and output it. represents the first mapping subfunction, represents the second product, ) represents the probability distribution of the corresponding sample category label in the current feature processing layer; represents the first difference; Represents the second mapping subfunction.

[0217] in, The details are as follows:

[0218] ;

[0219] Below, we will give an example based on the above formula:

[0220] Assume that the feature processing space corresponding to the OptOU of the preset model includes three feature processing layers, each with its own first scaling parameter , , , then the first product is , , In this way, the importance of each layer output can be adjusted between different feature processing layers of the preset model for subsequent feature enhancement.

[0221] Furthermore, assuming that the second image feature of the sample is , then the first mapping subfunction is In this way, the sample image features can be enhanced in different feature processing layers.

[0222] Exemplarily, for the first feature processing layer, the product of the subsequent first scaling parameter is ,Right now × Similarly, the second and third layers are calculated in this way.

[0223] Furthermore, after scaling the output of the previous feature processing layer by the first scaling parameter, the input of the current feature processing layer can be obtained, and the input can be converted into a probability distribution by a function to obtain For example, for the second feature processing layer, the corresponding probability distribution is ,in, The formula can be Calculated.

[0224] Furthermore, the first difference may be calculated based on the probability distribution of the sample second image feature received by the current feature processing layer under the corresponding sample category label and the probability distribution of the sample category label in the sample target label image.

[0225] Furthermore, according to the second product , the first learning parameter and the first difference , and obtain the second mapping sub-function.

[0226] Finally, the target mapping function can be obtained by subtracting the sum of multiple second mapping sub-functions corresponding to multiple feature processing layers in the feature processing space from the first mapping sub-function.

[0227] Through the above methods, the feature representation can be effectively enhanced and consistency can be maintained in multi-layer feature processing, thereby enhancing the model's explanatory power and prediction accuracy, while ensuring the reasonable transfer and optimization of features between layers, thereby significantly improving the overall performance and robustness of the model.

[0228] Please refer to Figure 6 , the convergence and effect of the proposed default model (LAM) in this application are as follows Figure 6 As shown by Figure 6 It can be seen that compared with other models, LAM achieves nearly 100% accuracy and converges extremely quickly (promoted by its extremely small size of learnable parameters).

[0229] See also Figure 7 The embodiment of the present application further provides an image annotation device based on an autonomous driving scenario, which can implement the above-mentioned image annotation method based on an autonomous driving scenario. The image annotation device based on an autonomous driving scenario includes:

[0230] An acquisition module 71 is used to acquire an image to be annotated from the autonomous driving data, and input the image to be annotated into a target model to extract potential features, thereby obtaining a first image feature;

[0231] An adaptation module 72, configured to adapt the first image feature through a plurality of feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label;

[0232] Enhancement module 73, used to perform feature enhancement representation of the corresponding second image features layer by layer according to each feature category label through mapping parameters corresponding to the multi-layer cascade feature processing space in the target model, to obtain a target label image corresponding to each feature category label, each target label image includes a target highlight area corresponding to the corresponding feature category label; wherein the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by minimizing the difference between the sample target label image and at least one sample predicted label image, and each sample predicted label image is obtained by feature enhancement representation of the corresponding sample image features by the feature processing space based on the corresponding sample feature type label;

[0233] The labeling module 74 is used to label the image to be labeled according to the target highlighted area in each target label image.

[0234] The specific implementation of the image annotation device based on the autonomous driving scene is basically the same as the specific implementation of the image annotation method based on the autonomous driving scene, and will not be repeated here. On the premise of meeting the requirements of the embodiment of this application, the image annotation device based on the autonomous driving scene can also be provided with other functional modules to implement the image annotation method based on the autonomous driving scene in the above embodiment.

[0235] The embodiment of the present application also provides a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned image annotation method based on the autonomous driving scene when executing the computer program. The computer device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0236] See also Figure 8 , Figure 8 The hardware structure of a computer device according to another embodiment is shown, and the computer device includes:

[0237] The processor 81 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0238] The memory 82 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 82 can store an operating system and other applications. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 82, and the processor 81 calls and executes the image annotation method based on the autonomous driving scene in the embodiment of this application;

[0239] Input / output interface 83, used to implement information input and output;

[0240] Communication interface 84, used to realize communication interaction between the device and other devices, which can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0241] A bus 85 that transmits information between the various components of the device (e.g., the processor 81, the memory 82, the input / output interface 83, and the communication interface 84);

[0242] The processor 81 , the memory 82 , the input / output interface 83 and the communication interface 84 are connected to each other in communication within the device via a bus 85 .

[0243] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned image labeling method based on the autonomous driving scenario is implemented.

[0244] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0245] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0246] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0247] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0248] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0249] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0250] It should be understood that in the present application, "at least one (item)" and "several" refer to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0251] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0252] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0253] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0254] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0255] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. An image annotation method based on an autonomous driving scenario, characterized in that: The method comprises: Acquire an image to be annotated from the autonomous driving data, and input the image to be annotated into a target model to extract potential features, thereby obtaining a first image feature; Adapting the first image feature through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label; By using mapping parameters corresponding to the multi-layer cascade feature processing space in the target model, the corresponding second image features are represented by feature enhancement layer by layer according to each feature category label, so as to obtain a target label image corresponding to each feature category label, wherein each target label image includes a target highlighting area corresponding to the corresponding feature category label; The feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by establishing a constraint model for minimization learning based on the difference between the sample target label image and at least one sample predicted label image, so as to optimize the mapping parameters so that the feature processing layer can identify a label image closer to the sample type label according to the adjusted mapping parameters; each sample predicted label image is obtained by the feature processing space performing feature enhancement representation on the corresponding sample image features based on the corresponding sample feature type label; The target model is trained in the following manner: obtaining a sample image to be annotated corresponding to the sample target label image, and inputting the sample image to be annotated into a preset model to extract potential features, thereby obtaining multiple sample first image features; adapting the multiple sample first image features through multiple feature category channels in the preset model to obtain sample second image features corresponding to each feature category channel; each feature category channel corresponds to a sample feature category label of the sample target label image; performing layer-by-layer feature enhancement representation of the corresponding sample second image features according to each sample feature category label through the sample first mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, thereby obtaining a sample prediction label image corresponding to each sample feature category label, wherein each sample prediction label image contains a sample highlight area corresponding to the corresponding sample feature category label; annotating the sample image to be annotated according to the sample highlight area in each sample prediction label image to obtain a sample annotated image; constructing a target loss according to the difference between the sample annotated image and the corresponding area of ​​the sample target label image; adjusting the parameters of the preset model based on the target loss to obtain a target model; The image to be labeled is labeled according to the target highlighted area in each target label image.

2. The image annotation method based on the autonomous driving scene according to claim 1, characterized in that: The sample first mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model are used to perform feature enhancement representation of the corresponding sample second image features layer by layer according to each sample feature category label to obtain a sample prediction label image corresponding to each sample feature category label, including: Sequentially through each feature processing layer of the multi-layer cascade of the feature processing space in the preset model, the sample second image feature is processed according to the sample first mapping parameter of each feature processing layer, to obtain a layer sample prediction label image corresponding to each feature processing layer; In each feature processing layer, based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image, determine the sample second mapping parameter of the current feature processing layer; By using the sample second mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are feature enhanced layer by layer according to each sample feature category label to obtain a sample prediction label image corresponding to each sample feature category label.

3. The image annotation method based on the autonomous driving scene according to claim 2, characterized in that: The method of determining the sample second mapping parameter of the current feature processing layer in each feature processing layer based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image comprises: In the current feature processing layer, based on the difference between the layer sample highlighted area of ​​the layer sample prediction label image corresponding to each sample feature category label and the corresponding area of ​​the corresponding sample feature category label in the sample target label image, determine the layer sub-loss of each sample feature category label in the current feature processing layer; Constructing a target constraint model based on the layer sub-loss of each sample feature category label at the current feature processing layer, the sample first mapping parameter, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the layer sample prediction label image; The target constraint model is solved to obtain the second mapping parameters of the samples of the current feature processing layer.

4. The image annotation method based on the autonomous driving scene according to claim 3 is characterized in that: The sample first mapping parameter includes a first scaling parameter and an inter-layer mapping parameter; the target constraint model is constructed based on the layer sub-loss of each sample feature category label at the current feature processing layer, the sample first mapping parameter, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the layer sample prediction label image, including: Based on multiple sample feature category labels, construct an objective function at multiple layer sub-losses corresponding to the current feature processing layer; For each sample feature category label, obtain the sample second image feature processed by the previous feature processing layer of the current feature processing layer, and the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and construct a first constraint function based on the sample second image feature processed by the previous feature processing layer, the sample second image feature processed by the previous feature processing layer received by the current feature processing layer, and the first scaling parameter; Constructing a second constraint function based on the second image feature of the sample processed by the previous feature processing layer received by the current feature processing layer, the layer sample prediction label image and the inter-layer mapping parameter; A target constraint model is constructed based on the target function, the first constraint function and the second constraint function.

5. The image annotation method based on the autonomous driving scene according to claim 2, characterized in that: The sample second mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model are used to perform feature enhancement representation of the corresponding sample second image features layer by layer according to each sample feature category label to obtain a sample prediction label image corresponding to each sample feature category label, including: Generate a target mapping function of the preset model through multiple sample second mapping parameters corresponding to multiple feature processing layers of the preset model; wherein the target mapping function is used to perform a layer-by-layer feature enhancement representation of the corresponding sample second image features according to each sample feature category label; The sample second image feature corresponding to each sample category label is calculated according to the target mapping function to obtain a sample predicted label image corresponding to each sample feature category label.

6. The image annotation method based on the autonomous driving scenario according to claim 5, characterized in that: The sample first mapping parameter includes a first scaling parameter and a first learning parameter; the target mapping function of the preset model is generated by using the plurality of sample second mapping parameters corresponding to the plurality of feature processing layers of the preset model, including: Obtaining a first product according to the product of a plurality of first scaling parameters corresponding to the plurality of feature processing layers; Obtaining a first mapping subfunction based on the product of the first product and the sample second image feature of the corresponding sample category label; For each current feature processing layer, obtaining a product of a first scaling parameter subsequent to the current feature processing layer in the feature processing space to obtain a second product; Determine the probability distribution of the corresponding sample category label at the current feature processing layer according to the second image feature of the sample processed by the previous feature processing layer of the current feature processing layer and the first scaling parameter; Obtaining a first difference based on the probability distribution corresponding to the sample category label and the difference between the corresponding area in the sample target label image; determining a second mapping sub-function according to the second product, the first learning parameter and the first difference; A target mapping function is obtained according to the sum of a plurality of second mapping sub-functions corresponding to a plurality of feature processing layers in the feature processing space and the first mapping sub-function.

7. An image annotation device based on an autonomous driving scenario, characterized in that: The device comprises: An acquisition module, used for acquiring an image to be annotated from the autonomous driving data, and inputting the image to be annotated into a target model to extract potential features, thereby obtaining a first image feature; An adaptation module, used for adapting the first image feature through multiple feature category channels in the target model to obtain a second image feature corresponding to each feature category channel; each feature category channel corresponds to a feature category label; An enhancement module is used to perform a layer-by-layer feature enhancement representation of the corresponding second image features according to each feature category label through the mapping parameters corresponding to the multi-layer cascade feature processing space in the target model, so as to obtain a target label image corresponding to each feature category label, and each target label image contains a target highlight area corresponding to the corresponding feature category label; wherein, the feature processing space is composed of a plurality of cascaded feature processing layers, and the mapping parameters corresponding to each feature processing layer are obtained by establishing a constraint model for minimization learning based on the difference between the sample target label image and at least one sample predicted label image, so as to optimize the mapping parameters so that the feature processing layer can identify a label image closer to the sample type label according to the adjusted mapping parameters; each sample predicted label image is obtained by the feature processing space performing feature enhancement representation on the corresponding sample image features based on the corresponding sample feature type label; wherein, the target model is trained in the following manner: obtaining a sample image to be labeled corresponding to the sample target label image, and The image is input into a preset model to extract potential features, and multiple sample first image features are obtained; the multiple sample first image features are adapted through multiple feature category channels in the preset model to obtain sample second image features corresponding to each feature category channel; each feature category channel corresponds to a sample feature category label of the sample target label image; through the sample first mapping parameters corresponding to the multi-layer cascade feature processing space in the preset model, the corresponding sample second image features are represented by feature enhancement layer by layer according to each sample feature category label, and the sample prediction label image corresponding to each sample feature category label is obtained, and each sample prediction label image contains a sample highlight area corresponding to the corresponding sample feature category label; the sample to-be-annotated image is annotated according to the sample highlight area in each sample prediction label image to obtain a sample annotated image; according to the difference between the sample annotated image and the corresponding area of ​​the sample target label image, a target loss is constructed; based on the target loss, the parameters of the preset model are adjusted to obtain a target model; The labeling module is used to label the to-be-labeled image according to the target highlighted area in each target label image.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the image annotation method based on the autonomous driving scene according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image annotation method based on the autonomous driving scenario according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-task joint perception network model and detection method for traffic road surface information

    US20240420487A1