Model training method and device, equipment, storage medium and program product

By introducing pavement feature topological relationships as constraints in the pavement feature segmentation model, and adjusting model parameters to match predefined topological relationships, the problem of low pavement feature segmentation accuracy in complex scenarios is solved, and higher segmentation accuracy and spatial relationship understanding are achieved.

CN120014580APending Publication Date: 2025-05-16NAVINFO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111959.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and segment different types of road surface elements in complex scenarios such as road environments, resulting in low segmentation accuracy of segmentation models.

Method used

By introducing pavement feature topological relationships as constraints, the model parameters are adjusted to match predefined topological relationships, and the accuracy of the model is improved at pixel level and spatial relationships.

Benefits of technology

The segmentation accuracy of the pavement feature segmentation model is significantly improved, so that the model is not only more accurate at the pixel level, but also follows spatial relationships and more accurately understand and segment the pavement features in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014580A_ABST
    Figure CN120014580A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, equipment, a storage medium and a program product, and particularly relates to the technical field of traffic information processing, and the method comprises the steps: inputting a sample image containing road surface elements into a to-be-trained model, carrying out the segmentation of the road surface elements, and obtaining a label mapping graph of the sample image, pixel values of pixels in the label mapping graph represent road surface element categories to which the corresponding pixels in the sample image belong; based on a predefined road surface element topological relation, target pixels are identified from the label mapping graph, and the matching degree of the target pixels and the road surface element topological relation meets a preset matching degree condition; and adjusting model parameters of the to-be-trained model based on the target pixel. The method is used for achieving the effect of improving the segmentation accuracy of the pavement element segmentation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of traffic information processing, and in particular to a model training method, device, equipment, storage medium and program product. Background Art

[0002] Pavement elements usually refer to the components of the road surface, including various traffic signs and markings. Traffic signs and markings refer to the markings on the road surface that use lines, arrows, text, elevation marks, raised road signs and contour marks to convey traffic information such as guidance, restrictions and warnings to traffic participants.

[0003] Pavement feature segmentation refers to the process of identifying and distinguishing different types of ground features or features in image or point cloud data. This usually involves computer vision and image processing techniques, especially semantic segmentation techniques, which are used to classify pixels in images or points in point clouds into different categories. Pavement feature segmentation has important applications in fields such as autonomous driving, robot navigation, urban planning, and geographic information systems (GIS).

[0004] The features of road surface elements may not be obvious due to various factors (such as lighting, background, shooting angle, wear, etc.), making it difficult for the segmentation model to accurately identify and segment different road surface elements. Summary of the invention

[0005] The embodiments of the present application provide a model training method, device, equipment, storage medium and program product to achieve the effect of improving the segmentation accuracy of the road feature segmentation model.

[0006] In a first aspect, an embodiment of the present application provides a model training method, comprising:

[0007] Inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs;

[0008] Based on the predefined road surface element topological relationship, identifying a target pixel from the label map, the target pixel being a pixel whose matching degree with the road surface element topological relationship satisfies a preset matching degree condition;

[0009] Based on the target pixel, a model parameter of the model to be trained is adjusted.

[0010] In a possible implementation, adjusting the model parameters of the to-be-trained model based on the target pixel includes:

[0011] Marking the target pixel and generating a topological relationship matching image, wherein the pixel value of the pixel in the topological relationship matching image represents the matching degree between the corresponding pixel in the label mapping image and the topological relationship of the road surface element;

[0012] Calculating model loss based on the topological relationship matching image;

[0013] Based on the model loss, the model parameters of the model to be trained are adjusted.

[0014] In a possible implementation manner, the road surface element topological relationship includes a mutually exclusive relationship, and the identifying the target pixel from the label map based on the predefined road surface element topological relationship includes:

[0015] Performing convolution processing on the label map to obtain a convolution image;

[0016] For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the convolution image, determine its matching degree with the mutually exclusive relationship;

[0017] A pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is determined as the identified first target pixel.

[0018] In a possible implementation, performing convolution processing on the label map to obtain a convolution image includes:

[0019] The label map is convolved to calculate the difference between a pixel in the label map and its neighboring pixels to obtain a convolved image.

[0020] In a possible implementation manner, the road surface element topological relationship includes an inclusion relationship, and the identifying the target pixel from the label map based on the predefined road surface element topological relationship includes:

[0021] Obtaining a true value label mapping diagram of the sample image;

[0022] For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the true value label map, determine its matching degree with the inclusion relationship;

[0023] A pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is determined as the identified second target pixel.

[0024] In a possible implementation, marking the target pixel and generating a topological relationship matching image includes:

[0025] Marking the position of the first target pixel to generate a first marked image;

[0026] marking the position of the second target pixel to generate a second marked image;

[0027] A union operation is performed on the first labeled image and the second labeled image to generate the topological relationship matching degree image.

[0028] In a possible implementation manner, the step of inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements and obtaining a label map of the sample image includes:

[0029] Inputting a sample image containing road surface elements into the model to be trained to extract features, and obtaining a feature map of the sample image;

[0030] Perform pixel-level category label mapping on the feature map to generate the label mapping map, where different category labels correspond to different road surface element categories.

[0031] In a possible implementation, the calculating the model loss based on the topological relationship matching image includes:

[0032] Using a cross entropy loss function to calculate the difference between the feature map of the sample image and the true value label mapping map of the sample image, to obtain a cross entropy loss;

[0033] Based on the topological relationship matching image, the cross entropy loss is weighted to calculate the model loss.

[0034] In a second aspect, an embodiment of the present application provides a model training device, comprising:

[0035] A segmentation module, used for inputting a sample image containing road surface elements into a model to be trained to perform road surface element segmentation, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs;

[0036] An identification module, for identifying target pixels from the label map based on a predefined road surface element topological relationship, wherein the target pixels are pixels whose matching degree with the road surface element topological relationship satisfies a preset matching degree condition;

[0037] An adjustment module is used to adjust the model parameters of the model to be trained based on the target pixel.

[0038] In a possible implementation manner, the adjustment module is specifically used to:

[0039] Marking the target pixel and generating a topological relationship matching image, wherein the pixel value of the pixel in the topological relationship matching image represents the matching degree between the corresponding pixel in the label mapping image and the topological relationship of the road surface element;

[0040] Calculating model loss based on the topological relationship matching image;

[0041] Based on the model loss, the model parameters of the model to be trained are adjusted.

[0042] In a possible implementation manner, the road surface element topological relationship includes a mutually exclusive relationship, and the identification module is specifically used to:

[0043] Performing convolution processing on the label map to obtain a convolution image;

[0044] For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the convolution image, determine its matching degree with the mutually exclusive relationship;

[0045] A pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is determined as the identified first target pixel.

[0046] In a possible implementation manner, the identification module is specifically used to:

[0047] The label map is convolved to calculate the difference between a pixel in the label map and its neighboring pixels to obtain a convolved image.

[0048] In a possible implementation manner, the predefined road surface element topological relationship includes an inclusion relationship, and the identification module is specifically configured to:

[0049] Obtaining a true value label mapping diagram of the sample image;

[0050] For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the true value label map, determine its matching degree with the inclusion relationship;

[0051] A pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is determined as the identified second target pixel.

[0052] In a possible implementation manner, the adjustment module is specifically used to:

[0053] Marking the position of the first target pixel to generate a first marked image;

[0054] marking the position of the second target pixel to generate a second marked image;

[0055] A union operation is performed on the first labeled image and the second labeled image to generate the topological relationship matching degree image.

[0056] In a possible implementation manner, the segmentation module is specifically used to:

[0057] Inputting a sample image containing road surface elements into the model to be trained to extract features, and obtaining a feature map of the sample image;

[0058] Perform pixel-level category label mapping on the feature map to generate the label mapping map, where different category labels correspond to different road surface element categories.

[0059] In a possible implementation manner, the adjustment module is specifically used to:

[0060] Using a cross entropy loss function to calculate the difference between the feature map of the sample image and the true value label mapping map of the sample image, to obtain a cross entropy loss;

[0061] Based on the topological relationship matching image, the cross entropy loss is weighted to calculate the model loss.

[0062] In a third aspect, an embodiment of the present application provides a model training device, including: a memory, a processor;

[0063] The memory stores computer-executable instructions;

[0064] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0065] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0066] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0067] The model training method, device, equipment, storage medium and program product provided by the embodiments of the present application can input a sample image containing road surface elements into the model to be trained to perform road surface element segmentation, and obtain a label map of the sample image. The pixel value of the pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs. The target pixel whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition can be extracted from the label map, and the model parameters of the model to be trained are adjusted based on the extracted target pixel. In the image segmentation task, especially when it involves the segmentation of complex scenes such as road environments, relying solely on pixel-level classification may not be enough to capture higher-level spatial relationships and structural information. Therefore, the present application introduces the topological relationship of road surface elements as a constraint, so that the segmentation result of the model is not only more accurate at the pixel level, but also follows the spatial relationship. This method can more accurately understand and segment the road surface elements in the image. By training the model using the above method, the segmentation accuracy of the road surface element segmentation model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0069] Figure 1 A schematic diagram of a scenario for the model training method provided in this application;

[0070] Figure 2 Schematic diagram of the model training method provided for this application Figure 1 ;

[0071] Figure 3 Schematic diagram of the model training method provided for this application Figure 2 ;

[0072] Figure 4 A schematic diagram of a label mapping diagram provided for this application;

[0073] Figure 5 A schematic diagram of the convolution kernel provided for this application;

[0074] Figure 6 A schematic diagram of a convolution image generated by convolution processing of a label map provided in the present application;

[0075] Figure 7 A schematic diagram of a first labeled image provided in this application;

[0076] Figure 8 A schematic diagram of a second labeled image provided in this application;

[0077] Fig. 9A schematic diagram of a topological relationship matching degree image generated by the union provided in this application;

[0078] Fig.10 A schematic diagram of the training process of the model training method provided in this application;

[0079] Fig.11 A schematic diagram of the calculation process of the model loss provided for this application;

[0080] Fig.12 A schematic diagram of the structure of the model training device provided in this application;

[0081] Fig.13 A schematic diagram of the structure of the model training device provided in this application.

[0082] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0083] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0084] In the vehicle-road cooperative system, the construction and update of the dynamic map of autonomous driving depends on the accurate semantic segmentation of road features. The segmentation of road features is affected by many factors, such as lighting, background, shooting angle, and wear, which may make the features of road features unclear, making it difficult for the segmentation model to accurately identify and segment different road features.

[0085] The model training method provided in the present application can input a sample image containing road surface elements into a model to be trained to segment the road surface elements, and obtain a label map of the sample image. The pixel value of the pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs. The target pixels whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition can be extracted from the label map, and the model parameters of the model to be trained are adjusted based on the extracted target pixels. In image segmentation tasks, especially when it involves the segmentation of complex scenes such as road environments, relying solely on pixel-level classification may not be sufficient to capture higher-level spatial relationships and structural information. Therefore, the present application introduces the topological relationship of road surface elements as a constraint, so that the segmentation result of the model is not only more accurate at the pixel level, but also follows the spatial relationship. This method can more accurately understand and segment the road surface elements in the image. By training the model using the above method, the technical problem of low segmentation accuracy of the road surface element segmentation model is solved.

[0086] Figure 1 A schematic diagram of the scenario of the model training method provided in this application, such as Figure 1 As shown, the training data set includes multiple training samples, and the training samples can specifically be sample images containing road surface elements. In the terminal device 101, the sample images containing road surface elements can be input into the road surface element segmentation model for model training. Among them, the terminal device can be a computing unit in an intelligent transportation system, such as a desktop computer, a laptop computer, and a vehicle. The trained road surface element segmentation model can be deployed in a vehicle driving recorder for road surface element recognition.

[0087] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0088] Figure 2 Schematic diagram of the model training method provided for this application Figure 1 ,like Figure 2 As shown, the method includes:

[0089] S201, inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs.

[0090] In an embodiment of the present application, the model to be trained may be a road surface element segmentation model. By inputting a sample image containing road surface elements into the road surface element segmentation model, a label map may be generated, and each pixel value in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs. Road surface elements may be divided into small categories such as stop signs, speed limit signs, roadway boundaries, pedestrian crossings, left turn guide lines, stop lines, etc. based on specific functions and uses. In addition, road surface elements may also be divided into large categories such as lines, arrows, text, symbols, and road surface areas based on shape and representation.

[0091] Among them, line-type pavement elements (which can be referred to as line elements) can include lane boundary lines, pedestrian crossing lines, left turn guide lines, stop lines, etc. These elements usually appear in the form of continuous lines. Arrow-type pavement elements (which can be referred to as arrow elements) are used to indicate the direction of travel and usually appear in the middle of the lane. Text-type pavement elements (which can be referred to as text elements) include speed limit signs, etc., which are traffic information in text form. Symbol-type pavement elements (which can be referred to as symbol elements) are various graphic symbols, such as the "P" symbol in the stop sign. The pavement area refers to the entire area of ​​the road surface.

[0092] S202: Based on the predefined road surface element topological relationship, identify target pixels from the label mapping image, where the target pixels are pixels whose matching degree with the road surface element topological relationship meets a preset matching degree condition.

[0093] In an embodiment of the present application, pixels whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition can be extracted from the label mapping map. Among them, the predefined road surface element topological relationship can include mutually exclusive relationships and inclusive relationships. Mutually exclusive relationships refer to that different types of road surface elements have no intersection and are not adjacent, for example, line-type road surface elements and arrow-type road surface elements are not adjacent, arrow-type road surface elements and text-type road surface elements are not adjacent, symbol-type road surface elements and text-type road surface elements are not adjacent, etc. Inclusion relationships refer to that multiple types of road surface elements are all within the road surface area, and line-type road surface elements, arrow-type road surface elements, text-type road surface elements and symbol-type road surface elements are all within the road surface area.

[0094] The degree of match refers to the similarity or consistency of a pixel value or a group of pixel values ​​with a predefined pattern or structure. By analyzing the similarity or consistency of the pixel value distribution of the label map and the predefined road feature topology, the target pixel can be extracted. When the preset match degree condition is set to mismatch, the target pixel is the pixel that does not conform to the predefined road feature topology. This setting is used to identify abnormal or irregular areas. When the preset match degree condition is set to match, the target pixel is the pixel that conforms to the predefined road feature topology. This setting is used to confirm and identify normal road features.

[0095] S203: Adjust the model parameters of the model to be trained based on the target pixel.

[0096] In one embodiment, the model loss can be calculated based on the target pixel and the model parameters can be adjusted to achieve a more targeted optimization of the model so that it performs better on a specific task. This method helps to improve the accuracy and robustness of the model. In another embodiment, the model parameters can be optimized by introducing an adaptive loss function. The hyperparameters in the adaptive loss function can be adjusted based on the target pixel, and these parameters can be dynamically updated as the training progresses. The calculated hyperparameters are used to reconstruct the adaptive loss function, and the model parameters can be adjusted through the back propagation algorithm based on the gradient information of the reconstructed adaptive loss function. Since the adaptive loss function is optimized according to the target pixel, the update of the model parameters will be more accurate and efficient.

[0097] The model training method provided in the embodiment of the present application can input a sample image containing road surface elements into the model to be trained to perform road surface element segmentation, and obtain a label map of the sample image. The pixel value of the pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs. The target pixel whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition can be extracted from the label map, and the model parameters of the model to be trained are adjusted based on the extracted target pixel. In the image segmentation task, especially when it involves the segmentation of complex scenes such as road environments, relying solely on pixel-level classification may not be enough to capture higher-level spatial relationships and structural information. Therefore, the present application introduces the topological relationship of road surface elements as a constraint, so that the segmentation result of the model is not only more accurate at the pixel level, but also follows the spatial relationship. This method can more accurately understand and segment the road surface elements in the image. By training the model using the above method, the segmentation accuracy of the road surface element segmentation model can be improved.

[0098] Figure 3 Schematic diagram of the model training method provided for this application Figure 2 ,like Figure 3 As shown, in this embodiment Figure 2Based on the embodiment, the model training method is described in detail, and the method includes:

[0099] S301, inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs.

[0100] The sample image may be a video frame image. Before the sample image is input into the model to be trained, the sample image may be floated, normalized, and the pixel arrangement may be adjusted.

[0101] In a possible implementation, a sample image containing road surface elements is input into a model to be trained to segment the road surface elements, and a label map of the sample image is obtained, which may specifically include:

[0102] Input a sample image containing road surface elements into the model to be trained to extract features, and obtain a feature map of the sample image;

[0103] Perform pixel-level category label mapping on the feature map to generate a label map. Different category labels correspond to different road feature categories.

[0104] The semantic segmentation network in the road feature segmentation model to be trained can be a ResNet18 network. ResNet18 is a commonly used convolutional neural network architecture with a shallow number of layers and low computational complexity, and is suitable for fast feature extraction in resource-limited environments.

[0105] The sample image is input into the semantic segmentation network for image feature extraction to generate a feature map, which contains the feature information of different areas in the sample image. Based on the feature map, each pixel is classified to predict the road feature category to which it belongs. According to the pixel-level classification results, a label map is generated. Each pixel value in the label map represents the road feature category to which the pixel belongs. For example, the category label 3 represents a symbol feature, 5 represents an arrow feature, and 7 represents a text feature.

[0106] In a possible implementation, performing pixel-level category label mapping on the feature map to generate a label mapping map may specifically include:

[0107] The feature map is mapped to pixel-level category labels through the argmax operation to generate a label map.

[0108] Each pixel in the feature map contains multiple channels, each channel corresponds to a predicted score for a road feature category, and this predicted score represents the probability that the pixel belongs to the corresponding road feature category. For example, if there are N road feature categories, each pixel in the feature map will have N channels, and the value of each channel can be regarded as the predicted score of the corresponding road feature category. For each pixel in the feature map, the argmax operation is used to select the channel with the highest score. Specifically, the argmax function traverses all channels of the pixel and returns the index of the channel with the highest score, which corresponds to the predicted category of the pixel. For example, if the channel index 3 has the highest score, the pixel is predicted to belong to road feature category 3. By applying the argmax operation to each pixel of the feature map, a label map is generated.

[0109] Reference Figure 4 As shown, it is a schematic diagram of the label map provided by the present application, and the width and height of the label map are equal to the width and height of the input sample image. The label map is a single-channel image, and its pixel value corresponds to the road surface feature category, and the road surface feature category is a predicted value. As shown in the figure, 2 represents the road surface area, 3 represents the symbol feature, 5 represents the arrow feature, and 7 represents the text feature.

[0110] S302: Based on the predefined road surface element topological relationship, identify the target pixel from the label map, where the target pixel is a pixel whose matching degree with the road surface element topological relationship meets a preset matching degree condition.

[0111] In a possible implementation, the road surface element topological relationship includes a mutually exclusive relationship, and based on the predefined road surface element topological relationship, identifying the target pixel from the label map may specifically include:

[0112] Perform convolution processing on the label map to obtain a convolution image;

[0113] For any pixel in the label map, determine its matching degree with the mutually exclusive relationship based on its pixel value in the label map and the pixel value of its corresponding pixel in the convolution image;

[0114] The pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is determined as the identified first target pixel.

[0115] In this embodiment, a pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is referred to as a first target pixel.

[0116] The label map is convolved to obtain a convolved image. For each pixel in the label map, its pixel value is checked, and at the same time, the pixel value of its corresponding pixel in the convolved image is checked. For example, if a pixel in the label map is checked and its pixel value is equal to 5, it means that it belongs to the arrow element. At the same time, the pixel value of the corresponding pixel in the convolution image is checked. Since multiple convolution images can be obtained by convolution processing of the label map, for this pixel, it is assumed that the pixel value of the corresponding pixel in the first convolution image is equal to -2 (the result of subtracting the text element label value from the arrow element label value), the pixel values ​​of the corresponding pixels in the second and third convolution images are also equal to -2, the pixel values ​​of the corresponding pixels in the fourth, fifth, sixth and seventh convolution images are equal to 0 (the result of subtracting the arrow element label value from the arrow element label value), and the pixel value of the corresponding pixel in the seventh convolution image is equal to 3 (the result of subtracting the road surface area label value from the arrow element label value). The number of convolution images that do not meet the mutually exclusive relationship is counted. In this example, there are 3 convolution images (the first, second and third) that do not meet the mutually exclusive relationship. A matching degree metric can be defined, such as matching degree = 1-(the number of convolution images that do not meet the mutually exclusive relationship / the total number of convolution images). The preset matching degree condition can be set to be less than a preset matching degree threshold. If the calculated matching degree is less than the preset matching degree threshold, it can be considered that the pixel does not meet the mutually exclusive relationship, and the pixel position needs to be paid attention to and needs to be extracted. Through the above method, all pixels that do not meet the mutually exclusive relationship in the label map can be extracted.

[0117] In a possible implementation, performing convolution processing on the label map to obtain a convolution image may specifically include:

[0118] The label map is convolved to calculate the difference between a pixel in the label map and its neighboring pixels to obtain a convolved image.

[0119] For the label map, a convolution kernel can be used to perform convolution to analyze the relationship between each pixel and its neighboring pixels. Figure 5 As shown, it is a schematic diagram of the convolution kernel provided in this application. The size of the convolution kernel is 3×3, which means that the convolution operation of each pixel will take into account itself and its surrounding 8 neighboring pixels. For each pixel in the label map (except the boundary pixels), the difference between it and the 8 neighboring pixels is calculated. For the boundary pixels, due to the lack of a complete neighborhood, the convolution cannot be calculated directly, so the pixel values ​​of these boundary pixels can be set to 255. Through the above operations, 8 convolution images can be generated. Each pixel value in the convolution image represents the difference between the pixel at the corresponding position in the label map and its neighboring pixels. Refer to Figure 6As shown, it is a schematic diagram of the convolution image generated by convolution processing of the label mapping map provided in the present application. The width, height and number of channels of the convolution image are the same as those of the label mapping map. It can be seen from the figure that each pixel value of each convolution image represents the difference between the pixel at the corresponding position in the label mapping map and its domain pixel.

[0120] It should be noted that the goal of convolution processing is to analyze the relationship between each pixel and its neighboring pixels. In addition to calculating the difference between a pixel and its neighboring pixels in the label map, the label map can also be processed in other ways to obtain a convolved image. For example, the sum of the pixel and its neighboring pixels in the label map can be calculated; the product of the pixel and its neighboring pixels in the label map can be calculated; the weighted convolution kernel can be used to perform weighted averaging on the neighboring pixels, etc. The specific processing method can be selected and combined according to the specific application requirements.

[0121] In a possible implementation, the road surface element topological relationship includes an inclusion relationship, and based on the predefined road surface element topological relationship, identifying the target pixel from the label map includes:

[0122] Get the true value label map of the sample image;

[0123] For any pixel in the label map, determine its matching degree with the inclusion relationship based on its pixel value in the label map and the pixel value of its corresponding pixel in the true value label map;

[0124] A pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is determined as the identified second target pixel.

[0125] In this embodiment, a pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is referred to as a second target pixel.

[0126] The true value label map can be a label map that is manually annotated or annotated using an annotation tool, which represents the real road feature category to which each pixel in the sample image belongs. The width and height of the true value label map are equal to the width and height of the sample image. The true value label map is a single-channel image.

[0127] For each pixel in the label map, obtain its pixel value, and at the same time, obtain the pixel value of its corresponding pixel in the true value label map, and determine whether the two pixel values ​​are consistent. If they are consistent, it can be determined that the pixel meets the inclusion relationship. If they are inconsistent, it can be determined that the pixel does not meet the inclusion relationship. For example, if a pixel in the label map has a pixel value equal to 3, indicating that it belongs to a symbol element, and at the same time, the pixel value of the corresponding pixel in the true value label map is not 3, it can be determined that the matching degree of the pixel and the inclusion relationship is 0, the pixel does not meet the inclusion relationship, and the pixel position needs to be paid attention to and needs to be extracted. When the preset matching degree condition is set to mismatch, all pixels in the label map that do not meet the inclusion relationship can be extracted. Since there is usually only one true value label map, the matching degree of the pixel and the inclusion relationship is usually binary (0 or 1). A matching degree of 1 indicates that the inclusion relationship is met; a matching degree of 0 indicates that the inclusion relationship is not met.

[0128] S303 , marking target pixels and generating a topological relationship matching degree image, wherein the pixel value of a pixel in the topological relationship matching degree image represents the matching degree between the corresponding pixel in the label mapping image and the topological relationship of the road surface element.

[0129] The identified target pixels are marked to generate a topological relationship matching image. In the topological relationship matching image, the pixel value of each pixel represents the matching degree between the corresponding pixel in the label map and the topological relationship of the road surface element.

[0130] In a possible implementation, marking target pixels and generating a topological relationship matching image may specifically include:

[0131] Marking the position of the first target pixel to generate a first marked image;

[0132] Marking the position of the second target pixel to generate a second marked image;

[0133] A union operation is performed on the first labeled image and the second labeled image to generate a topological relationship matching degree image.

[0134] The target pixels include a first target pixel and a second target pixel.

[0135] In this embodiment, a binary image of the same size as the label map can be created, and the binary image is used to mark the pixel positions whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition. For the first target pixel, the pixel value of the corresponding pixel in the binary image can be set to 1, so as to mark the position of the first target pixel in the binary image and generate a first marked image.

[0136] Another binary image of the same size as the label map is created. For the second target pixel, the pixel value of the corresponding pixel in the binary image can be set to 1, thereby marking the position of the second target pixel in the binary image and generating a second labeled image.

[0137] A pixel-by-pixel union operation is performed on the first labeled image and the second labeled image. For example, for each pixel position, if the value of either the first labeled image or the second labeled image is 1, the value of the position is set to 1 in the topological relationship matching image.

[0138] In the generated topological relationship matching image, the pixel value of each pixel indicates the matching degree between the corresponding pixel of the pixel in the label map and the predefined road surface element topological relationship, or the pixel value of each pixel indicates whether the matching degree between the corresponding pixel of the pixel in the label map and the predefined road surface element topological relationship meets the preset matching degree condition. If the preset matching degree condition is met, it can be marked as 0 (or other values ​​indicating compliance), and if it is not met, it can be marked as 1 (or other values ​​indicating non-compliance).

[0139] Reference Figure 7 and Figure 8 , which are schematic diagrams of the first marked image and the second marked image provided by the present application, respectively. A pixel value of 1 indicates that the preset matching degree condition is met, and a pixel value of 0 indicates that the preset matching degree condition is not met. Fig. 9 , which is a schematic diagram of a topological relationship matching degree image generated by finding the union provided in this application.

[0140] S304: Calculate the model loss based on the topological relationship matching image.

[0141] S305: Adjust model parameters of the model to be trained based on the model loss.

[0142] Based on the topological relationship matching image, the model loss is calculated. The loss function can be designed to consider both classification error and topological relationship error. Classification error measures the difference between the predicted category and the true category of each pixel. Topological relationship error measures the degree of match between the predicted result and the predefined topological relationship of the road surface elements. For example, if the spatial relationship between the crosswalk and the stop line does not meet the expectations, the loss will increase. According to the calculated model loss, the model parameters can be updated through the back propagation algorithm. Through repeated iterative training, the model gradually learns not only to correctly classify each pixel, but also to ensure that the spatial relationship between these pixels meets the predefined topological relationship constraints, thereby improving the model's ability to segment complex road surface elements. This approach enables the model to have high classification accuracy not only at the pixel level, but also to understand and maintain the reasonable topological relationship between road surface elements at a higher level.

[0143] In a possible implementation, the model loss is calculated based on the topological relationship matching image, which may specifically include:

[0144] Use the cross entropy loss function to calculate the difference between the feature map of the sample image and the true value label map of the sample image to obtain the cross entropy loss;

[0145] Based on the topological relationship matching image, the cross entropy loss is weighted and the model loss is calculated.

[0146] In this embodiment, the cross entropy loss can be weighted using the topological relationship matching image. For example, locations that do not conform to the topological relationship of the road surface elements can be given higher weights to emphasize the impact of these errors in the loss.

[0147] Reference Fig.10 As shown, it is a schematic diagram of the training process of the model training method provided in the present application, a sample image is obtained from a training data set, the sample image is input into the semantic segmentation network in the road feature segmentation model to extract image features, a feature map is obtained, a true value label mapping map of the sample image is obtained, the cross entropy loss is calculated based on the feature map and the true value label mapping map, a series of processing is performed on the feature map, the topological relationship loss is determined in combination with the true value label mapping map, the model loss is determined based on the cross entropy loss and the topological relationship loss, and according to the calculated model loss, the model parameters can be updated by acting on the semantic segmentation network through the back propagation algorithm.

[0148] Reference Fig.11 As shown, it is a schematic diagram of the calculation process of the model loss provided in this application. After the image features of the sample image are extracted through the semantic segmentation network and the feature map is obtained, the feature map can be mapped to the category label at the pixel level to generate a label map. For the label map, a 3×3 convolution kernel can be used for convolution processing to obtain an 8-neighborhood convolution image. Based on the 8-neighborhood convolution image, combined with the predefined road element mutual exclusion relationship, pixels that do not conform to the mutual exclusion relationship can be extracted from the label map, and a first labeled image can be generated. On the other hand, based on the true value label map and the predefined road element inclusion relationship, pixels that do not conform to the inclusion relationship can be extracted from the label map, and a second labeled image can be generated. The first labeled image and the second labeled image are unioned to generate a topological relationship matching image. The topological relationship matching image is used to adjust the cross entropy loss to form the final model loss. The trained road element segmentation model can perform road element segmentation, and the output results include the location, shape, and attributes of the road element.

[0149] Traditional road feature segmentation models improve segmentation accuracy by designing effective neural network architectures and loss functions. For example, the STDC model designs an STDC module and embeds it into the U-net architecture. This design utilizes the feature extraction capabilities of STDC while retaining the structural advantages of U-net. The STDC model also designs a special STDC loss function to guide the training of the model, thereby achieving higher segmentation accuracy in semantic segmentation tasks.

[0150] Traditional road feature segmentation models usually use pixel-level loss functions to guide model training. Although this method can capture detailed information of the image to a certain extent, it often ignores the topological relationship between elements in a specific scene. The embodiments of the present application incorporate these topological relationships into the model training process, which can significantly improve the training efficiency and reasoning accuracy of the model.

[0151] The model training method provided by the embodiment of the present application can input a sample image containing road surface elements into the model to be trained to perform road surface element segmentation, and obtain a label map of the sample image. The pixel value of the pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs. The target pixel whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition can be extracted from the label map, and the extracted target pixel is marked to generate a topological relationship matching degree image. The topological relationship matching degree image marks the pixel position whose matching degree with the predefined road surface element topological relationship meets the preset matching degree condition. Through the topological relationship matching degree image, the model loss can be calculated, and the model parameters can be adjusted based on the model loss. In the image segmentation task, especially when it involves the segmentation of complex scenes such as road environments, relying solely on pixel-level classification may not be enough to capture higher-level spatial relationships and structural information. Therefore, the present application introduces the topological relationship of road surface elements as a constraint, so that the segmentation result of the model is not only more accurate at the pixel level, but also follows the spatial relationship. This method can more accurately understand and segment the road surface elements in the image. By training the model using the above method, the segmentation accuracy of the road surface element segmentation model can be improved.

[0152] Fig.12 A schematic diagram of the structure of the model training device provided in this application, such as Fig.12 As shown, the model training device 120 provided in this embodiment includes:

[0153] The segmentation module 1201 is used to input the sample image containing the road surface elements into the model to be trained to perform road surface element segmentation, and obtain a label map of the sample image, wherein the pixel value of the pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs;

[0154] The identification module 1202 is used to identify target pixels from the label map based on the predefined road surface element topological relationship, where the target pixels are pixels whose matching degree with the road surface element topological relationship meets a preset matching degree condition;

[0155] The adjustment module 1203 is used to adjust the model parameters of the model to be trained based on the target pixel.

[0156] In a possible implementation, the adjustment module is specifically used to:

[0157] Mark the target pixel and generate a topological relationship matching image. The pixel value of the pixel in the topological relationship matching image represents the matching degree between the corresponding pixel in the label mapping image and the topological relationship of the road surface element.

[0158] Calculate the model loss based on the topological relationship matching image;

[0159] Based on the model loss, adjust the model parameters of the model to be trained.

[0160] In a possible implementation, the topological relationship of the road surface elements includes a mutually exclusive relationship, and the identification module is specifically used to:

[0161] Perform convolution processing on the label map to obtain a convolution image;

[0162] For any pixel in the label map, determine its matching degree with the mutually exclusive relationship based on its pixel value in the label map and the pixel value of its corresponding pixel in the convolution image;

[0163] The pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is determined as the identified first target pixel.

[0164] In a possible implementation, the identification module is specifically used to:

[0165] The label map is convolved to calculate the difference between a pixel in the label map and its neighboring pixels to obtain a convolved image.

[0166] In a possible implementation, the predefined road surface element topological relationship includes an inclusion relationship, and the identification module is specifically used to:

[0167] Get the true value label map of the sample image;

[0168] For any pixel in the label map, determine its matching degree with the inclusion relationship based on its pixel value in the label map and the pixel value of its corresponding pixel in the true value label map;

[0169] A pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is determined as the identified second target pixel.

[0170] In a possible implementation, the adjustment module is specifically used to:

[0171] Marking the position of the first target pixel to generate a first marked image;

[0172] Marking the position of the second target pixel to generate a second marked image;

[0173] A union operation is performed on the first labeled image and the second labeled image to generate a topological relationship matching degree image.

[0174] In a possible implementation, the segmentation module is specifically used to:

[0175] Input a sample image containing road surface elements into the model to be trained to extract features, and obtain a feature map of the sample image;

[0176] Perform pixel-level category label mapping on the feature map to generate a label map. Different category labels correspond to different road feature categories.

[0177] In a possible implementation, the adjustment module is specifically used to:

[0178] Use the cross entropy loss function to calculate the difference between the feature map of the sample image and the true value label map of the sample image to obtain the cross entropy loss;

[0179] Based on the topological relationship matching image, the cross entropy loss is weighted and the model loss is calculated.

[0180] The model training device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and will not be described in detail in this embodiment.

[0181] Fig.13 This is a schematic diagram of the structure of the model training device provided in this application. Fig.13 As shown, the model training device 130 provided in this embodiment includes: at least one processor 1301 and a memory 1302. Optionally, the device 130 also includes a communication component 1303. The processor 1301, the memory 1302 and the communication component 1303 are connected via a bus.

[0182] In a specific implementation process, at least one processor 1301 executes the computer-executable instructions stored in the memory 1302, so that at least one processor 1301 executes the above method.

[0183] The specific implementation process of the processor 1301 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0184] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0185] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0186] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0187] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0188] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0189] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0190] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0191] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0192] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0194] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0195] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0196] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A model training method, characterized in that: include: Inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs; Based on the predefined road surface element topological relationship, identifying a target pixel from the label map, the target pixel being a pixel whose matching degree with the road surface element topological relationship satisfies a preset matching degree condition; Based on the target pixel, a model parameter of the model to be trained is adjusted.

2. The method according to claim 1, characterized in that The adjusting the model parameters of the to-be-trained model based on the target pixel includes: Marking the target pixel and generating a topological relationship matching image, wherein the pixel value of the pixel in the topological relationship matching image represents the matching degree between the corresponding pixel in the label mapping image and the topological relationship of the road surface element; Calculating model loss based on the topological relationship matching image; Based on the model loss, the model parameters of the model to be trained are adjusted.

3. The method according to claim 2, characterized in that The road surface element topological relationship includes a mutually exclusive relationship, and the identifying of the target pixel from the label map based on the predefined road surface element topological relationship includes: Performing convolution processing on the label map to obtain a convolution image; For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the convolution image, determine its matching degree with the mutually exclusive relationship; A pixel whose matching degree with the mutually exclusive relationship satisfies a preset matching degree condition is determined as the identified first target pixel.

4. The method according to claim 3, characterized in that The road surface element topological relationship includes an inclusion relationship, and the identifying of target pixels from the label map based on the predefined road surface element topological relationship includes: Obtaining a true value label mapping diagram of the sample image; For any pixel in the label map, based on its pixel value in the label map and the pixel value of its corresponding pixel in the true value label map, determine its matching degree with the inclusion relationship; A pixel whose matching degree with the inclusion relationship satisfies a preset matching degree condition is determined as the identified second target pixel.

5. The method according to claim 4, characterized in that The step of marking the target pixel and generating a topological relationship matching image includes: Marking the position of the first target pixel to generate a first marked image; marking the position of the second target pixel to generate a second marked image; A union operation is performed on the first labeled image and the second labeled image to generate the topological relationship matching degree image.

6. The method according to any one of claims 1 to 5, characterized in that: The step of inputting a sample image containing road surface elements into a model to be trained to segment the road surface elements and obtaining a label map of the sample image includes: Inputting a sample image containing road surface elements into the model to be trained to extract features, and obtaining a feature map of the sample image; Perform pixel-level category label mapping on the feature map to generate the label mapping map, where different category labels correspond to different road surface element categories.

7. The method according to any one of claims 2 to 5, characterized in that: The calculating the model loss based on the topological relationship matching image includes: Using a cross entropy loss function to calculate the difference between the feature map of the sample image and the true value label mapping map of the sample image, to obtain a cross entropy loss; Based on the topological relationship matching image, the cross entropy loss is weighted to calculate the model loss.

8. A model training device, characterized in that: include: A segmentation module, used for inputting a sample image containing road surface elements into a model to be trained to perform road surface element segmentation, and obtaining a label map of the sample image, wherein the pixel value of a pixel in the label map represents the road surface element category to which the corresponding pixel in the sample image belongs; An identification module, for identifying target pixels from the label map based on a predefined road surface element topological relationship, wherein the target pixels are pixels whose matching degree with the road surface element topological relationship satisfies a preset matching degree condition; An adjustment module is used to adjust the model parameters of the model to be trained based on the target pixel.

9. A model training device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.

10. A computer-readable storage medium / computer program product, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor; and / or, The computer program product comprises a computer program, which implements the method according to any one of claims 1 to 7 when executed by a processor.