Model training method and device, electronic equipment, medium, product and vehicle
By modifying the bounding box of the bird's-eye view feature map of the environment image in the assisted driving scenario, training samples containing scarce elements are generated, which solves the problem of high cost of acquiring images of scarce environmental elements and achieves efficient model training and recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CONTINENTAL SMART CORE TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-28
AI Technical Summary
In assisted driving scenarios, the cost of acquiring images of scarce environmental elements is high, and existing technologies are insufficient to effectively train recognition models.
By acquiring the bird's-eye view feature map of the environment image, the outline of the element to be adjusted is determined, and the outline of the target element is modified according to the transformation parameters to generate an environment image containing scarce elements, which is used as a training sample for the model.
It eliminates the need for extensive environmental image collection, reducing model training costs and improving the accuracy of rare element identification.
Smart Images

Figure CN121505396B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of driver assistance technology, specifically to a model training method, device, electronic equipment, medium, product, and vehicle. Background Technology
[0002] In assisted driving scenarios, some environmental elements are relatively scarce, such as solid and dashed lane lines, and herringbone patterns. The richness of these environmental elements is much less than that of ordinary lane lines. If a recognition model for these scarce environmental elements needs to be trained specifically, a large number of images of these environmental elements need to be collected. Understandably, collecting a large number of images containing scarce environmental elements is quite difficult and costly. Summary of the Invention
[0003] This application provides a model training method, apparatus, electronic device, medium, product, and vehicle, which solves the problem of high model training costs caused by collecting a large number of images containing scarce environmental elements.
[0004] In a first aspect, embodiments of this application propose a model training method, which includes: acquiring a first bird's-eye view feature map corresponding to a first environmental image; determining a first outline of an element to be adjusted corresponding to a target element in the first bird's-eye view feature map; modifying the first outline to a second outline of the target element according to the conversion parameters between the element to be adjusted and the target element, thereby obtaining a second bird's-eye view feature map, wherein the second outline includes at least a portion of the first outline; generating a second environmental image based on the second bird's-eye view feature map, and using the second environmental image as a sample image of a first model.
[0005] It is understood that the aforementioned elements to be adjusted can include various targets to be identified in the assisted driving scenario, such as lane structure features, such as lane lines, curbs, and obstacles. The target element and the element to be adjusted are objects of the same type but different shapes. For example, both the target element and the element to be adjusted may be lane lines, but the element to be adjusted is a white solid lane line, while the target element is a herringbone pattern.
[0006] This allows for the editing of a second environmental image containing relatively scarce target elements from a single environmental image that includes the elements to be adjusted. The edited second environmental image can then be used as a training sample image for the first model used for subsequent road element recognition. Therefore, it eliminates the need to collect a large number of environmental images or deploy an image generation model, thus effectively reducing costs.
[0007] In one possible implementation of the first aspect above, modifying the first outline of the target element to the second outline of the target element according to the conversion parameters between the element to be adjusted and the target element includes: corresponding to the element to be adjusted being a solid lane line and the target element being a dashed lane line, converting the first outline of the element to be adjusted from a solid lane line to a dashed lane line corresponding to the second outline of the target element based on the first conversion parameters, wherein the first conversion parameters include the line type of the dashed lane line and the discontinuity dimension parameters of the dashed lane line; or, corresponding to the element to be adjusted being a lane line of a first size and the target element being a lane line of a second size, modifying the first outline of the element to be adjusted from the first size parameter to the second size parameter corresponding to the second outline of the target element based on the second conversion parameters. In this case, the first size parameter is smaller than the second size parameter, and the second conversion parameter includes the second size parameter; or, corresponding to the element to be adjusted being a single-lane line and the target element being a two-lane line, the first outer frame line of the element to be adjusted is copied and translated based on the third conversion parameter to obtain the second outer frame line of the target element, where the third conversion parameter is a preset distance parameter between the two-lane lines; or, corresponding to the element to be adjusted being a solid lane line and the target element being a herringbone line, an outer frame line area corresponding to the target element is added outside the first outer frame line of the element to be adjusted based on the fourth conversion parameter to obtain the second outer frame line of the target element, where the fourth conversion parameter includes the relative position parameter between the outer frame line area corresponding to the herringbone line and the solid lane line, and the size parameter corresponding to the herringbone line.
[0008] In some embodiments, when the element to be adjusted is a solid lane line and the target element is a dashed lane line, the electronic device can determine the difference between the two based on the size and shape definitions of the solid lane line and the dashed lane line in traffic rules. For example, it can determine a first conversion parameter including the line type of the dashed lane line and the discontinuity size parameter of the dashed lane line, and then convert the first outer frame line of the element to be adjusted from a solid lane line to a dashed lane line corresponding to the second outer frame line of the target element based on the location of the first outer frame line of the element to be adjusted (corresponding to the position parameter of the element to be adjusted).
[0009] In some embodiments, when the difference between the feature to be adjusted and the target feature is a size difference, for example, corresponding to a lane line of a first size for the feature to be adjusted and a lane line of a second size for the target feature, the size of the lane line of the first size can be modified to the second size parameter based on the second size parameter indicated by the second conversion parameter. For example, if the feature to be adjusted is a narrow line and the target feature is a wide line, the electronic device can modify the first outer frame line of the feature to be adjusted from the first size parameter to the second size parameter corresponding to the second outer frame line of the target feature, so that the first outer frame line corresponding to the narrow line is adjusted to the second outer frame line corresponding to the wide line.
[0010] In some embodiments, when the difference between the element to be adjusted and the target element is large, for example, the element to be adjusted is a single-lane line and the target element is a two-lane line, the electronic device can copy and translate the single-lane line of the first outer frame of the element to be adjusted based on the preset distance parameter between the two-lane lines indicated by the third conversion parameter, so as to obtain the second outer frame line of the target element as a two-lane line.
[0011] In one possible implementation of the first aspect described above, generating a second environmental image based on a second bird's-eye view feature map includes: projecting the second outline of the target element from the second bird's-eye view feature map onto a two-dimensional plane corresponding to the first environmental image according to the camera parameters used to capture the first environmental image, to obtain a third environmental image including the second outline; and filling the second outline in the third environmental image with corresponding pixels in the pixel region of the element to be adjusted in the first environmental image, according to the pixel region of the element to be adjusted in the first environmental image, to obtain a second environmental image including the target element.
[0012] It is understood that the electronic device can project the second outline of the second bird's-eye view feature map from the BEV plane to the image plane where the first environment image is located, so that the second outline is imaged in the image plane where the first environment image is located, and fill the corresponding pixels within the second outline to generate a second environment image including the target element. Furthermore, the second environment image may include relatively rare road elements, which can be used as sample images for training the first model, such as a road element recognition model.
[0013] In one possible implementation of the first aspect described above, based on the pixel region of the element to be adjusted in the first environmental image, the corresponding pixels in the pixel region are filled in the second outer frame of the third environmental image to obtain a second environmental image including the target element. This includes: determining the position of the pixel point to be filled in the second outer frame; determining each first pixel point in the pixel region of the element to be adjusted in the first environmental image that satisfies a first distance condition with respect to the position of each pixel point to be filled, and taking each first pixel point as a pixel point to be filled; filling each pixel point to be filled into the corresponding pixel point position in the third environmental image to obtain a second environmental image including the target element.
[0014] It is understood that the second outline of the target element can be determined on the BEV plane, and the second outline can be filled based on the pixels of the element to be adjusted in the first environment image, so that the outline of the target element in the obtained second environment image is more natural.
[0015] In one possible implementation of the first aspect above, the first distance condition includes: the pixel in the pixel region corresponding to the element to be adjusted that is closest to the pixel in the second outer frame that is to be filled.
[0016] In one possible implementation of the first aspect above, each pixel to be filled is filled into the corresponding pixel position in the third environment image to obtain a second environment image including the target element. The method further includes: filling each pixel to be filled into the corresponding pixel position in the third environment image to obtain a filled fourth environment image including the target element; and performing noise addition processing and / or color channel value adjustment processing on the target element in the fourth environment image to obtain a second environment image including the target element. The noise addition processing includes at least one of the following: dilation processing, erosion processing, and weighted aging processing based on pixels in the pixel region that satisfy the first distance condition.
[0017] Dilation can be understood as expanding the bright areas of a target element from the center outwards, while erosion can shrink the bright areas of a target element inwards. Bright areas can be pixel regions whose corresponding grayscale values are less than or equal to a preset grayscale threshold. Weighted aging processing can be used to simulate image aging and wear effects. For example, electronic devices can add aging features such as noise, scratches, and fading to target elements in a fourth-environment image by weighting pixel values, making the target elements appear to have a more natural wear effect.
[0018] In one possible implementation of the first aspect described above, corresponding to the solid lane line as the element to be adjusted and the dashed lane line as the target element, the corresponding pixels in the pixel region of the element to be adjusted in the first environmental image are filled in the second outer frame of the third environmental image to obtain a second environmental image including the target element. The implementation further includes: determining the position of each road surface pixel to be filled within the discontinuous region of the dashed lane line; determining each second pixel in the pixel region corresponding to the road surface in the first environmental image that satisfies a second distance condition with respect to the position of each road surface pixel to be filled, and using each second pixel as a road surface pixel to be filled; and filling each road surface pixel to be filled into the corresponding position of the road surface pixel to be filled in the third environmental image to obtain a second environmental image including the dashed lane line.
[0019] The second distance condition may include: a pixel region in the pixel region corresponding to the road surface whose distance to each road surface pixel point to be filled is less than or equal to a second preset distance threshold, or the pixel point in the pixel region corresponding to the road surface whose nearest neighbor is each road surface pixel point to be filled.
[0020] It can be understood that the discontinuous region is the area obtained after removing the dashed lane lines from the solid lane lines. In order to make the final second environmental image, which includes the dashed lane lines, appear natural, the road surface pixels of the first outer frame can be filled into the discontinuous region.
[0021] In one possible implementation of the first aspect above, the method further includes: recognizing the second environment image based on the second model to obtain a recognition result, wherein the second model includes a lane line detection model or a semantic segmentation model; and determining the second environment image as a sample image of the first model when the recognition result is a target feature.
[0022] It is understandable that electronic devices can use lane detection models or semantic segmentation models to verify whether the target elements in the second environment image meet the recognition requirements, so as to ensure that the second environment image can be used in the subsequent training tasks of the first model.
[0023] In one possible implementation of the first aspect above, obtaining the first bird's-eye view feature map corresponding to the first environmental image includes: projecting the first environmental image onto a bird's-eye view plane to obtain a bird's-eye view image; and performing feature recognition on the bird's-eye view image to obtain the first bird's-eye view feature map.
[0024] The bird's-eye view (BEV) is a top-down perspective representing the vehicle's surrounding environment through data fusion from multiple sensors (such as cameras and LiDAR). Therefore, the first bird's-eye view feature map is a feature map identified from the bird's-eye view image (corresponding to the top-down view) of the first environmental image on the bird's-eye view plane. Thus, the first bird's-eye view feature map can include various elements from the vehicle's surrounding environment image, such as static road elements and dynamic road elements. In this embodiment, the elements to be adjusted are static road elements, such as lane lines, curbs, and road barriers (e.g., traffic cones).
[0025] In one possible implementation of the first aspect described above, the camera parameters include camera intrinsic parameters, camera extrinsic parameters, and / or camera pose parameters.
[0026] Secondly, this application also proposes a model training device, which includes an image acquisition module, a feature recognition module, a preprocessing module, and a data processing module. The image acquisition module is used to acquire a first bird's-eye view feature map corresponding to a first environmental image. The feature recognition module is used to determine the first outer frame line of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map. The preprocessing module is used to modify the first outer frame line to a second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element, thereby obtaining a second bird's-eye view feature map, wherein the second outer frame line includes at least a portion of the first outer frame line. The data processing module is used to generate a second environmental image based on the second bird's-eye view feature map, and use the second environmental image as a sample image of the first model.
[0027] Thirdly, embodiments of this application also provide an electronic device, including: one or more processors; one or more memories; the one or more memories storing one or more programs, which, when executed by one or more processors, cause the electronic device to execute the model training method proposed in the first aspect and various implementations of the first aspect.
[0028] Fourthly, embodiments of this application also provide a computer-readable medium storing instructions that, when executed on a machine, cause the machine to perform the model training method proposed in the first aspect and various implementations thereof.
[0029] Fifthly, embodiments of this application also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the model training method proposed in the first aspect and various implementations of the first aspect.
[0030] Sixthly, embodiments of this application also provide a vehicle including the electronic equipment proposed in the second aspect above.
[0031] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be referred to the first aspect and the beneficial effects of various implementations of the first aspect, which will not be elaborated here.
[0032] The technical solution provided in this application has at least the following beneficial effects:
[0033] This application embodiment realizes the editing of a second environmental image containing relatively scarce target elements from a single environmental image including the elements to be adjusted. The edited second environmental image is then used as a training sample image for a first model for subsequent road element recognition. Therefore, it is not necessary to collect a large number of environmental images or deploy an image generation model, thereby effectively reducing costs. Attached Figure Description
[0034] Figure 1 A schematic diagram illustrating the implementation process of a model training method according to an embodiment of this application is shown.
[0035] Figure 2 A scene diagram is shown, illustrating a fishbone line as a target element according to an embodiment of this application.
[0036] Figure 3 A schematic diagram of a nearest neighbor pixel filling proposed according to some embodiments of this application is shown;
[0037] Figure 4 A schematic diagram of nearest neighbor road surface pixel filling according to some embodiments of this application is shown;
[0038] Figure 5 A schematic diagram of the framework structure of a model training device according to this application is shown;
[0039] Figure 6 This application provides a schematic diagram illustrating a possible functional framework of a vehicle according to an embodiment. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0041] It is understood that in some implementations, image generation models based on deep learning algorithms can also be used to generate road images. However, directly generating road images may result in insufficient stability and geometric consistency of the generated images, such as incorrect scale or unstable clarity. Furthermore, deploying the aforementioned image generation models is costly. Therefore, directly generating road images based on image generation models is not the preferred solution.
[0042] To address the issue of excessively high model training costs due to the large number of environmental images collected, this application proposes a model training method. The method includes: acquiring a first bird's-eye view feature map corresponding to a first environmental image; determining the first outline of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map; modifying the first outline to a second outline of the target element based on the conversion parameters between the element to be adjusted and the target element, thereby obtaining a second bird's-eye view feature map, wherein the second outline includes at least a portion of the first outline; generating a second environmental image based on the second bird's-eye view feature map, and using the second environmental image as a sample image for the first model.
[0043] The technical solution provided in this application has at least the following beneficial effects:
[0044] This application embodiment realizes the editing of a second environmental image containing relatively scarce target elements from a single environmental image including the elements to be adjusted. The edited second environmental image is then used as a training sample image for a first model for subsequent road element recognition. Therefore, it is not necessary to collect a large number of environmental images or deploy an image generation model, thereby effectively reducing costs.
[0045] It is understood that the aforementioned elements to be adjusted can include various targets to be identified in the assisted driving scenario, such as lane structure features, such as lane lines, curbs, and obstacles. The target element and the element to be adjusted are objects of the same type but different shapes. For example, both the target element and the element to be adjusted may be lane lines, but the element to be adjusted may be a white single lane line, while the target element may be a herringbone pattern.
[0046] Therefore, by adjusting the geometry of the first outer frame of the element to be adjusted to the geometry of the second outer frame of the target element, and by filling the second outer frame of the target element based on the pixels of the element to be adjusted, a target element of the same color can be obtained.
[0047] In other embodiments, the pixels of the element to be adjusted may also be color modified, for example, by adjusting the color channel values.
[0048] The following example, which uses lane lines as both the element to be adjusted and the target element, will be used to illustrate the embodiments of this application in detail.
[0049] Figure 1 A schematic diagram illustrating the implementation process of a model training method proposed according to an embodiment of this application is shown.
[0050] It is understood that the subject executing this model training method can be any form of electronic device, such as a terminal or server. In some embodiments, the electronic device can be an in-vehicle electronic device, but for the sake of simplicity, it will not be elaborated here.
[0051] refer to Figure 1 The implementation process includes the following steps:
[0052] S101, Obtain the first bird's-eye view feature map corresponding to the first environment image.
[0053] In some embodiments, when a vehicle is driving on a road, it can acquire a first environmental image of the surrounding environment through sensors. Electronic devices can determine a first bird's-eye view feature map based on the first environmental image. For example, the first environmental image can be projected onto a bird's-eye view plane to obtain a bird's-eye view image; and feature recognition can be performed on the bird's-eye view image to obtain the first bird's-eye view feature map.
[0054] Among them, the bird's eye view (BEV) is a top-down view of the vehicle's surrounding environment, which is represented by data fusion from multiple sensors (such as cameras, lidar, etc.). Therefore, the first bird's eye feature map is the feature map identified by the bird's eye image (corresponding to the top view) of the first environment image on the bird's eye view plane.
[0055] Therefore, the first bird's-eye view feature map can include various elements in the image of the vehicle's surrounding environment, such as static road elements and dynamic road elements. The elements to be adjusted proposed in this application embodiment are static road elements, such as lane lines, curbs, and roadblocks (e.g., traffic cones).
[0056] S102, determine the first outer frame of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map.
[0057] It can be understood that the target element is the final element style that needs to be modified. Therefore, the electronic device first determines the first outline of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map, so as to modify the first outline based on the second outline of the target element.
[0058] In some embodiments, if the element to be adjusted is a solid white lane line, the outer frame of the area corresponding to the solid white lane line can be selected as the first outer frame, which appears as a long rectangle in the bird's-eye view. Then, the first rectangular outer frame can be modified based on the specific shape of the second outer frame. For example, if the second outer frame is a dashed lane line, the first outer frame can be modified to a dashed lane line according to the size definition of dashed lane lines in traffic rules. For instance, the rectangular range of the first outer frame can be broken at preset intervals, modifying it into an intermittent rectangular frame.
[0059] S103, based on the conversion parameters between the element to be adjusted and the target element, the first outer frame line is modified to the second outer frame line of the target element to obtain the second bird's-eye view feature map, wherein the second outer frame line includes at least a part of the first outer frame line.
[0060] It is understood that this transformation parameter can be determined based on the difference between the feature to be adjusted and the target feature. Therefore, the electronic device can modify the first outline to a second outline, and the second outline may include at least a portion of the first outline.
[0061] For example, modifying the first outer frame line to the second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element includes: corresponding to the element to be adjusted being a solid lane line and the target element being a dashed lane line, converting the first outer frame line of the element to be adjusted from a solid lane line to a dashed lane line corresponding to the second outer frame line of the target element based on the first conversion parameters, wherein the first conversion parameters include the line type of the dashed lane line and the discontinuity dimension parameters of the dashed lane line; or, corresponding to the element to be adjusted being a lane line of a first size and the target element being a lane line of a second size, modifying the first outer frame line of the element to be adjusted from the first size parameter to the second size parameter corresponding to the second outer frame line of the target element based on the second conversion parameters, wherein the first The size parameter is smaller than the second size parameter, and the second conversion parameter includes the second size parameter; or, corresponding to the element to be adjusted being a single-lane line and the target element being a two-lane line, the first outer frame line of the element to be adjusted is copied and translated based on the third conversion parameter to obtain the second outer frame line of the target element, where the third conversion parameter is a preset distance parameter between the two-lane lines; or, corresponding to the element to be adjusted being a solid lane line and the target element being a herringbone line, an outer frame line area corresponding to the target element is added outside the first outer frame line of the element to be adjusted based on the fourth conversion parameter to obtain the second outer frame line of the target element, where the fourth conversion parameter includes the relative position parameter between the outer frame line area corresponding to the herringbone line and the solid lane line, and the size parameter corresponding to the herringbone line.
[0062] In some embodiments, when the element to be adjusted is a solid lane line and the target element is a dashed lane line, the electronic device can determine the difference between the two based on the size and shape definitions of the solid lane line and the dashed lane line in traffic rules. For example, it can determine a first conversion parameter including the line type of the dashed lane line and the discontinuity size parameter of the dashed lane line, and then convert the first outer frame line of the element to be adjusted from a solid lane line to a dashed lane line corresponding to the second outer frame line of the target element based on the location of the first outer frame line of the element to be adjusted (corresponding to the position parameter of the element to be adjusted).
[0063] In some embodiments, when the difference between the feature to be adjusted and the target feature is a size difference, for example, corresponding to a lane line of a first size for the feature to be adjusted and a lane line of a second size for the target feature, the size of the lane line of the first size can be modified to the second size parameter based on the second size parameter indicated by the second conversion parameter. For example, if the feature to be adjusted is a narrow line and the target feature is a wide line, the electronic device can modify the first outer frame line of the feature to be adjusted from the first size parameter to the second size parameter corresponding to the second outer frame line of the target feature, so that the first outer frame line corresponding to the narrow line is adjusted to the second outer frame line corresponding to the wide line.
[0064] In some embodiments, when the difference between the element to be adjusted and the target element is large, for example, the element to be adjusted is a single-lane line and the target element is a two-lane line, the electronic device can copy and translate the single-lane line of the first outer frame of the element to be adjusted based on the preset distance parameter between the two-lane lines indicated by the third conversion parameter, so as to obtain the second outer frame line of the target element as a two-lane line.
[0065] For example, refer to Figure 2 When the element to be adjusted is a solid lane line and the target element is a herringbone line, outside the first outer frame of the element to be adjusted (e.g., Figure 2 (Example solid lane line A1) The electronic device needs to add the outer frame area corresponding to the target feature (corresponding to the "fishbone" of the herringbone line, for example) Figure 2 Example fishbone T1), to obtain the second outer frame of the target element.
[0066] It is understandable that the above conversion rules can be determined based on the preset setting standards of the target elements and the setting standards of the elements to be adjusted.
[0067] S104, Generate a second environment image based on the second bird's-eye view feature map, and use the second environment image as a sample image of the first model.
[0068] It is understood that the electronic device can project the second outline of the second bird's-eye view feature map from the BEV plane to the image plane where the first environment image is located, so that the second outline is imaged in the image plane where the first environment image is located, and fill the corresponding pixels within the second outline to generate a second environment image including the target element. Furthermore, the second environment image may include relatively rare road elements, which can be used as sample images for training the first model.
[0069] In some embodiments, the first model may be, for example, a road element recognition model, which can be used to identify road elements in the driving environment during vehicle operation, such as lane lines, curbs, and road barriers (e.g., traffic cones).
[0070] For example, generating a second environmental image based on a second bird's-eye view feature map includes: projecting the second outline of the target element from the second bird's-eye view feature map onto a two-dimensional plane corresponding to the first environmental image according to the camera parameters used to capture the first environmental image, to obtain a third environmental image including the second outline; and filling the second outline in the third environmental image with corresponding pixels in the pixel region of the element to be adjusted in the first environmental image, according to the pixel region of the element to be adjusted in the first environmental image, to obtain a second environmental image including the target element.
[0071] In some embodiments, the second outline of the target feature can be projected from the BEV plane onto the image plane corresponding to the first environment image based on the camera parameters used to capture the first environment image. The camera parameters include camera intrinsic parameters, camera extrinsic parameters, and / or camera pose parameters.
[0072] In some embodiments, based on the pixel region of the element to be adjusted in the first environmental image, the corresponding pixels in the pixel region are filled in the second outer frame of the third environmental image to obtain a second environmental image including the target element. This includes: determining the position of the pixel point to be filled in the second outer frame; determining each first pixel point in the pixel region of the element to be adjusted in the first environmental image that satisfies a first distance condition with respect to the position of each pixel point to be filled, and using each first pixel point as a pixel point to be filled; filling each pixel point to be filled into the corresponding pixel point position in the third environmental image to obtain a second environmental image including the target element.
[0073] It is understood that the second outline of the target element can be determined on the BEV plane, and the second outline can be filled based on the pixels of the element to be adjusted in the first environment image, so that the outline of the target element in the obtained second environment image is more natural.
[0074] In some embodiments, the first distance condition includes: a pixel region in the region corresponding to the element to be adjusted whose distance from the pixel position to be filled in the second outer frame is less than or equal to a first preset distance threshold; or, the pixel in the region corresponding to the element to be adjusted that is closest to the pixel position to be filled in the second outer frame. This results in a more natural contour of the target element in the obtained second environmental image.
[0075] The following is combined with Figure 3 The method of filling the nearest neighbor pixel is illustrated.
[0076] Figure 3 A schematic diagram of a nearest neighbor pixel filling proposed according to some embodiments of this application is shown.
[0077] refer to Figure 3 In the fishbone pattern, the first pixel C1, the second pixel C2, the third pixel C3, the fourth pixel C4, and the fifth pixel C5 to be filled can all obtain their corresponding pixel values from the nearest point on the solid lane line B1 (corresponding to the element to be adjusted). This ensures that the target element in the modified second environment image maintains the same resolution as the first environment image, thus achieving the result shown in the example above. Figure 2 The style includes a second environmental image with fishbone lines.
[0078] In some embodiments, when the element to be adjusted is a solid lane line and the target element is a dashed lane line, the electronic device can fill the pixels of the road surface into the positions where the solid lane line needs to be broken to become a dashed lane line, so that the dashed lane line in the second environmental image blends naturally with the environment.
[0079] For example, corresponding to a solid lane line as the element to be adjusted and a dashed lane line as the target element, the second environmental image is obtained by filling the corresponding pixels in the second outer frame of the third environmental image with the pixel region of the element to be adjusted in the first environmental image, based on the pixel region of the element to be adjusted. The second environmental image includes the target element. The method further includes: determining the position of each road surface pixel point to be filled in the discontinuous region of the dashed lane line; determining each second pixel point in the pixel region corresponding to the road surface in the first environmental image that satisfies the second distance condition with the position of each road surface pixel point to be filled, and taking each second pixel point as the road surface pixel point to be filled; filling each road surface pixel point to be filled into the corresponding position of the road surface pixel point to be filled in the third environmental image, thereby obtaining the second environmental image including the dashed lane line.
[0080] The second distance condition may include: a pixel region in the pixel region corresponding to the road surface whose distance to each road surface pixel point to be filled is less than or equal to a second preset distance threshold, or the pixel point in the pixel region corresponding to the road surface whose nearest neighbor is each road surface pixel point to be filled.
[0081] The following is combined with Figure 4 The filling method for the nearest road surface pixels is illustrated.
[0082] Figure 4 A schematic diagram of nearest neighbor road surface pixel filling according to some embodiments of this application is shown.
[0083] refer to Figure 4 This illustrates an example of a dashed lane line 400 obtained by modifying a solid lane line, with the outer perimeter of the dashed lane line 400 being a road surface pixel region 406. The second outer frame of the dashed lane line 400 includes a first region 401, a second region 403, and a third region 405, while the discontinuous regions include a fourth region 402 and a sixth region 404.
[0084] It can be understood that the discontinuous region is the area obtained after removing the dashed lane lines from the solid lane lines. In order to make the final second environmental image, which includes the dashed lane lines, appear natural, the road surface pixels of the first outer frame can be filled into the discontinuous region.
[0085] Continue to refer to Figure 4Taking the first road surface pixel position D0 to be filled in the fourth region 402 as an example, its nearest neighbor is the first road surface pixel E1 in the road surface pixel region 406. Therefore, the first road surface pixel E1 can be filled into the first road surface pixel position D0. Similarly, taking the second road surface pixel position D1 to be filled in the fourth region 402 as an example, its nearest neighbor is the second road surface pixel E2 in the road surface pixel region 406. Therefore, the second road surface pixel E2 can be filled into the second road surface pixel position D1.
[0086] In some embodiments of this application, the pixels to be filled are filled into the corresponding pixel positions in the third environment image to obtain a filled fourth environment image including the target elements; after the target elements in the fourth environment image are subjected to noise addition processing and / or color channel value adjustment processing, a second environment image including the target elements is obtained; wherein, the noise addition processing includes at least one of the following: dilation processing, erosion processing, and weighted aging processing based on pixels in the pixel region that satisfy the first distance condition.
[0087] It can be understood that dilation expands the bright areas of a target feature outwards from the center, while erosion shrinks the bright areas inwards. The bright areas can be pixel regions whose corresponding grayscale values are less than or equal to a preset grayscale threshold.
[0088] Weighted aging processing can be used to simulate the effects of image aging and wear. For example, electronic devices can add aging features such as noise, scratches, and fading to target elements in a fourth environmental image by weighting pixel values, making the target elements appear to have a more natural wear effect. The aforementioned first distance condition can include pixel regions in the area corresponding to the element to be adjusted whose distance from the pixel position to be filled in the second outer frame is less than or equal to a first preset distance threshold, or the pixel in the area corresponding to the element to be adjusted that is closest to the pixel position to be filled in the second outer frame. Therefore, based on the pixel aging status of the element to be adjusted, weighted aging processing is applied to the target element.
[0089] In some embodiments, the electronic device can also adjust the color channel values of the target element so that the target element can be adjusted to different preset colors. For example, a double yellow solid line lane line (corresponding to the target element) can be edited based on a single white solid line lane line (corresponding to the element to be adjusted), a yellow road barrier (corresponding to the target element) can be edited based on an orange road barrier (corresponding to the element to be adjusted), or a black and yellow alternating road edge (corresponding to the target element) can be edited based on a yellow curb (corresponding to the element to be adjusted).
[0090] For example, the method further includes: recognizing a second environmental image based on a second model to obtain a recognition result, wherein the second model includes a lane line detection model or a semantic segmentation model; and determining the second environmental image as a sample image of the first model when the recognition result is a target feature.
[0091] It is understandable that electronic devices can use lane detection models or semantic segmentation models to verify whether the target elements in the second environment image meet the recognition requirements, so as to ensure that the second environment image can be used in the subsequent training tasks of the first model.
[0092] In some embodiments, the first model can be, for example, a road element recognition model, which can be used to identify road elements in the driving environment during vehicle operation, such as lane lines, curbs, and road barriers (e.g., traffic cones). Therefore, based on the model training method proposed in this application, the electronic device can use a second environmental image containing rare target elements to train the aforementioned road element recognition model, thereby improving the accuracy of the vehicle's recognition of road elements in the driving environment during operation.
[0093] Figure 5 A schematic diagram of the framework structure of a model training device 500 according to this application is shown. The model training device 500 includes an image acquisition module 501, a feature recognition module 502, a preprocessing module 503, and a data processing module 504. The image acquisition module 501 is used to acquire a first bird's-eye view feature map corresponding to a first environmental image. The feature recognition module 502 is used to determine the first outer frame line of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map. The preprocessing module 503 is used to modify the first outer frame line to a second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element, obtaining a second bird's-eye view feature map, wherein the second outer frame line includes at least a portion of the first outer frame line. The data processing module 504 is used to generate a second environmental image based on the second bird's-eye view feature map, and use the second environmental image as a sample image of the first model.
[0094] It is understood that the specific execution process of the image acquisition module 501 can be referred to the specific implementation process of step S101 above, and will not be repeated here.
[0095] It is understood that the specific execution process of the feature recognition module 502 can be referred to the specific implementation process of step S102 above, and will not be repeated here.
[0096] It is understood that the specific execution process of the preprocessing module 503 can be referred to the specific implementation process of step S103 above, and will not be repeated here.
[0097] It is understood that the specific execution process of the above data processing module 504 can refer to the specific implementation process of step S105 in the above text, and will not be repeated here.
[0098] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to implement the model training method proposed in any of the above embodiments.
[0099] According to the method provided in the embodiments of this application, this application also provides a computer-readable medium storing program code, which, when run on a computer, enables the computer to implement the model training method proposed in any of the above embodiments.
[0100] According to the method provided in the embodiments of this application, this application also provides a vehicle that includes the electronic equipment proposed in any of the above embodiments.
[0101] The specific structure of the above-mentioned vehicle will be described in detail below with reference to the accompanying drawings.
[0102] It is understood that the model training method mentioned in the embodiments of this application can be applied to electronic devices in vehicles. Figure 6 This application provides a schematic diagram illustrating a possible functional framework of a vehicle according to an embodiment.
[0103] like Figure 6 As shown, the functional framework of a vehicle may include various subsystems, such as Figure 6 The system includes a sensor system 1510, a control system 1520, one or more peripheral devices 1530 (one is shown as an example), a power supply 1540, and a computer system 1550. Optionally, the vehicle may also include other functional systems, such as an engine system that provides power to the vehicle, etc., which are not limited herein.
[0104] The sensor system 1510 may include several detection devices that can sense the measured information and convert the sensed information into electrical signals or other desired forms of information output according to a certain rule. For example... Figure 6 As shown, these detection devices may include a global positioning system (GPS), a vehicle speed sensor (VPS) (1512), an inertial measurement unit (IMU) (1513), etc., and this application does not limit them.
[0105] The Global Positioning System (GPS) 1511 is a system that uses GPS positioning satellites to perform real-time positioning and navigation globally. In this application, the GPS 1511 can be used to achieve real-time vehicle positioning and provide the vehicle's geographical location information. The vehicle speed sensor 1512 is used to detect the vehicle's speed. The inertial measurement unit 1513 may include a combination of an accelerometer and a gyroscope, and is a device for measuring the vehicle's angular rate and acceleration. For example, during vehicle movement, the inertial measurement unit can measure changes in the vehicle's position and angle based on the vehicle's inertial acceleration, such as measuring the vehicle's acceleration and angular rate.
[0106] The control system 1520 may include a steering unit 1521 and a braking unit 1522, etc.
[0107] Steering unit 1521 can represent a system for adjusting the direction of travel of a vehicle, which may include, but is not limited to, a steering wheel or other structural device for adjusting or controlling the direction of travel of a vehicle. Braking unit 1522 can represent a system for slowing down the speed of a vehicle, and may also be referred to as a vehicle braking system. It may include, but is not limited to, a brake controller, a reducer, or other structural device for slowing down a vehicle. In practical applications, braking unit 1522 can use friction to slow down the vehicle tires, thereby slowing down the vehicle's speed.
[0108] Peripheral device 1530 may include several components, such as Figure 6 The diagram shows a communication system 1531, a touchscreen 1532, a user interface 1533, etc. The communication system 1531 is used to enable network communication between the vehicle and other devices besides the vehicle.
[0109] In practical applications, the communication system 1531 can employ wireless or wired communication technologies to achieve network communication between vehicles and other devices. The wired communication technology can refer to communication between vehicles and other devices via network cables or fiber optic cables. The wireless communication technologies include, but are not limited to, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies, etc.
[0110] The touchscreen 1532 can be used to detect operation commands on the touchscreen 1532. For example, a user can perform touch operations on the content data displayed on the touchscreen 1532 according to actual needs to achieve the corresponding function, such as playing music, video, or other multimedia files. The user interface 1533 can specifically be a touch panel for detecting operation commands on the touch panel. The user interface 1533 can also be a physical button or a mouse. The user interface 1533 can also be a display screen for outputting data and displaying images or data. Optionally, the user interface 1533 can also be at least one device belonging to the category of peripheral devices, such as a touchscreen, microphone, and speaker.
[0111] Several functions of the vehicle are controlled and implemented by the computer system 1550. The computer system 1550 may include multiple processors such as processor 1551, a continuous damping control system (CDC) 1552, a mobile data center (MDC) 1553, an onboard telematics box (T-BOX) 1554, as well as a memory 1555 (also referred to as a storage device) and a gateway 1556. In practical applications, the memory 1555 may be located inside the computer system 1550 or outside the computer system 1550, for example, as a cache in the vehicle; this application does not limit this.
[0112] Among them, processors 1551, CDC 1552, MDC 1553, and T-BOX 1554 can be used to run relevant programs or instructions corresponding to programs stored in memory 1555 to realize the corresponding functions of the vehicle, such as the function of calling the vehicle camera.
[0113] Memory 1555 may include volatile memory, such as RAM; it may also include non-volatile memory, such as ROM, flash memory, HDD, or SSD; or it may include a combination of the above types of memory. Memory 1555 can be used to store a set of program code or instructions corresponding to program code, so that processor 1551 can call the program code or instructions stored in memory 1555 to implement the corresponding functions of the vehicle. This function includes, but is not limited to, […]. Figure 6 The illustrated vehicle functional framework diagram shows some or all of the functions. In this application, the memory 1555 can store a set of program code for vehicle control. The processors 1551, CDC 1552, MDC 1553, and T-BOX 1554 can call this program code to control the vehicle to perform the model training method shown in the example of this application.
[0114] Optionally, in addition to storing program code or instructions, memory 1555 may also store information such as road maps, driving routes, and sensor data. Computer system 1550 can be combined with other components in the vehicle functional framework diagram, such as sensors and GPS in the sensor system, to realize relevant vehicle functions. For example, computer system 1550 can control the vehicle's driving direction or speed based on data input from sensor system 1510; this application does not impose limitations on this.
[0115] It is understood that the various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0116] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0117] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used to implement the program code when needed. The mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0118] The above describes the possible hardware structures of electronic devices. It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of both.
[0119] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0120] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.
[0121] Although this application has been illustrated and described with reference to certain embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of this application.
Claims
1. A model training method, characterized in that, The method includes: Obtain the first bird's-eye view feature map corresponding to the first environmental image; Determine the first outer frame of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map; Based on the conversion parameters between the element to be adjusted and the target element, the first outer frame line is modified to the second outer frame line of the target element to obtain a second bird's-eye view feature map, wherein the second outer frame line includes at least a portion of the first outer frame line; A second environment image is generated based on the second bird's-eye view feature map, and the second environment image is used as a sample image of the first model; The step of modifying the first outer frame line to the second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element includes: If the element to be adjusted is a solid lane line and the target element is a dashed lane line, the first outer frame of the element to be adjusted is converted from the solid lane line to the dashed lane line corresponding to the second outer frame of the target element based on a first conversion parameter. The first conversion parameter includes the line type of the dashed lane line and the discontinuity dimension parameter of the dashed lane line; or... When the element to be adjusted is a solid lane line and the target element is a herringbone line, a new outer frame region corresponding to the target element is added outside the first outer frame region of the element to be adjusted based on the fourth conversion parameter to obtain the second outer frame region of the target element. The fourth conversion parameter includes the relative position parameter between the outer frame region corresponding to the herringbone line and the solid lane line, and the size parameter corresponding to the herringbone line.
2. The model training method according to claim 1, characterized in that, The generation of the second environmental image based on the second bird's-eye view feature map includes: Based on the camera parameters used to capture the first environmental image, the second outer frame line of the target element is projected from the second bird's-eye view feature map onto the two-dimensional plane corresponding to the first environmental image to obtain a third environmental image including the second outer frame line; Based on the pixel region of the element to be adjusted in the first environmental image, the corresponding pixels in the pixel region are filled into the second outer frame in the third environmental image to obtain the second environmental image including the target element.
3. The model training method according to claim 2, characterized in that, The step of filling the second outer frame in the third environment image with the corresponding pixels in the pixel region of the element to be adjusted in the first environment image to obtain the second environment image including the target element includes: Determine the positions of the pixels to be filled within the second outer frame; First pixels that satisfy the first distance condition between the pixel region of the element to be adjusted and the position of each pixel to be filled in the first environmental image are identified, and each first pixel is taken as the pixel to be filled. Each of the pixels to be filled is filled into the corresponding pixel position in the third environment image to obtain the second environment image including the target element.
4. The model training method according to claim 3, characterized in that, The first distance condition includes: the pixel in the pixel region corresponding to the element to be adjusted that is closest to the pixel in the second outer frame that is to be filled.
5. The model training method according to claim 3, characterized in that, The step of filling each of the pixels to be filled into the corresponding pixel positions in the third environment image to obtain the second environment image including the target element further includes: Each of the pixels to be filled is filled into the corresponding pixel position in the third environment image to obtain a fourth environment image including the target element; After adding noise and / or adjusting the color channel values of the target elements in the fourth environmental image, a second environmental image including the target elements is obtained. The noise-adding process includes at least one of the following: dilation processing, erosion processing, and weighted aging processing based on pixels in a pixel region that satisfy a first distance condition.
6. The model training method according to claim 2, characterized in that, Corresponding to the element to be adjusted being a solid lane line and the target element being a dashed lane line, the step of filling the second outer frame line in the third environment image with the corresponding pixels in the pixel region of the element to be adjusted in the first environment image to obtain the second environment image including the target element further includes: The position of each road surface pixel to be filled is determined within the discontinuous area of the dashed lane line; Identify each second pixel in the pixel region corresponding to the road surface within the first environmental image that satisfies the second distance condition between it and the location of each road surface pixel to be filled, and use each second pixel as the road surface pixel to be filled; Each of the road surface pixels to be filled is filled into the corresponding position of the road surface pixel in the third environment image to obtain the second environment image including the dashed lane line.
7. The model training method according to claim 1, characterized in that, The method further includes: The second environmental image is identified based on the second model to obtain the identification result, wherein the second model includes a lane line detection model or a semantic segmentation model; If the recognition result is the target element, the second environmental image is determined as a sample image of the first model.
8. The model training method according to claim 1, characterized in that, The step of obtaining the first bird's-eye view feature map corresponding to the first environment image includes: The first environmental image is projected onto the bird's-eye view plane to obtain the bird's-eye view image; The bird's-eye view image is subjected to feature recognition to obtain the first bird's-eye view feature map.
9. The model training method according to claim 2, characterized in that, The camera parameters include camera intrinsic parameters, camera extrinsic parameters, and / or camera pose parameters.
10. A model training device, characterized in that, The model training device includes an image acquisition module, a feature recognition module, a preprocessing module, and a data processing module, wherein... The image acquisition module is used to acquire a first bird's-eye view feature map corresponding to the first environment image; The feature recognition module is used to determine the first outer frame line of the element to be adjusted corresponding to the target element in the first bird's-eye view feature map; The preprocessing module is used to modify the first outer frame line to the second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element, thereby obtaining a second bird's-eye view feature map, wherein the second outer frame line includes at least a portion of the first outer frame line. The step of modifying the first outer frame line to the second outer frame line of the target element according to the conversion parameters between the element to be adjusted and the target element includes: If the element to be adjusted is a solid lane line and the target element is a dashed lane line, the first outer frame of the element to be adjusted is converted from the solid lane line to the dashed lane line corresponding to the second outer frame of the target element based on a first conversion parameter. The first conversion parameter includes the line type of the dashed lane line and the discontinuity dimension parameter of the dashed lane line; or... When the element to be adjusted is a solid lane line and the target element is a herringbone line, a new outer frame area corresponding to the target element is added outside the first outer frame of the element to be adjusted based on the fourth conversion parameter to obtain the second outer frame of the target element. The fourth conversion parameter includes the relative position parameter between the outer frame area corresponding to the herringbone line and the solid lane line, and the size parameter corresponding to the herringbone line. The data processing module is used to generate a second environmental image based on the second bird's-eye view feature map, and use the second environmental image as a sample image of the first model.
11. An electronic device, characterized in that, include: One or more processors; One or more memories; the one or more memories storing one or more programs that, when executed by the one or more processors, cause the electronic device to perform the model training method of any one of claims 1-9.
12. A computer-readable medium, characterized in that, The computer-readable medium stores instructions that, when executed on a machine, cause the machine to perform the model training method of any one of claims 1-9.
13. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the model training method according to any one of claims 1-9.
14. A vehicle, characterized in that, Includes the electronic device as described in claim 11.
Citation Information
Patent Citations
Image processing method and device, vehicle, medium and chip
CN115205311A