Road element recognition method and device, terminal equipment and storage medium
By splicing the images collected by the camera device under multiple visions into a stitched image and converting them into target top view features, the problem of low efficiency and low accuracy of vehicle identification of road elements is solved, and more efficient and accurate recognition is achieved.
Patent Information
- Application Number
- CN202311836519.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the vehicle has low efficiency and low accuracy in identifying road elements during driving, mainly due to feature characterization errors and information omissions caused by multiple independent extraction of image features.
The images captured by the camera device under multiple visions are stitched into a stitched image, the image features at multiple scales of the stitched image are determined, and converted into target top view features to identify road elements.
By reducing the number of feature extractions, the efficiency of road element recognition is improved, and feature representation errors are avoided, the representation effect of spatial position information is enhanced, and the recognition accuracy is improved.
Smart Images

Figure CN120236258A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of vehicles, and particularly relates to a method, device, terminal device, and storage medium for road element recognition. Background Art
[0002] During the driving process of a vehicle, it is usually necessary to collect surrounding images along the horizontal direction based on camera devices under different visions to perceive road elements in the surrounding environment.
[0003] Specifically, for each surrounding image, feature extraction is independently performed to obtain the image features corresponding to each surrounding image. Then, multiple image features are fused to identify road elements.
[0004] However, before recognition, it is necessary to perform feature extraction on the surrounding images multiple times, resulting in low recognition efficiency. Moreover, when independently extracting image features multiple times, it is easy for the image features corresponding to some road element information in the image to be misrepresented, leading to low accuracy in recognition based on the fused image features. Summary of the Invention
[0005] Embodiments of this application provide a method, device, terminal device, and storage medium for road element recognition, which can solve the problems of low efficiency and low accuracy in road element recognition.
[0006] In a first aspect, embodiments of this application provide a method for road element recognition, which includes:
[0007] Stitch the images containing roads collected by the camera device under multiple visions into a stitched image; the shooting direction of the camera device is horizontal;
[0008] Determine the image features at multiple scales of the stitched image;
[0009] Convert the multiple image features into target top-view features respectively;
[0010] Recognize road elements on the road based on the target top-view features.
[0011] In a second aspect, embodiments of this application provide a device for road element recognition, which includes:
[0012] A stitching module, configured to stitch the images containing roads collected by the camera device under multiple visions into a stitched image; the shooting direction of the camera device is horizontal;
[0013] A determination module, configured to determine the image features at multiple scales of the stitched image;
[0014] A conversion module, configured to convert the multiple image features into target top-view features respectively;
[0015] An identification module, configured to identify road elements on a road based on target top - view features.
[0016] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to the first aspect above is implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a computer - readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to the first aspect above is implemented.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is enabled to execute the method according to the first aspect above.
[0019] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: Since the images are images containing roads collected by a camera device from multiple views, when multiple images are stitched into a stitched image, the stitched image can accurately present the global information of each road element in the surrounding environment. Then, the terminal device only needs to extract the image features of the stitched image. For example, the terminal device can extract the image features of the stitched image at multiple scales, and convert the image features at multiple scales to obtain target top - view features. In this way, it is not necessary to extract image features from the surrounding images multiple times, improving the recognition efficiency. And, since the stitched image contains the global information of the road elements, it is possible to avoid the situation where the image features corresponding to some image information are misrepresented when independently extracting image features. Furthermore, the terminal device can improve the accuracy of subsequent recognition based on the fused target top - view features. Also, the camera device usually collects images from multiple views along the horizontal direction, and the image features extracted based on these images are more inclined to describe the projection information of road elements in the perspective view. When representing the spatial position information of road elements, some detailed information is likely to be omitted. As a result, the fused features directly obtained by fusing image features have a poor effect when representing the spatial position information of road elements. Based on this, the terminal device can convert multiple image features into target top - view features for road element recognition to improve the representation effect of the spatial position information of road elements. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of the implementation of a road element recognition method provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic diagram of an implementation method for obtaining a spliced image in a road element recognition method provided by an embodiment of the present application;
[0023] Figure 3 It is an initial image under the rear vision in a road element recognition method provided by an embodiment of the present application;
[0024] Figure 4 It is the initial images under each vision in a road element recognition method provided by an embodiment of the present application;
[0025] Figure 5 It is a spliced image in a road element recognition method provided by an embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of an implementation method for obtaining the target top view feature in a road element recognition method provided by an embodiment of the present application;
[0027] Figure 7 It is a schematic structural diagram of a road element recognition device provided by an embodiment of the present application;
[0028] Figure 8 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0029] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0030] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0031] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance.
[0032] During the driving process of a vehicle, it is usually necessary to collect surrounding images horizontally based on camera devices under different visions to perceive road elements in the surrounding environment.
[0033] Specifically, one or more camera devices are usually provided on a vehicle under multiple visions such as the front side, the left front side, the left rear side, the right front side, the right rear side, and the rear side to collect surrounding images. Then, for each surrounding image, feature extraction is independently performed to obtain the image features corresponding to each surrounding image. Then, multiple image features are fused to identify road elements.
[0034] Exemplarily, a feature extractor (backbone) is usually provided on a terminal device of a vehicle to extract the image features (2D features) of surrounding images.
[0035] However, before recognition, the terminal device needs to perform feature extraction on the surrounding images multiple times, resulting in low recognition efficiency. Moreover, when independently extracting image features multiple times, it is easy for the image features corresponding to some road elements in the image to be misrepresented, leading to low accuracy in recognition based on the fused image features.
[0036] Specifically, in an actual application scenario, if the coverage area of a road element in a certain surrounding image is small, the feature information of the road element in the surrounding image is usually difficult to be accurately represented. Furthermore, the road element cannot be jointly represented with the same road elements in other surrounding images. At this time, when fusing multiple image features, the incorrect image features of the road element will also be transmitted to the fused features. Furthermore, the recognition accuracy of the road element is reduced.
[0037] Based on this, in order to improve the efficiency and accuracy of road element recognition, an embodiment of this application provides a road element recognition method, which can be applied to terminal devices such as in-vehicle terminals or intelligent driving controllers. The specific type of the terminal device is not limited in the embodiment of this application.
[0038] Please refer to Figure 1 , Figure 1The figure shows a flowchart of the implementation of a road element recognition method provided by an embodiment of the present application. The method includes the following steps:
[0039] S101. Stitch the images containing roads collected by the camera device under multiple perspectives into a stitched image; the shooting direction of the camera device is the horizontal direction.
[0040] In one embodiment, the above perspective has been explained and will not be elaborated here. Among them, the number of the above camera devices can be one or multiple. When the number is one, the terminal device can control the camera device to rotate to capture images under different perspectives. When the number is multiple, one or more camera devices can be respectively arranged corresponding to each perspective to capture images.
[0041] In one embodiment, the above road is usually a road with road elements. Among them, road elements include but are not limited to lane lines, curbs, stop lines, crosswalks, and road surface arrows, etc., and are not limited thereto.
[0042] It should be noted that a navigation map module is usually set on the terminal device to be used for path planning and navigation. Among them, the navigation map module usually also includes road elements of each road. At this time, the terminal device can first determine whether there are road elements on the driving road based on the position information of the vehicle and the navigation map module. If it is determined that there are road elements on the driving road, the terminal device can execute the above road element recognition method. Otherwise, the terminal device can stop executing the road element recognition method. Furthermore, the terminal device does not need to always execute the road element recognition method, reducing the power consumption of the terminal device.
[0043] In one embodiment, the terminal device can stitch multiple images into a stitched image according to a preset stitching order. For example, the images corresponding to the perspectives of the left front side and the right front side can be preset to be arranged above, the images corresponding to the front side and the rear side perspectives can be arranged in the middle, and the images corresponding to the left rear side and the right rear side perspectives can be respectively arranged below to form a stitched image.
[0044] It should be noted that since the shooting direction of the camera device is the horizontal direction when shooting images, and road elements are usually distributed on the ground. Therefore, in the captured images, usually mainly the lower part of the image contains road elements.
[0045] Based on this, in order to increase the proportion of the coverage of road elements in the stitched image, so that the terminal device can extract more detailed feature information of road elements from the stitched image, the terminal device can perform stitching on multiple images according to the steps S201 - S203 as Figure 2 shown to obtain a stitched image. Details are as follows:
[0046] S201. Obtain the initial images containing roads collected by the imaging device under multiple visions.
[0047] S202. For any one of the initial images, according to the preset height cropping ratio, crop the lower part of the initial image as the target cropped image; the height direction of the initial image points from the lower part to the upper part of the initial image.
[0048] In one embodiment, the above-mentioned initial image is the original image captured by the imaging device. The above-mentioned preset height cropping ratio can be set according to the actual situation, and there is no limitation thereto. Exemplarily, the above-mentioned height cropping ratio can be 2 / 3.
[0049] It should be noted that the height of the target cropped image is the height of the initial image multiplied by the height cropping ratio, while the width in the target cropped image is the same as the width of the initial image. That is, at the position where the height of the initial image is multiplied by the height cropping ratio, the initial image is cropped horizontally.
[0050] It should be added that the height direction of the initial image points from the lower part to the upper part of the initial image, and generally only the lower part of the image contains road elements. Therefore, the terminal device can use the lower part of the cropped initial image as the target cropped image.
[0051] In one embodiment, the above-mentioned cropping method includes but is not limited to cropping horizontally along a preset cropping program, or cropping the lower part of the initial image using a cropping frame that meets the above-mentioned preset cropping ratio, and there is no limitation thereto.
[0052] Refer to Figure 3 , Figure 3 is the initial image under the rear vision in a road element recognition method provided by an embodiment of the present application. Among them, Figure 3 The line segment L1 in represents the cropping direction (the direction of horizontal cropping along the width of the initial image). Based on Figure 3 it can be known that the height direction of the initial image is L2 (pointing from the lower part to the upper part of the initial image). Therefore, it can be considered that the image corresponding to D is the target cropped image (the lower part of the initial image).
[0053] S203. Stitch multiple target cropped images to obtain a stitched image.
[0054] In one embodiment, when generating the stitched image, the target cropped images corresponding to the front and rear visions can be set in the middle, the target cropped images corresponding to the left front and right front visions can be set at the upper end, and the images corresponding to the left rear and right rear visions can be set at the lower end to form a stitched image.
[0055] It should be specifically noted that since the driving direction of the vehicle is usually forward or backward, it can be considered that the initial images corresponding to the front and rear views mainly include the road. That is, the target cropped images corresponding to the front and rear views mainly contain road elements. Based on this, in order to further increase the proportion of the coverage of road elements in the stitched image, the terminal device can also adjust the resolution of each target cropped image.
[0056] As an example, for any target cropped image, the terminal device can determine the target resolution of the target cropped image according to the view corresponding to the target cropped image, and adjust the target cropped image according to the target resolution. Then, multiple adjusted target cropped images are stitched to obtain a stitched image.
[0057] Among them, the terminal device can pre-store the correspondence between each view and the resolution. Then, based on the view corresponding to the target cropped image, the target resolution of the target cropped image is determined.
[0058] Exemplarily, it can be pre-set that the target resolutions corresponding to the views of the left front side, the right front side, the left rear side, and the right rear side are equal, and the target resolutions corresponding to the front view and the rear view are equal. And, the target resolutions corresponding to the front view and the rear view are greater than the target resolutions corresponding to the views of the left front side, the right front side, the left rear side, and the right rear side. Thus, the proportion of the coverage of road elements in the stitched image can be increased.
[0059] Refer to Figure 4 and Figure 5 , Figure 4 are the initial images under each view in a road element recognition method provided by an embodiment of the present application. Figure 5 is the stitched image in a road element recognition method provided by an embodiment of the present application. From Figure 4 it can be determined that the resolutions of the initial images under each view are the same (the image sizes are the same). In Figure 5 , taking the line segment L3 as the dividing line, the image on the left side of L3 (L left) is the target cropped image corresponding to the left front view; the image on the right side of L3 (L right) is the target cropped image corresponding to the right front view; the target cropped image corresponding to the front view is located below the target cropped images corresponding to the left front view and the right front view respectively; the target cropped image corresponding to the rear view is located below the target cropped image corresponding to the front view; taking the line segment L4 as the dividing line, the view on the left side of L4 (L left) is the target cropped image corresponding to the left rear view; the view on the right side of L4 (L right) is the target cropped image corresponding to the right rear view. From Figure 5It can be known that when the image size of the stitched image is preset, the target cropped images corresponding to the front vision and the rear vision can be enlarged according to the target resolutions corresponding to the front vision and the rear vision; and, according to the target resolutions corresponding to the front vision and the rear vision, the target cropped images corresponding to the left front, right front, left rear, and right rear visions can be reduced.
[0060] S102. Determine the image features of the stitched image at multiple scales.
[0061] In one embodiment, the above image features are obtained by processing the pixels in the stitched image through a convolutional network. Among them, the image features can usually be represented by feature maps. And, each feature value in the feature map is used to represent the response of the strength of each feature in the stitched image. Among them, the image features can be one or more features representing the texture information, color information, and shape information of the image, and this is not limited.
[0062] As an example, the terminal device can use the backbone to extract the image features of the stitched image. Further, in order to obtain image features with stronger expressive power, after using the backbone to extract the initial image features of the stitched image. Then, the initial image features are input into a preset Feature Pyramid Networks (FPN) for feature processing to obtain image features at multiple scales.
[0063] Among them, the FPN can use shallow features (for example, single-layer scale features) to distinguish simple targets, and use deep features (for example, multi-layer scale features) to distinguish complex targets. Furthermore, the FPN can not only focus on the detailed information of the target based on the shallow features, but also focus on the semantic information of the target based on the deep features, so that the obtained image features can further represent the feature information of the stitched image.
[0064] S103. Convert the multiple image features into target top-view features respectively.
[0065] In one embodiment, the above target top-view feature is a BEV (Bird's-Eye View) feature. It should be noted that the imaging device usually collects images at multiple perspectives along the horizontal direction, and the image features extracted based on this image are more inclined to describe the projection information of the real road element world in the perspective view. However, when representing the spatial position information of road elements, this image feature is likely to miss some detailed information. Furthermore, the fusion feature directly obtained by fusing based on the image features has a poor effect when representing the spatial position information of road elements.
[0066] Based on this, the terminal device can convert multiple image features into target top-view features for road element recognition, so as to improve the representation effect of the spatial position information of road elements.
[0067] In one embodiment, the terminal device can separately convert the image features corresponding to each layer scale to obtain the top-view features corresponding to each image feature respectively. Then, fuse the multiple top-view features to obtain the target top-view feature. Alternatively, the multiple image features can be first fused, and then the fused image features are converted into the target top-view feature, and there is no limitation on this.
[0068] However, it should be noted that when fusing multiple image features first, since the data spaces (visions) of each camera device for collecting images are different, more post-processing rules are required to associate the perception results (image features) of each camera device for feature fusion. However, this method is not only relatively complex in operation, but also the effect of feature fusion is average.
[0069] Based on this, the terminal device can first separately convert the image features to obtain the top-view features corresponding to each image feature respectively. At this time, the data space where the road elements represented by each top-view feature are located is the data space under the BEV perspective. Furthermore, when fusing in the data space corresponding to the BEV perspective, the effect of feature fusion is better.
[0070] In another embodiment, the terminal device can obtain the target top-view feature through voxels, or VPN (Virtual Private Network), or Perspective View (PV). Exemplarily, voxels are used to discretize the 3D space to construct a regular structure for feature transformation, so as to provide a more effective representation for 3D scene understanding.
[0071] However, by generating the target top-view feature through the above method, although the structural information of large-scale scenes in the stitched image can be effectively covered, there may be a loss of local spatial accuracy.
[0072] Based on this, in order to obtain a target top-view feature with stronger spatial accuracy representation ability, the terminal device can process the image features according to the steps S601-S605 as Figure 6 shown. Details are as follows:
[0073] S601: For any one image feature, perform convolution on the image feature using a preset first convolution module to obtain a first target feature.
[0074] In one embodiment, the convolution stride and convolution kernel of the above first convolution module can be set in advance, and there is no limitation thereto. Exemplarily, the convolution stride can be 2, and the convolution kernel can be a 3*3 convolution kernel.
[0075] Among them, the above first convolution module is used to perform downsampling processing on the image features to reduce the number of feature values included in the feature map and enable each feature value to represent more image information. Exemplarily, the above first convolution module can use a max pooling layer to perform downsampling convolution processing on the image features to obtain a convolution feature map. Then, based on the convolution feature map, a first target feature is determined.
[0076] In one embodiment, the first target feature can be represented by a 2D matrix of [H1, W1]. Wherein, H1 and W1 are respectively the height and width of the convolution feature map, and each value in the 2D matrix is a feature value in the convolution feature map.
[0077] In another embodiment, the first target feature can also be represented by a 4D matrix of [B, C1, H1, W1]. Wherein, B is the number of image features processed in each batch, and C1 is the channel corresponding to the image features.
[0078] S602. Establish image position features corresponding to the image features.
[0079] In one embodiment, as described in the above S102 step, the feature map includes multiple values, and each feature value in the feature map is used to represent the response of the strength of each feature in the spliced image. Based on this, the terminal device can also determine the position information of each feature corresponding to each feature value in the spliced image at the same time, so as to establish image position features corresponding to the entire image features based on the respective position information. It should be noted that the dimension of the image position features should be consistent with the dimension of the image features.
[0080] S603. Generate a second target feature based on the image features and the image position features.
[0081] In one embodiment, since the image features and the image position features are of the same dimension, therefore, a corresponding image feature matrix can be generated based on the image features, and an image position feature matrix can be generated based on the image position features. Then, the image feature matrix and the image position feature matrix are subjected to matrix addition or multiplication operations to obtain a second target feature matrix. At this time, the second target feature matrix can be used to represent the second target feature.
[0082] As an example, the terminal device can also first fuse the image features and the image position features to obtain a fused feature. Then, a preset second convolution module is used to perform convolution on the fused feature to obtain a second target feature.
[0083] Among them, the process of the second convolution module convolving the fused features is similar to the process of the first convolution module convolving the image features, and thus will not be described in detail herein. In addition, the fused features can be obtained by adding or multiplying the image feature matrix corresponding to the image features and the image position feature matrix corresponding to the image position features.
[0084] It should be noted that in the feature map corresponding to the second target feature, each eigenvalue can simultaneously represent the position information and the feature information of the feature at this time.
[0085] It should be noted that the first convolution module and the second convolution module can be different convolution modules. For example, in the first convolution module and the second convolution module, at least one of the convolution stride and the convolution kernel is different.
[0086] In this embodiment, the purpose of using different second convolution modules to convolve the fused features is as follows: the second target feature different from the first target feature in source can be obtained from the image features corresponding to the stitched image, so as to enrich the feature information of the stitched image and represent the feature information of interest in the image from a global perspective.
[0087] In one embodiment, the matrix corresponding to the second target feature should be similar in form to the matrix corresponding to the first target feature, and both are represented by a 2D matrix of [H1, W1] or a 4D matrix of [B, C1, H1, W1]. Therefore, it can be considered that each eigenvalue in the first target feature corresponds one-to-one to each eigenvalue in the second target feature.
[0088] S604. According to the preset top view position features, convert the first target feature and the second target feature to obtain the top view features corresponding to the image features.
[0089] In one embodiment, the top view position features can be represented by a 2D matrix of [H1, W1], or by a 4D matrix of [B, C2, H2, W2]. Moreover, C2, H2, and W2 can be the same as or different from C1, H1, and W1, and this is not limited herein.
[0090] Among them, each eigenvalue in the top view position features can be set in advance, and this is not limited herein. It should be noted that each eigenvalue at this time is only used to represent the spatial position relationship from the perspective of the top view.
[0091] Exemplarily, the above top view position features can be a preset space with a size of C2 * H2 * W2. Among them, H2 and W2 are the length and width of the spatial dimensions of the BEV plane (the size of the feature map corresponding to the top view position features), and C is the height perpendicular to this plane, which is used to correspond to the channels of the first target feature and the second target feature during the conversion process.
[0092] Based on this, it can be considered that converting the second target feature and the first target feature into the target top view feature is to find the feature position mapping relationship from the feature map to the BEV space, so that the semantic information represented by each feature value of the feature map (the feature maps corresponding to the first target feature and the second target feature respectively) can be completely retained in the BEV space.
[0093] In a specific embodiment, the terminal device may first determine the feature position mapping relationship between the second target feature and the target top view feature according to the second target feature and the top view position feature. Then, based on the feature position mapping relationship, the first target feature is converted to obtain the top view feature.
[0094] Among them, in the feature map corresponding to the second target feature, each feature value can simultaneously represent the position information and feature information of the corresponding feature. Therefore, the terminal device can determine the feature position mapping relationship between the second target feature and the target top view feature according to the image position feature in the second target feature.
[0095] In a specific embodiment, the matrix obtained by multiplying the matrix corresponding to the target top view feature by the matrix corresponding to the second target feature can be determined as the above feature position mapping relationship. That is, the feature position mapping relationship is Y = K T *Q.
[0096] Among them, Y is the transformation matrix, K is the second target feature, Q is the top view position feature, and T is a preset transpose matrix.
[0097] It should be noted that since the second target feature is obtained based on the image feature and the image position feature, in the obtained transformation matrix Y, it not only contains the feature position mapping relationship, but also includes the feature information corresponding to the second target feature.
[0098] It can be understood that each feature value in the above first target feature corresponds one-to-one with each feature value in the second target feature. Therefore, when determining the above feature position mapping relationship, each feature value in the first target feature can be fused with the feature value corresponding to the transformation matrix to obtain the final top view feature. At this time, each feature value included in the top view feature is formed by the first target feature and the second target feature from different sources, enriching the global feature information of the stitched image in the BVE space.
[0099] Among them, the fusion method can be the matrix addition or multiplication operations described above, and will not be described again.
[0100] In a specific embodiment, the terminal device may input the top view position feature, the second target feature, and the first target feature into a preset feature calculation formula to obtain the top view feature. The feature calculation formula is as follows:
[0101]
[0102] Among them, out QKV is the top view feature, V is the first target feature, K is the second target feature, Q is the top view position feature, T is a preset transpose matrix, d is a preset value, and dim is the dimension for softmax. It can be understood that K T *Q is the conversion matrix Y described above.
[0103] In another embodiment, the top view position feature, the second target feature, and the first target feature may also be input into the following calculation formula to obtain the top view feature. The calculation formula is as follows:
[0104]
[0105] However, it should be noted that compared with the preset feature calculation formula, although the above calculation formula can obtain the top view feature with the same effect, it requires more matrix operation operations and the processing time to obtain the top view feature is longer.
[0106] S605. Fuse multiple top view features to obtain the target top view feature.
[0107] In one embodiment, the above top view feature fusion method may be the matrix addition or multiplication operation described above, and no further description will be given here.
[0108] Among them, fusing multiple top view features to obtain the target top view feature can enable the target top view feature to represent the semantic information of the road elements in the stitching image in the global scene.
[0109] S104. Identify road elements on the road based on the target top view feature.
[0110] In one embodiment, the target top view feature may be input into a preset perception module to identify the above road elements. Exemplarily, the perception module includes but is not limited to a lane line perception module, a road edge perception module, a stop line perception module, a crosswalk perception module, and a road surface arrow perception module, etc., and is not limited thereto.
[0111] In this embodiment, since the images are images containing roads collected by the imaging device from multiple perspectives, when stitching multiple images into a stitched image, the stitched image can accurately present the global information of each road element in the surrounding environment. Then, the terminal device only needs to extract the image features of the stitched image. For example, the terminal device can extract the image features of the stitched image at multiple scales to obtain the target top view features based on the image features at multiple scales. In this way, it is not necessary to extract image features from the surrounding images multiple times, improving the recognition efficiency. Moreover, since the stitched image contains the global information of the road elements, the situation where the image features corresponding to some image information are misrepresented when independently extracting image features can be avoided. Furthermore, the terminal device can improve the accuracy of subsequent recognition based on the fused target top view features. In addition, the imaging device usually collects images from multiple perspectives in the horizontal direction, and the image features extracted based on these images are more inclined to describe the projection information of road elements in the perspective view. When representing the spatial position information of road elements, some detailed information is likely to be missed. As a result, the fused features directly obtained by fusing the image features have a poor effect when representing the spatial position information of road elements. Based on this, the terminal device can convert multiple image features into target top view features for road element recognition to improve the representation effect of the spatial position information of road elements.
[0112] Please refer to Figure 7 , Figure 7 which is a structural block diagram of a road element recognition device provided in an embodiment of the present application. In this embodiment, each module included in the road element recognition device is used to execute Figure 1 , Figure 2 and Figure 6 the respective steps in the corresponding embodiments. Specifically, please refer to Figure 1 , Figure 2 and Figure 6 and Figure 1 , Figure 2 and Figure 6 the relevant descriptions in the corresponding embodiments corresponding to Figure 7 . For the sake of convenience, only the parts related to this embodiment are shown. Referring to Figure 7 , the road element recognition device 700 may include: a stitching module 710, a determination module 720, a conversion module 730, and an identification module 740, where:
[0113] The stitching module 710 is configured to stitch the images containing roads collected by the imaging device from multiple perspectives into a stitched image; the shooting direction of the imaging device is the horizontal direction.
[0114] The determination module 720 is configured to determine the image features of the stitched image at multiple scales.
[0115] A conversion module 730 for converting multiple image features into target top - view features respectively.
[0116] An identification module 740 for identifying road elements on a road based on the target top - view features.
[0117] In one embodiment, the splicing module 710 is further configured to:
[0118] Obtain initial images containing roads collected by a camera device under multiple visions; for any one of the initial images, crop the lower part of the initial image as a target cropped image according to a preset height cropping ratio; the height direction of the initial image points from the lower part to the upper part of the initial image; splice multiple target cropped images to obtain a spliced image.
[0119] In one embodiment, the splicing module 710 is further configured to:
[0120] For any one of the target cropped images, determine the target resolution of the target cropped image according to the vision corresponding to the target cropped image; adjust the target cropped image according to the target resolution; splice multiple adjusted target cropped images to obtain a spliced image.
[0121] In one embodiment, the conversion module 730 is further configured to:
[0122] For any one of the image features, perform convolution on the image feature using a preset first convolution module to obtain a first target feature; establish an image position feature corresponding to the image feature; generate a second target feature based on the image feature and the image position feature; perform conversion on the first target feature and the second target feature according to a preset top - view position feature to obtain a top - view feature corresponding to the image feature; fuse multiple top - view features to obtain a target top - view feature.
[0123] In one embodiment, the conversion module 730 is further configured to:
[0124] Fuse the image feature and the image position feature to obtain a fused feature; perform convolution on the fused feature using a preset second convolution module to obtain a second target feature; the first convolution module and the second convolution module are different convolution modules.
[0125] In one embodiment, the conversion module 730 is further configured to:
[0126] Determine a feature position mapping relationship between the second target feature and the target top - view feature according to the second target feature and the top - view position feature; perform conversion on the first target feature based on the feature position mapping relationship to obtain a top - view feature.
[0127] In one embodiment, the conversion module 730 is further configured to:
[0128] Input the top view position feature, the second target feature, and the first target feature into a preset feature calculation formula to obtain the top view feature. The feature calculation formula is as follows:
[0129]
[0130] where out QKV is the top view feature, V is the first target feature, K is the second target feature, Q is the top view position feature, T is a preset transpose matrix, d is a preset value, and dim is the dimension for softmax.
[0131] It should be understood that Figure 7 in the structural block diagram of the road element recognition device shown, each module is used to execute Figure 1 、 Figure 2 and Figure 6 the respective steps in the corresponding embodiments. For Figure 1 、 Figure 2 and Figure 6 the respective steps in the corresponding embodiments have been explained in detail in the above embodiments. For details, please refer to Figure 1 、 Figure 2 and Figure 6 and Figure 1 、 Figure 2 and Figure 6 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.
[0132] Figure 8 is the structural block diagram of a terminal device provided by an embodiment of the present application. As Figure 8 shown, the terminal device 800 in this embodiment includes: a processor 810, a memory 820, and a computer program 830 stored in the memory 820 and executable on the processor 810, such as a program for the road element recognition method. When the processor 810 executes the computer program 830, it implements the steps in the respective embodiments of the above-mentioned various road element recognition methods, such as Figure 1 the S101 to S104 shown. Alternatively, when the processor 810 executes the computer program 830, it implements the functions of each module in the corresponding embodiments of the above Figure 7 For example, Figure 7 the functions of each module shown. For details, please refer to Figure 7 the relevant descriptions in the corresponding embodiments.
[0133] Exemplarily, the computer program 830 may be divided into one or more modules. One or more modules are stored in the memory 820 and executed by the processor 810 to implement the road element recognition method provided by the embodiments of the present application. One or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 830 in the terminal device 800. For example, the computer program 830 may implement the road element recognition method provided by the embodiments of the present application.
[0134] The terminal device 800 may include, but is not limited to, a processor 810 and a memory 820. Those skilled in the art can understand that Figure 8 merely examples of the terminal device 800, which do not constitute a limitation on the terminal device 800, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, buses, etc.
[0135] The so-called processor 810 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0136] The memory 820 may be an internal storage unit of the terminal device 800, such as the hard disk or memory of the terminal device 800. The memory 820 may also be an external storage device of the terminal device 800, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the terminal device 800. Further, the memory 820 may also include both the internal storage unit and the external storage device of the terminal device 800.
[0137] The embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by the processor to implement the road element recognition method in the above various embodiments.
[0138] The embodiments of the present application provide a computer program product. When the computer program product runs on the terminal device, the terminal device is enabled to execute the road element recognition method in the above various embodiments.
[0139] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for identifying road elements, characterized in that, The method includes: Stitching the images containing roads collected by the imaging device under multiple visions into a stitched image; the shooting direction of the imaging device is the horizontal direction; Determining the image features of the stitched image at multiple scales; Converting the multiple image features into target top-view features respectively; Identifying the road elements on the road based on the target top-view features.
2. The method according to claim 1, characterized in that, The step of stitching the images containing roads collected by the imaging device under multiple visions into a stitched image includes: Obtaining the initial images containing roads collected by the imaging device under the multiple visions; For any one of the initial images, cropping the lower part of the initial image as a target cropped image according to a preset height cropping ratio; the height direction of the initial image points from the lower part to the upper part of the initial image; Stitching the multiple target cropped images to obtain the stitched image.
3. The method according to claim 2, wherein The step of stitching the multiple target cropped images to obtain the stitched image includes: For any one of the target cropped images, determining the target resolution of the target cropped image according to the vision corresponding to the target cropped image; Adjusting the target cropped image according to the target resolution; Stitching the multiple adjusted target cropped images to obtain the stitched image.
4. The method according to any one of claims 1 to 3, characterized in that The step of converting the multiple image features into target top-view features respectively includes: For any one of the image features, performing convolution on the image feature by using a preset first convolution module to obtain a first target feature; Establishing an image position feature corresponding to the image feature; Generating a second target feature based on the image feature and the image position feature; Converting the first target feature and the second target feature according to a preset top-view position feature to obtain the top-view feature corresponding to the image feature; Fusing the multiple top-view features to obtain the target top-view feature.
5. The method according to claim 4, characterized in that The step of generating a second target feature based on the image feature and the image position feature includes: Fusing the image feature and the image position feature to obtain a fused feature; Performing convolution on the fused feature by using a preset second convolution module to obtain the second target feature; the first convolution module and the second convolution module are different convolution modules.
6. The method according to claim 4, characterized in that, The step of converting the first target feature and the second target feature according to a preset top-view position feature to obtain the top-view feature corresponding to the image feature includes: Determining a feature position mapping relationship between the second target feature and the target top-view feature according to the second target feature and the top-view position feature; Converting the first target feature based on the feature position mapping relationship to obtain the top-view feature.
7. The method according to claim 4, characterized in that, The step of converting the first target feature and the second target feature according to a preset top-view position feature to obtain the top-view feature corresponding to the image feature includes: Inputting the top-view position feature, the second target feature, and the first target feature into a preset feature calculation formula to obtain the top-view feature; the feature calculation formula is as follows: Among them, out QKV is the top view feature, V is the first target feature, K is the second target feature, Q is the top view position feature, T is a preset transposed matrix, d is a preset value, and dim is the dimension for softmax.
8. A road element recognition device, characterized in that, The device includes: A stitching module, configured to stitch images containing roads collected by a camera device under multiple visions into a stitched image; the shooting direction of the camera device is horizontal; A determination module, configured to determine image features at multiple scales of the stitched image; A conversion module, configured to convert the multiple image features into target top-view features respectively; An identification module, configured to identify road elements on the road based on the target top-view features.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.