Road information recognition method, device and storage medium
By integrating the road segmentation detection results of multiple fisheye diagrams and top view, the problems of missing segmentation and miss segmentation in road information identification in the prior art are solved, and higher detection accuracy and stability are achieved.
Patent Information
- Application Number
- CN202111272784.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The prior art has missed segmentation and missed segmentation in road information identification, resulting in inaccurate road information.
By acquiring the road segmentation detection result diagram of multiple fisheye diagrams and corresponding top view views, the fusion process is performed to generate the fusion detection result diagram, thereby improving the accuracy of the road information.
By integrating the detection results of fisheye diagram and top view, the phenomenon of missing segmentation and missegment is reduced, and the accuracy and robustness of road information detection are improved.
Smart Images

Figure CN114120254B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of road traffic, and particularly to a road information recognition method, apparatus, and storage medium. Background Art
[0002] In the field of intelligent driving, the safe driving of vehicles depends on accurate road information. Based on this, the recognition of road information such as lane lines, road signs, drivable areas, curbs, and the road scene where the vehicle is currently located in the road becomes particularly important.
[0003] In the related art, first, an in-vehicle camera is used to collect images to obtain a fisheye image, and then the fisheye image is subjected to image semantic segmentation to obtain road information, or the fisheye image is converted into a top view, and then the top view is subjected to image semantic segmentation to obtain road information.
[0004] However, when segmenting the fisheye image, there will be a certain degree of missed segmentation or missegmentation of road elements with fixed structural information such as lane lines and road markings, and when segmenting the top view, there will be more missed segmentation or missegmentation of road elements with height information such as railings and walls. Therefore, no matter which of the above methods is used for road information detection, there will be phenomena of missed segmentation and missegmentation, resulting in inaccurate road information finally obtained. Summary of the Invention
[0005] Embodiments of this application provide a road information recognition method, apparatus, and storage medium, which can improve the problem of the accuracy of the obtained road information. The technical solutions are as follows:
[0006] On the one hand, a road information recognition method is provided, and the method includes:
[0007] Obtain a plurality of fisheye images collected by a target vehicle, and obtain a first top view corresponding to the plurality of fisheye images, where the plurality of fisheye images are images collected from multiple perspectives;
[0008] Obtain a road segmentation detection result map of each of the plurality of fisheye images and a road segmentation detection result map of the first top view;
[0009] Fuse the road segmentation detection result maps of the plurality of fisheye images and the road segmentation detection result map of the first top view to obtain a fused detection result map;
[0010] Obtain road information according to the fused detection result map.
[0011] Optionally, the obtaining the road segmentation detection result map of each of the plurality of fisheye images includes:
[0012] Perform image semantic segmentation on each of the multiple fisheye images to obtain a segmentation result image for each fisheye image;
[0013] Perform object detection on each fisheye image to obtain an object detection result for each fisheye image, where the object detection result is used to indicate whether the corresponding image contains an object detection box;
[0014] Generate a road segmentation detection result image for the corresponding fisheye image according to the segmentation result image and the object detection result of each fisheye image.
[0015] Optionally, the obtaining of the first top view corresponding to the multiple fisheye images includes:
[0016] Convert the multiple fisheye images to the top view coordinate system through coordinate transformation to obtain the first top view.
[0017] Optionally, the fusing of the road segmentation detection result images of the multiple fisheye images and the road segmentation detection result image of the first top view to obtain a fused detection result image includes:
[0018] Convert the road segmentation detection result images of the multiple fisheye images to the top view coordinate system through coordinate transformation to obtain a third top view;
[0019] Fuse the third top view and the road segmentation detection result image of the first top view to obtain a fused detection result image.
[0020] Optionally, the converting of the road segmentation detection result images of the multiple fisheye images to the top view coordinate system through coordinate transformation to obtain a third top view includes:
[0021] Convert the road segmentation detection result image of each fisheye image to the top view coordinate system to obtain a top view sub-image corresponding to each fisheye image;
[0022] Stitch the top view sub-images corresponding to the multiple fisheye images to obtain a stitched top view;
[0023] If the road segmentation detection result images of the multiple fisheye images also include object detection boxes, convert the object detection boxes to the stitched top view to obtain the third top view.
[0024] Optionally, the road segmentation detection result image includes the class attribute of each pixel point, and the stitching of the top view sub-images corresponding to the multiple fisheye images to obtain a stitched top view includes:
[0025] If the first region in the first top-down sub-map and the second region in the second top-down sub-map are overlapping regions, when the class attributes of two pixel points at the same position within the first region and the second region are different, determine the pixel point with the highest class attribute priority from the two pixel points;
[0026] Use the class attribute of the determined pixel point as the class attribute of the pixel point at the corresponding position in the stitched top-down view.
[0027] Optionally, the road segmentation detection result map includes the class attribute of each pixel point. The process of fusing the road segmentation detection result maps of the third top-down view and the first top-down view to obtain the fused detection result map includes:
[0028] If the distance represented by each pixel point in the third top-down view is different from the distance represented by each pixel point in the road segmentation detection result map of the first top-down view, convert the pixel points in the third top-down view so that the distance represented by each pixel point in the converted third top-down view is the same as the distance represented by each pixel point in the road segmentation detection result map of the first top-down view;
[0029] For multiple third pixel points in the converted third top-down view that have corresponding pixel points in the road segmentation detection result map of the first top-down view, use the class attribute with the highest priority among the class attributes of each third pixel point and the corresponding pixel point as the class attribute of the corresponding pixel point;
[0030] If the converted third top-down view also includes target detection frames, fuse the target detection frames in the converted third top-down view into the road segmentation detection result map of the first top-down view to obtain the fused detection result map.
[0031] Optionally, the fused detection result map includes the class attribute of each pixel point and target detection frames. The process of obtaining road information based on the fused detection result map includes:
[0032] Identify the road elements included in the fused detection result map according to the class attribute of each pixel point in the fused detection result map. The road elements refer to objects having an association relationship with the road;
[0033] Identify the road scene category in which the target vehicle is currently located according to the class attribute of each pixel point and the target detection frames in the fused detection result map;
[0034] Use the identified road elements and the road scene category as the road information.
[0035] On the other hand, a road information recognition device is provided. The device includes:
[0036] A first acquisition module, configured to acquire a plurality of fisheye images collected by a target vehicle, and acquire a first top view corresponding to the plurality of fisheye images, where the plurality of fisheye images are images collected from multiple perspectives;
[0037] A second acquisition module, configured to acquire a road segmentation detection result map for each of the plurality of fisheye images and a road segmentation detection result map for the first top view;
[0038] A fusion module, configured to fuse the road segmentation detection result maps of the plurality of fisheye images and the road segmentation detection result map of the first top view to obtain a fusion detection result map;
[0039] A third acquisition module, configured to acquire road information according to the fusion detection result map.
[0040] Optionally, the second acquisition module includes:
[0041] A segmentation sub-module, configured to perform image semantic segmentation on each of the plurality of fisheye images to obtain a segmentation result map for each fisheye image;
[0042] A detection sub-module, configured to perform object detection on each fisheye image to obtain an object detection result for each fisheye image, where the object detection result is used to indicate whether a target detection frame is included in the corresponding image;
[0043] A generation sub-module, configured to generate a road segmentation detection result map for the corresponding fisheye image according to the segmentation result map and the object detection result of each fisheye image.
[0044] Optionally, the first acquisition module includes:
[0045] A first conversion sub-module, configured to convert the plurality of fisheye images to a top view coordinate system through coordinate conversion to obtain the first top view.
[0046] Optionally, the fusion module includes:
[0047] A second conversion sub-module, configured to convert the road segmentation detection result maps of the plurality of fisheye images to a top view coordinate system through coordinate conversion to obtain a third top view;
[0048] A fusion sub-module, configured to fuse the third top view and the road segmentation detection result map of the first top view to obtain a fusion detection result map.
[0049] Optionally, the second conversion sub-module is mainly configured to:
[0050] Convert the road segmentation detection result map of each fisheye image to the top view coordinate system to obtain a top view sub-map corresponding to each fisheye image;
[0051] Stitch the top - down sub - graphs corresponding to the multiple fisheye images respectively to obtain a stitched top - down view;
[0052] If the road segmentation detection result images of the multiple fisheye images further include target detection frames, convert the target detection frames to the stitched top - down view to obtain the third top - down view.
[0053] Optionally, the road segmentation detection result image includes the class attribute of each pixel point. The second conversion sub - module is mainly used for:
[0054] If the first region in the first top - down sub - graph and the second region in the second top - down sub - graph are overlapping regions, when the class attributes of two pixel points at the same position in the first region and the second region are different, determine a pixel point with the highest class - attribute priority from the two pixel points;
[0055] Use the class attribute of the determined pixel point as the class attribute of the pixel point at the corresponding position in the stitched top - down view.
[0056] Optionally, the road segmentation detection result image includes the class attribute of each pixel point. The fusion sub - module is mainly used for:
[0057] If the distance represented by each pixel point in the third top - down view is different from the distance represented by each pixel point in the road segmentation detection result image of the first top - down view, convert the pixel points in the third top - down view so that the distance represented by each pixel point in the converted third top - down view is the same as the distance represented by each pixel point in the road segmentation detection result image of the first top - down view;
[0058] For multiple third pixel points in the converted third top - down view that have corresponding pixel points in the road segmentation detection result image of the first top - down view, use the class attribute with the highest priority among the class attributes of each third pixel point and the corresponding pixel point as the class attribute of the corresponding pixel point;
[0059] If the converted third top - down view further includes target detection frames, fuse the target detection frames in the converted third top - down view into the road segmentation detection result image of the first top - down view to obtain the fused detection result image.
[0060] Optionally, the fused detection result image includes the class attribute of each pixel point and target detection frames. The third acquisition module includes:
[0061] A first recognition sub - module, configured to recognize road elements included in the fused detection result image according to the class attribute of each pixel point in the fused detection result image, where the road elements refer to objects having an associated relationship with the road;
[0062] A second recognition sub-module, configured to recognize the road scene category in which the target vehicle is currently located according to the category attribute of each pixel point and the target detection frame in the fused detection result map;
[0063] A third determination sub-module, configured to use the recognized road elements and the road scene category as the road information.
[0064] On the other hand, a computer-readable storage medium is provided, in which a computer program is stored, and when the computer program is executed by a computer, the steps of the above-mentioned road information recognition method are implemented.
[0065] On the other hand, a computer program product containing instructions is provided, and when it runs on a computer, the computer is caused to execute the steps of the above-mentioned road information recognition method.
[0066] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:
[0067] In the embodiments of the present application, the road segmentation detection result map of the fisheye image and the road segmentation detection result map of the first top view are fused to obtain a fused detection result map, and then road information is obtained according to the fused detection result map. Since the fisheye image segmentation detection is more accurate in recognizing elements with height information, and the top view segmentation detection has a stronger ability to capture elements with structural and global information, by fusing the road segmentation detection results of the fisheye image and the top view, the more accurate elements with height information detected by the fisheye image and the more accurate elements with structural and global information detected by the top view can be fused, thereby reducing the phenomena of missed segmentation and mis-segmentation, as well as missed detection and mis-detection in road information recognition, and improving the accuracy of road information detection. Description of the Drawings
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 It is a system architecture diagram related to a road information recognition method provided by an embodiment of the present application;
[0070] Figure 2 It is a flowchart of a road information recognition method provided by an embodiment of the present application;
[0071] Figure 3 It is a flowchart of another road information recognition method provided by an embodiment of the present application;
[0072] Figure 4 It is a schematic structural diagram of a road information recognition device provided by an embodiment of the present application;
[0073] Figure 5 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific embodiments
[0074] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0075] Before explaining the embodiments of the present application in detail, the system architecture related to the embodiments of the present application will be introduced first.
[0076] Figure 1 It is a system architecture diagram related to a road information recognition method provided by an embodiment of the present application. As Figure 1 shown, the system includes multiple in-vehicle cameras 101 and a server 102. Among them, the multiple in-vehicle cameras 101 can communicate with the server 102 through a wireless network.
[0077] Among them, the multiple in-vehicle cameras 101 can be deployed around the vehicle body. Moreover, the viewing angles of the respective in-vehicle cameras 101 are different, and there may or may not be a partially overlapping area between the coverage ranges of two adjacent in-vehicle cameras 101. Exemplarily, the multiple in-vehicle cameras 101 may include a front-view fisheye camera deployed at the front end of the vehicle body, a rear-view fisheye camera deployed at the rear of the vehicle body, a left-view fisheye camera deployed on the left side of the vehicle body, and a right-view fisheye camera deployed on the right side of the vehicle body. Of course, in some possible implementation manners, the multiple in-vehicle cameras 101 may further include fewer or more in-vehicle cameras. For example, it may further include fisheye cameras deployed at the left front and right front of the vehicle body, etc. Or, there may be other deployment manners for the multiple in-vehicle cameras 101, and the embodiments of the present application do not limit this.
[0078] It should be noted that each in-vehicle camera 101 on the vehicle can collect a fisheye image in its own viewing angle and transmit the collected fisheye image to the server 102 through a wireless network.
[0079] The server 102 is configured to receive the fisheye images sent by the respective in-vehicle cameras of the vehicle, obtain the first top view corresponding to the fisheye images, obtain the road segmentation detection result images of each image in the multiple fisheye images and the first top view, and fuse the road segmentation detection results of the multiple fisheye images and the road segmentation detection result images of the first top view to obtain a fused detection result image. Furthermore, based on the fused detection result image, road information is obtained.
[0080] Among them, the server 102 can be a single server, a server cluster, or a cloud platform, and the embodiments of the present application do not limit this.
[0081] Next, the road information recognition method provided by the embodiments of the present application will be introduced.
[0082] Figure 2 It is a road information recognition method provided by the embodiments of the present application. This method can be applied to Figure 1 the server shown in Figure 2 As shown, this method includes the following steps:
[0083] Step 201: Obtain multiple fisheye images collected by a target vehicle, and obtain a first top view corresponding to the multiple fisheye images. The multiple fisheye images are images collected from multiple perspectives.
[0084] In the embodiments of the present application, a plurality of on-vehicle cameras are installed around the body of the target vehicle, and the perspectives of each on-vehicle camera are different. On this basis, each on-vehicle camera can collect a fisheye image from its own perspective, and then transmit the collected fisheye image from its own perspective to the server through a wireless network. The server receives the fisheye images collected from multiple perspectives sent by the multiple on-vehicle cameras.
[0085] Exemplarily, the multiple on-vehicle cameras can include a front-view fisheye camera, a rear-view fisheye camera, a left-view fisheye camera, and a right-view fisheye camera. Correspondingly, the fisheye images collected from multiple perspectives can include a front-view fisheye image, a rear-view fisheye image, a left-view fisheye image, and a right-view fisheye image.
[0086] After the server receives the multiple fisheye images, it can convert the multiple fisheye images to the top view coordinate system through coordinate transformation to obtain the first top view.
[0087] In a possible implementation manner, taking any one of the multiple fisheye images as an example, the server can determine a road surface coordinate system according to the current position of the target vehicle, and then determine the image coordinate system corresponding to the road surface coordinate system to obtain the top view coordinate system. Then, according to the conversion relationship between the image coordinate system of the fisheye image and the top view coordinate system, each pixel point in the fisheye image is converted to the top view coordinate system, so as to obtain the first top view.
[0088] Optionally, in some possible implementation manners, the server can also convert the multiple fisheye images to a second top view to obtain the first top view. The second top view is a blank top view generated according to the image acquisition range of the target vehicle.
[0089] Exemplarily, the blank top view generated by the server according to the image acquisition range of the target vehicle may be a blank top view corresponding to the image acquisition range of the target vehicle; alternatively, the server may also determine an area larger than the image acquisition range of the target vehicle with the image acquisition range of the target vehicle as the center. Then, a blank top view is generated according to the size of the area, and the above blank top view is the second top view. Among them, the pixel value corresponding to each pixel point in the second top view is a specified pixel value. For example, it may be 0 or 255 for all.
[0090] After generating the second top view, the server may convert the pixel points in each fisheye image to the second top view. Taking any one of the multiple fisheye images as an example, the server may convert the position coordinates of each pixel point in the fisheye image according to the conversion relationship between the image coordinate system of the fisheye image and the coordinate system of the image acquisition range, so as to obtain the position coordinates of the target position points corresponding to each pixel point in the image acquisition range of the target vehicle. Then, according to the conversion relationship between the coordinate system of the image acquisition range of the target vehicle and the image coordinate system of the second top view, the position coordinates of the target position points are converted, so as to obtain the position coordinates of the pixel points corresponding to the target position points in the second top view. In this way, the pixel point corresponding to each target position point in the second top view is the pixel point corresponding to the corresponding target position point in the fisheye image in the second top view. On this basis, the server may use the pixel value of each pixel point in the fisheye image as the pixel value of the corresponding pixel point in the second top view, so as to realize the conversion of the fisheye image to the second top view. After converting each fisheye image to the second top view through the above method, the converted one is the first top view.
[0091] Step 202: Obtain the road segmentation detection result map of each fisheye image among the multiple fisheye images and the road segmentation detection result map of the first top view.
[0092] In some embodiments, taking obtaining the road segmentation detection result map of each fisheye image among the multiple fisheye images as an example, the server may perform image semantic segmentation on each fisheye image among the multiple fisheye images to obtain the segmentation result map of each fisheye image; then perform target detection on each fisheye image to obtain the target detection result of each fisheye image, where the target detection result is used to indicate whether a target detection frame is included in the corresponding image; afterwards, according to the segmentation result map and the target detection result of each fisheye image, generate the road segmentation detection result map of the corresponding fisheye image.
[0093] Among them, taking any one of the multiple fisheye images as an example, for the convenience of description, it is called the first fisheye image. The server can input the first fisheye image into an image semantic segmentation network, and perform image semantic segmentation processing on the first fisheye image through the image semantic segmentation network to determine the class attributes of each pixel point in the first fisheye image. Then, according to the class attributes of each pixel point in the first fisheye image, a segmentation result image of the first fisheye image is generated. Among them, the pixel points with the same class attribute in the segmentation result image of the first fisheye image are divided into one region. Among them, the class attribute of each pixel point is used to indicate the road element to which the corresponding pixel point belongs. Exemplarily, the road element can be a road surface, a lane line, a road sign, a curb, a railing, a wall, etc.
[0094] Optionally, in some possible implementation manners, the server can also input the first fisheye image into a target detection network, perform target detection on the first fisheye image through the target detection network, and thus obtain the target detection result of the first fisheye image. Among them, the target detection result can include one or more target detection frames, or the target detection result can be used to indicate that the first fisheye image does not contain the target to be detected, that is, does not contain a target detection frame. Then, the server can fuse the target detection result of the first fisheye image with the segmentation result image of the first fisheye image to obtain the road segmentation detection result image of the first fisheye image.
[0095] It should be noted that if the target detection result of the first fisheye image contains one or more target detection frames, the target detection frames of the first fisheye image are converted into the segmentation result image of the first fisheye image to obtain the road segmentation detection result image of the first fisheye image; if the target detection result of the first fisheye image does not contain a target detection frame, the segmentation result image of the first fisheye image is directly used as the road segmentation detection result image of the first fisheye image.
[0096] Among them, when converting the target detection frame of the first fisheye image into the segmentation result image of the first fisheye image, since the target detection result of the first fisheye image and the image coordinate system corresponding to the segmentation result image of the first fisheye image are the same, the server can find the position coordinates of the corresponding position points in the segmentation result image of the first fisheye image according to the position coordinates of the center point and the four vertexes of the target detection frame in the target detection result of the first fisheye image. Then, the target detection frame in the segmentation result of the first fisheye image is converted into the segmentation result image of the first fisheye image to obtain the road segmentation detection result image of the first fisheye image.
[0097] Optionally, in some other embodiments, after the server performs image semantic segmentation on each of the multiple fisheye images, the segmentation result image containing the class attributes of each pixel point can also be directly used as the road segmentation detection result image of the corresponding fisheye image.
[0098] According to the same method, the server can obtain the road segmentation detection result map of the first top view. This is not elaborated in the embodiments of this application.
[0099] Optionally, before performing image semantic segmentation on each of the multiple fisheye images, each fisheye image can also be cropped and scaled according to actual needs, so that each of the cropped and scaled fisheye images has a fixed size. Then, the fisheye images with the fixed size after cropping and scaling are input into the image semantic segmentation network for semantic segmentation to determine the class attributes of each pixel point in each fisheye image. Among them, the sizes of the cropped and scaled fisheye images can be the same or different. Of course, the server can also crop and scale the first top view according to actual needs before performing image semantic segmentation on the first top view. This is not elaborated in the embodiments of this application.
[0100] Step 203: Fuse the road segmentation detection result maps of the multiple fisheye images and the road segmentation detection result map of the first top view to obtain a fused detection result map.
[0101] In some embodiments, the server converts the road segmentation detection result maps of the multiple fisheye images to the top view coordinate system through coordinate transformation to obtain a third top view; the third top view and the road segmentation detection result map of the first top view are fused to obtain a fused detection result map.
[0102] Among them, the server converts the road segmentation detection result map of each fisheye image in the multiple fisheye images to the top view coordinate system to obtain a top view sub-map corresponding to each fisheye image; then, the top view sub-maps corresponding to the multiple fisheye images are stitched together to obtain a stitched top view; if the road segmentation detection result maps of the multiple fisheye images also include target detection frames, the target detection frames are converted into the stitched top view to obtain a third top view.
[0103] Exemplarily, taking the road segmentation detection result map of any one of the multiple fisheye images as an example, referring to the method introduced in the foregoing step 201, the server can convert each pixel point in the road segmentation detection result map of the fisheye image to the top view coordinate system according to the conversion relationship between the image coordinate system of the road segmentation detection result map of the fisheye image and the top view coordinate system, so as to obtain a top view sub-map corresponding to the road segmentation detection result map of the fisheye image. After converting the road segmentation detection result maps of each fisheye image to the same top view coordinate system according to the same method, the top view sub-maps corresponding to the road segmentation detection result maps of each fisheye image can be obtained.
[0104] After obtaining the top-down sub-images corresponding to the respective fisheye images, since the top-down sub-images are in the same top-down coordinate system, the server can splice the top-down sub-images to obtain a spliced top-down view. Among them, since there may be an overlap in the coverage areas of two adjacent in-vehicle cameras on the vehicle body of the target vehicle, there may be an overlapping area in the top-down sub-images converted from the respective fisheye images collected by the in-vehicle cameras. Based on this, the server can process the overlapping area.
[0105] Exemplarily, taking any two adjacent top-down sub-images as an example, and calling them the first top-down sub-image and the second top-down sub-image, the server can first determine whether there is an overlapping area between the first top-down sub-image and the second top-down sub-image. If the first area in the first top-down sub-image and the second area in the second top-down sub-image are overlapping areas, then when the category attributes of two pixel points at the same position in the first area and the second area are different, determine a pixel point with the highest category attribute priority from the two pixel points; use the category attribute of the determined pixel point as the category attribute of the pixel point at the corresponding position in the spliced top-down view.
[0106] Among them, the server can determine whether there are pixel points with the same position coordinates in the first top-down sub-image and the second top-down sub-image. If there are pixel points with the same position coordinates in these two top-down sub-images, the area composed of these pixel points with the same position coordinates is the overlapping area. In this case, the server can determine whether the category attributes of the two pixel points at each position coordinate in this overlapping area are the same. If the category attributes of these two pixel points are different, determine a pixel point with the highest category attribute priority from these two pixel points, and use the pixel value and category attribute of the determined pixel point as the pixel value and category attribute of the pixel point at the corresponding position in the spliced top-down view. Optionally, if the category attributes of two pixel points with the same position coordinates in the overlapping area are the same, directly use the pixel value and category attribute corresponding to this pixel point as the pixel value and category attribute of the pixel point at the corresponding position in the spliced top-down view.
[0107] Among them, the priority of the category attribute can be set according to user needs. For example, when the user pays more attention to obstacles, the priority of the category attribute of obstacles can be set higher.
[0108] After determining the pixel values and category attributes of each pixel point in the spliced top-down view, if the road segmentation detection result images of multiple fisheye images also include target detection frames, the server can also convert the target detection frames to the spliced top-down view, so as to obtain a third top-down view.
[0109] Among them, taking the road segmentation detection result map of any one of the multiple fisheye maps as an example, the server can obtain the position coordinates of the center point and the position coordinates of the four vertices of the target detection box in the road segmentation detection result map of the fisheye map according to the image coordinate system of the road segmentation detection result map of the fisheye map. Then, according to the conversion relationship between the image coordinate system and the top-down coordinate system of the road segmentation detection result map of the fisheye map, the position coordinates of the center point of the target detection box and the position coordinates of the four vertices are converted to the stitched top-down view, so as to obtain the third top-down view.
[0110] Optionally, if the road segmentation detection result map of each fisheye map does not contain a target detection box, after converting the pixel values and class attributes of the pixel points in the road segmentation detection result map of each fisheye map to the stitched top-down view through the above method, the server uses the converted stitched top-down view as the third top-down view.
[0111] Optionally, in another possible implementation, the server can also convert the road segmentation detection result maps of multiple fisheye maps to the second top-down view to obtain the third top-down view, and the second top-down view is a blank top-down view generated according to the image acquisition range of the target vehicle.
[0112] Among them, the server divides the second top-down view into multiple regions. According to the positions of each pixel point in each region among the multiple regions, the pixel point corresponding to each pixel point in each region in the road segmentation detection result maps of the multiple fisheye maps is determined; according to the pixel values and class attributes of the pixel points corresponding to each pixel point in each region in the road segmentation detection result maps of the multiple fisheye maps, the pixel values and class attributes of each pixel point in each region are determined; if the road segmentation detection result maps of the multiple fisheye maps also include a target detection box, the target detection box is converted to the second top-down view to obtain the third top-down view.
[0113] Exemplarily, since the second top-down view is generated according to the image acquisition range of the target vehicle, that is, the second top-down view includes the image acquisition range around the body of the target vehicle. Based on this, the server can divide the second top-down view into eight corresponding regions according to the front area, the left front area, the left area, the left rear area, the rear area, the right rear area, the right area, and the right front area of the image acquisition range of the target vehicle. Of course, in some possible implementations, the second top-down view can also be divided into more or fewer regions. For example, the second top-down view can also be divided into four regions: the front area, the left area, the right area, and the rear area. Or, there can be other division methods for the multiple regions, and the embodiments of the present application do not limit this.
[0114] After dividing into multiple regions, the server determines, according to the positions of each pixel in each of the multiple regions, the corresponding pixel of each pixel in each region in the road segmentation detection result maps of the multiple fisheye images.
[0115] Among them, taking any one of the multiple regions as an example, the server can convert the position coordinates of each pixel in the region according to the conversion relationship between the image coordinate system of the region and the coordinate system of the image acquisition range corresponding to the region, so as to obtain the position coordinates of the target position points corresponding to each pixel in the region in the image acquisition range of the target vehicle. Then, according to the conversion relationship between the coordinate system of the image acquisition range of the target vehicle and the coordinate system of the road segmentation detection result maps of the multiple fisheye images, the server converts the position coordinates of the target position points, so as to obtain the position coordinates of the pixels corresponding to the target position points in the road segmentation detection result maps of the multiple fisheye images. In this way, the pixels corresponding to each target position point in the multiple fisheye images are the pixels corresponding to the corresponding target position point in the region in the multiple fisheye images.
[0116] After that, the server can determine the pixel value and category attribute of each pixel in each region according to the pixel value and category attribute of the pixel corresponding to each pixel in each region in the road segmentation detection result maps of the multiple fisheye images.
[0117] Among them, taking any one of the multiple regions as an example, for the sake of convenience of description, it is called the first region, and any pixel in the first region is called the first pixel. As can be seen from the foregoing introduction, there may be an overlap in the coverage ranges of two adjacent on-vehicle cameras on the body of the target vehicle. Therefore, the position point corresponding to the first pixel in the image acquisition range may be captured by both on-vehicle cameras at the same time. In this way, the first pixel may correspond to two pixels in two fisheye images. Based on this, in the embodiment of the present application, the server can first determine whether the pixels corresponding to the determined first pixel in the multiple fisheye images are one or two. If the first pixel in the first region corresponds to two pixels in the multiple fisheye images, and the category attributes corresponding to the two pixels are different, then a second pixel with the highest priority of category attribute is determined from the two pixels; the pixel value and category attribute of the second pixel are used as the pixel value and category attribute of the first pixel. Among them, the priority of the category attribute can be set according to user requirements, and this is not elaborated in the embodiment of the present application.
[0118] Optionally, if the first pixel in the first region corresponds to one pixel in the multiple fisheye images, the pixel value and category attribute of this pixel are directly used as the pixel value and category attribute of the first pixel.
[0119] After determining the class attributes and pixel values of the pixel points in each region of the second top view, if the road segmentation detection result maps of multiple fisheye images also include target detection frames, the server can also convert the target detection frames in the road segmentation detection result map of each fisheye image among the multiple fisheye images to the second top view, so as to obtain a third top view. The conversion method refers to the coordinate conversion method for converting a fisheye image to the second top view to obtain the first top view, which will not be elaborated in this embodiment of the present application.
[0120] After obtaining the third top view, since the third top view contains the road segmentation detection results of multiple fisheye images, the third top view and the road segmentation detection result map of the first top view can be fused to obtain a fused detection result map.
[0121] Among them, since the sizes of the third top view and the road segmentation detection result map of the first top view may not be the same, the distance represented by each pixel point included in the third top view may be different from the distance represented by each pixel point included in the road segmentation detection result map of the first top view. Based on this, in this embodiment of the present application, the server can first determine whether the distance represented by each pixel point included in the third top view is the same as the distance represented by each pixel point included in the road segmentation detection result map of the first top view. If the distance represented by each pixel point in the third top view is different from the distance represented by each pixel point in the road segmentation detection result map of the first top view, the pixel points in the third top view are converted so that the distance represented by each pixel point in the converted third top view is the same as the distance represented by each pixel point in the road segmentation detection result map of the first top view; for multiple third pixel points in the converted third top view that have corresponding pixel points in the road segmentation detection result map of the first top view, the class attribute with the highest priority among the class attributes of each third pixel point and the corresponding pixel point is used as the class attribute of the corresponding pixel point. If the converted third top view still includes target detection frames, the target detection frames in the converted third top view are fused into the road segmentation detection result map of the first top view to obtain a fused detection result map.
[0122] In one implementation, the server can calculate the distance represented by each pixel point included in the third top view according to the size of the third top view and the size of the image acquisition range corresponding to the third top view. According to the same method, the server can also obtain the distance represented by each pixel point included in the road segmentation detection result map of the first top view. Then, the server compares the distance represented by each pixel point in the third top view with the distance represented by each pixel point in the road detection result map of the first top view. If the distance represented by each pixel point in the third top view is greater than the distance represented by each pixel point in the road detection result map of the first top view, the third top view is enlarged so that the distance represented by each pixel point in the enlarged third top view is the same as the distance represented by each pixel point in the road detection map of the first top view. If the distance represented by each pixel point in the third top view is less than the distance represented by each pixel point in the road detection result map of the first top view, the third top view is reduced so that the distance represented by each pixel point in the reduced third top view is the same as the distance represented by each pixel point in the road detection map of the first top view.
[0123] After converting the distance represented by each pixel point in the third top view to be the same as the distance represented by each pixel point in the road segmentation detection result map of the first top view, the server can obtain a plurality of third pixel points from the third top view that have corresponding pixel points in the road detection result map of the first top view.
[0124] Among them, the server can refer to the method of determining the corresponding pixel points of each pixel point in the first region in the road segmentation detection result map of the fisheye view introduced above to determine the corresponding pixel points of each pixel point in the third top view in the first top view. If the position coordinates of a certain pixel point in the third top view are not within the coordinate range of the road segmentation detection result map of the first top view after conversion, it means that there is no corresponding pixel point for this pixel point in the road segmentation detection result map of the first top view. If the position coordinates of a certain pixel point in the third top view are within the coordinate range of the road segmentation detection result map of the first top view after conversion, it means that there is a corresponding pixel point for this pixel point in the road segmentation detection result map of the first top view. That is, this pixel point is a third pixel point. At this time, the corresponding pixel point of this third pixel point in the road segmentation detection result map of the first top view can be determined, that is, the corresponding pixel point of this third pixel point can be determined.
[0125] Since the class attributes of each third pixel point and the corresponding pixel point may be the same or different, for any third pixel point and the pixel point corresponding to the third pixel point, if their class attributes are different, then the class attribute with the highest priority among the class attributes of the third pixel point and the corresponding pixel point is used as the class attribute of the corresponding pixel point. If their class attributes are the same, the class attribute of the pixel point corresponding to the third pixel point is not updated. The method for determining the priority of the class attribute is the same as the method for determining the priority of the class attribute described above, and this will not be elaborated in the embodiments of this application.
[0126] Optionally, if the distance represented by each pixel point in the third top view is different from the distance represented by each pixel point in the road segmentation detection result map of the first top view, the pixel points corresponding to each pixel point in the third top view can also be directly found in the road segmentation detection result map of the first top view by means of coordinate transformation. Then, the server can obtain multiple third pixel points corresponding to pixel points in the road detection result map of the first top view from the third top view, and fuse the class attribute of each third pixel point among the multiple third pixel points in the third top view into the class attribute of the pixel point corresponding to it in the road detection result map of the first top view.
[0127] When fusing the class attribute of the pixel point in the third top view into the class attribute of the corresponding pixel point in the road detection result map of the first top view, if the converted third top view also includes a target detection frame, the target detection frame in the converted third top view can also be fused with the target detection frame in the road segmentation detection result map of the first top view.
[0128] Among them, the server first converts the target detection frame in the third top view to the road segmentation detection result map of the first top view. The conversion method refers to the above-mentioned conversion of the target detection frame to the second top view to obtain the coordinate conversion method of the third top view, and this will not be elaborated in the embodiments of this application.
[0129] After converting the target detection box in the third top view to the road segmentation detection result map of the first top view, for the target detection box in the road segmentation detection result map of the first top view, the server determines whether there is a first target detection box that intersects with the converted target detection box among the target detection boxes included in the road segmentation detection result map of the first top view. If the road segmentation detection result map of the first top view contains a first target detection box that intersects with the converted target detection box, the server further determines the intersection area of the two target detection boxes. If the ratio of the intersection area of the two target detection boxes to the area of one of the target detection boxes exceeds a preset ratio, the two target detection boxes are further fused. If the ratio of the intersection area of the two target detection boxes to the area of one of the target detection boxes does not exceed the preset ratio, the two target detection boxes are regarded as two independent detection boxes and no fusion process is performed.
[0130] Among them, the preset ratio can be set in advance. For example, the preset ratio is 60%, and the embodiments of the present application do not limit this.
[0131] It should be noted that when fusing two target detection boxes, the server can weight the position coordinates of the target detection box in the third top view and the corresponding target detection box in the road segmentation detection result map of the first top view to obtain the position coordinates of the fused target detection box. Among them, the position coordinates of the target detection box can include the position coordinates of the center point of the target detection box and / or the position coordinates of the four vertices. Of course, other methods can also be used to obtain the position coordinates of the fused target detection box, and the embodiments of the present application do not limit this.
[0132] Optionally, after converting the target detection box in the third top view to the road segmentation detection result map of the first top view, if there is no target detection box in the road segmentation detection result map of the first top view that intersects with the converted target detection box, the converted target detection box is used as a target detection box in the road segmentation detection result map of the first top view alone.
[0133] By converting the category attributes and target detection boxes of the pixels in the third top view to the road segmentation detection result map of the first top view, the fusion of the road segmentation detection result maps of the third top view and the first top view is realized, and thus a fused detection result map is obtained.
[0134] Optionally, the server may also fuse the class attributes corresponding to each pixel point in the road segmentation detection result map of the first top view with the target detection frame into the third top view, so as to obtain a fused detection result map. The fusion method may refer to the method of fusing the class attributes corresponding to each pixel point in the third top view with the target detection frame into the road segmentation detection result map of the first top view as described above, which will not be elaborated in this embodiment of the present application.
[0135] Optionally, when the road segmentation detection result map of the third top view and / or the first top view does not contain a target detection frame, the class attributes of the pixel points in the road segmentation detection result maps of the third top view and the first top view are directly fused, and the fused detection result map can be obtained.
[0136] In addition, the server may also perform erosion and dilation operations, Gaussian filtering operations, and multi-frame fusion operations on the fused detection result map to smooth the fused detection result map. Among them, the main function of erosion in the erosion and dilation operation is to eliminate the boundary points of the object and shrink the target, and noise points smaller than the structural element can be eliminated; the main function of dilation is to merge all background points in contact with the object into the object, enlarge the target, and fill the holes in the target. The main functions of these two operations are to make the boundary of the object in the fused detection result map smoother; Gaussian filtering is mainly used to eliminate Gaussian noise to reduce the noise in the fused detection result map and make the fused detection result map smoother; the multi-frame fusion operation refers to fusing the currently processed fused detection result map with the previously processed fused detection result map to make the fused detection result map smoother. By performing the above post-processing operations such as smoothing on the fused detection result map, the segmentation of pixel points is made continuous and the edge segmentation is more accurate. Of course, in some possible implementation manners, the vehicle's motion information and auxiliary information may also be used to preprocess the segmentation detection result map before fusion. The preprocessing includes correcting the road information in the segmentation detection result map. Exemplarily, the segmentation masks of the fisheye map and the first top view, as well as the position of the target detection frame, may be preprocessed and the position information corrected.
[0137] Step 204: Obtain road information according to the fused detection result map.
[0138] In some embodiments, the server identifies road elements in the fused detection result map according to the class attribute of each pixel point in the fused detection result map; identifies the road scene category where the target vehicle is currently located according to the class attribute of each pixel point in the fused detection result map and the target detection frame; and uses the identified road elements and the road scene category as road information.
[0139] Among them, the server may use an element composed of multiple pixel points with the same category attribute in the fusion detection result map as a road element, and the category attribute of this road element is the category attribute of these multiple pixel points. Exemplarily, the server may identify an element composed of all pixel points with the category attribute of road surface as the road surface, or the server may also identify an element composed of multiple consecutive pixel points with the category attribute of vehicle as the vehicle.
[0140] In addition, the server may also input the fusion detection result map into a deep learning classification network. The deep learning classification network may identify the road scene category where the target vehicle is currently located according to the category attribute of each pixel point and the target detection frame in the fusion detection result map. Exemplarily, the road scene category may include intersections, fork roads, lane line merges, parking lots, etc., and the embodiments of the present application do not limit this.
[0141] Optionally, when the target detection frame is not included in the fusion detection result map, the server may also input the fusion detection result map into the deep learning classification network, so that the deep learning classification network identifies the road scene category where the target vehicle is currently located based on the category attribute of each pixel point in the fusion detection result map.
[0142] In the embodiments of the present application, the road segmentation detection result map of the fisheye image and the road segmentation detection result map of the first top view are fused to obtain a fusion detection result map, and then road information is obtained according to the fusion detection result map. Since when performing fisheye image segmentation detection, the recognition of elements with height information is relatively accurate, and when detecting the top view, the ability to capture elements with structural and global information is stronger, so by fusing the road segmentation detection results of the fisheye image and the road segmentation detection results of the top view, the more accurate elements with height information detected by the fisheye image and the more accurate elements with structural and global information detected by the top view can be fused to achieve complementary advantages, thereby reducing missed segmentation, mis-segmentation, missed detection, and mis-detection in road information recognition, and enhancing the robustness and stability of the detection results. Subsequently, post-processing operations such as smoothing the fusion detection result map also make the segmentation of each pixel point continuous and the edge segmentation more accurate, improving the accuracy and comprehensiveness of road information detection.
[0143] In addition, in the embodiments of the present application, while segmenting the fisheye image and the first top view, target detection may also be performed. Then, while fusing the segmentation results, the target detection results are also fused. On this basis, the two fused detection results are identified to improve the recognition accuracy of road information.
[0144] Figure 3 It is a flowchart of an exemplary road information recognition method shown in the embodiments of the present application. See Figure 3, the server first obtains the front view, left view, right view, and rear view fisheye images, then performs coordinate transformation on the obtained front view, left view, right view, and rear view fisheye images to obtain the first top view, and then performs image semantic segmentation and object detection on the first top view to obtain the road segmentation detection result image of the first top view. At the same time, the server performs image semantic segmentation and object detection on the obtained front view, left view, right view, and rear view fisheye images respectively to obtain the road segmentation detection result images of multiple fisheye images, and then performs coordinate transformation on the road segmentation detection result images of multiple fisheye images to obtain the third top view. After that, the server fuses the road segmentation detection result image of the first top view with the third top view to obtain the fused detection result image, and then performs post-processing operations such as smoothing on the fused detection result image to make the segmentation of pixel points continuous and the edge segmentation more accurate. Finally, the server performs road element analysis and road scene recognition on the smoothed fused detection result image to obtain road information.
[0145] Next, the road information recognition device provided by the embodiments of the present application will be introduced.
[0146] See Figure 4 , an embodiment of the present application provides a road information recognition device 400, and the device 400 includes: a first acquisition module 401, a second acquisition module 402, a fusion module 403, and a third acquisition module 404.
[0147] The first acquisition module 401 is configured to acquire a plurality of fisheye images collected by a target vehicle, and acquire a first top view corresponding to the plurality of fisheye images, and the plurality of fisheye images are images collected from multiple perspectives;
[0148] The second acquisition module 402 is configured to acquire the road segmentation detection result image of each fisheye image in the plurality of fisheye images and the road segmentation detection result image of the first top view;
[0149] The fusion module 403 is configured to fuse the road segmentation detection result images of the plurality of fisheye images and the road segmentation detection result image of the first top view to obtain a fused detection result image;
[0150] The third acquisition module 404 is configured to acquire road information according to the fused detection result image.
[0151] Optionally, the second acquisition module 402 includes:
[0152] The segmentation sub-module is configured to perform image semantic segmentation on each fisheye image in the plurality of fisheye images to obtain the segmentation result image of each fisheye image;
[0153] The detection sub-module is configured to perform object detection on each fisheye image to obtain the object detection result of each fisheye image, and the object detection result is used to indicate whether a target detection box is included in the corresponding image;
[0154] A generation sub-module, configured to generate a road segmentation detection result map of a corresponding fisheye image according to the segmentation result map and the target detection result of each fisheye image.
[0155] Optionally, the first acquisition module 401 includes:
[0156] A first conversion sub-module, configured to convert multiple fisheye images to an overhead coordinate system through coordinate conversion to obtain a first overhead view.
[0157] Optionally, the fusion module 403 includes:
[0158] A second conversion sub-module, configured to convert the road segmentation detection result maps of multiple fisheye images to an overhead coordinate system through coordinate conversion to obtain a third overhead view;
[0159] A fusion sub-module, configured to fuse the road segmentation detection result maps of the third overhead view and the first overhead view to obtain a fused detection result map.
[0160] Optionally, the second conversion sub-module is mainly configured to:
[0161] Convert the road segmentation detection result map of each fisheye image to an overhead coordinate system to obtain an overhead sub-map corresponding to each fisheye image;
[0162] Stitch the overhead sub-maps corresponding to multiple fisheye images respectively to obtain a stitched overhead view;
[0163] If the road segmentation detection result maps of multiple fisheye images further include target detection frames, convert the target detection frames to the stitched overhead view to obtain a third overhead view.
[0164] Optionally, the road segmentation detection result map includes the class attribute of each pixel point, and the second conversion sub-module is mainly configured to:
[0165] If a first area in the first overhead sub-map and a second area in the second overhead sub-map are overlapping areas, when the class attributes of two pixel points at the same position in the first area and the second area are different, determine a pixel point with the highest class attribute priority from the two pixel points;
[0166] Use the class attribute of the determined pixel point as the class attribute of the pixel point at the corresponding position in the stitched overhead view.
[0167] Optionally, the road segmentation detection result map includes the class attribute of each pixel point, and the fusion sub-module is mainly configured to:
[0168] If the distance represented by each pixel point in the third top view is different from the distance represented by each pixel point in the road segmentation detection result map of the first top view, then the pixel points in the third top view are converted so that the distance represented by each pixel point in the converted third top view is the same as the distance represented by each pixel point in the road segmentation detection result map of the first top view;
[0169] For multiple third pixel points in the converted third top view that have corresponding pixel points in the road segmentation detection result map of the first top view, the category attribute with the highest priority among the category attributes of each third pixel point and the corresponding pixel point is used as the category attribute of the corresponding pixel point;
[0170] If the converted third top view further includes target detection frames, then the target detection frames in the converted third top view are fused into the road segmentation detection result map of the first top view to obtain a fused detection result map.
[0171] Optionally, the fused detection result map includes the category attribute of each pixel point and target detection frames. The third acquisition module 404 includes:
[0172] A first recognition sub-module, configured to recognize road elements included in the fused detection result map according to the category attribute of each pixel point in the fused detection result map, where the road elements refer to objects having an association relationship with the road;
[0173] A second recognition sub-module, configured to recognize the road scene category in which the target vehicle is currently located according to the category attribute of each pixel point and the target detection frames in the fused detection result map;
[0174] A determination sub-module, configured to use the recognized road elements and road scene category as road information.
[0175] In the embodiment of the present application, the road segmentation detection result map of the fisheye image and the road segmentation detection result map of the first top view are fused to obtain a fused detection result map, and then road information is obtained according to the fused detection result map. Since when performing fisheye image segmentation detection, elements with height information can be recognized more accurately, while when detecting the top view, the ability to capture elements with structural and global information is stronger. Therefore, by fusing the road segmentation detection results of the fisheye image and the road segmentation detection results of the top view, elements with more accurate height information detected by the fisheye image and elements with more accurate structural and global information detected by the top view can be fused to achieve complementary advantages, thereby reducing missed segmentation, false segmentation, missed detection, and false detection in road information recognition, and enhancing the robustness and stability of the detection results. Subsequently, post-processing operations such as smoothing are performed on the fused detection result map, which also makes the segmentation of each pixel point continuous and the edge segmentation more accurate, improving the accuracy and comprehensiveness of road information detection.
[0176] It should be noted that when the road information recognition device provided in the above embodiment recognizes road information, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the road information recognition device provided in the above embodiment and the embodiment of the road information recognition method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0177] Figure 5 It is a schematic structural diagram of a server shown according to an exemplary embodiment. The functions of road information recognition in the above embodiment can be implemented by Figure 5 the server shown in. This server can be a server in a background server cluster. Specifically:
[0178] The server 500 includes a central processing unit (CPU) 501, a system memory 504 including a random access memory (RAM) 502 and a read-only memory (ROM) 503, and a system bus 505 connecting the system memory 504 and the central processing unit 501. The server 500 also includes a basic input / output system (I / O system) 506 for facilitating the transmission of information between various components in the computer, and a mass storage device 507 for storing an operating system 513, application programs 514, and other program modules 515.
[0179] The basic input / output system 506 includes a display 508 for displaying information and input devices 509 such as a mouse, keyboard, etc. for user input of information. Both the display 508 and the input devices 509 are connected to the central processing unit 501 through an input / output controller 510 connected to the system bus 505. The basic input / output system 506 may also include an input / output controller 510 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 510 also provides outputs to a display screen, printer, or other types of output devices.
[0180] The mass storage device 507 is connected to the central processing unit 501 through a mass storage controller (not shown) connected to the system bus 505. The mass storage device 507 and its associated computer-readable medium provide non-volatile storage for the server 500. That is, the mass storage device 507 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0181] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage devices, CD-ROM, DVD (Digital Versatile Disc) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that computer storage media is not limited to the above several types. The above system memory 504 and mass storage device 507 may be collectively referred to as memory.
[0182] According to various embodiments of the present application, the server 500 may also run by connecting to a remote computer on a network such as the Internet. That is, the server 500 may be connected to the network 512 through a network interface unit 511 connected to the system bus 505, or in other words, the network interface unit 511 may also be used to connect to other types of networks or remote computer systems (not shown).
[0183] The above-mentioned memory further includes one or more programs, and the one or more programs are stored in the memory and configured to be executed by the CPU. The one or more programs include instructions for performing the road information recognition method provided in the embodiments of the present application.
[0184] The embodiments of the present application further provide a computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the server, the server can execute the road information recognition method provided in the above embodiments. For example, the computer-readable storage medium may be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0185] It should be understood that all or part of the steps for implementing the above embodiments can be realized by software, hardware, firmware or any combination thereof. When implemented by software, it can be realized in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-mentioned computer-readable storage medium.
[0186] That is, in some embodiments, a computer program product including instructions is further provided. When it runs on a computer, the computer is enabled to execute the road information recognition method provided in the above embodiments.
[0187] The above description is not intended to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A road information recognition method, characterized in that, The method includes: Obtaining a plurality of fisheye images collected by a target vehicle, and obtaining a first top view corresponding to the plurality of fisheye images, where the plurality of fisheye images are images collected from multiple perspectives; Obtaining a road segmentation detection result map for each of the plurality of fisheye images and a road segmentation detection result map for the first top view; Converting the road segmentation detection result maps of the plurality of fisheye images to a top-down coordinate system through coordinate transformation to obtain a third top view; Fusing the road segmentation detection result maps of the third top view and the first top view to obtain a fused detection result map; Obtaining road information according to the fused detection result map.
2. The method according to claim 1, wherein The obtaining of the road segmentation detection result map for each of the plurality of fisheye images includes: Performing image semantic segmentation on each of the plurality of fisheye images to obtain a segmentation result map for each fisheye image; Performing object detection on each fisheye image to obtain an object detection result for each fisheye image, where the object detection result is used to indicate whether an object detection frame is included in the corresponding image; Generating a road segmentation detection result map for the corresponding fisheye image according to the segmentation result map and the object detection result of each fisheye image.
3. The method according to claim 1, wherein The obtaining of the first top view corresponding to the plurality of fisheye images includes: Converting the plurality of fisheye images to a top-down coordinate system through coordinate transformation to obtain the first top view.
4. The method according to claim 1, characterized in that The converting of the road segmentation detection result maps of the plurality of fisheye images to a top-down coordinate system through coordinate transformation to obtain a third top view includes: Converting the road segmentation detection result map of each fisheye image to the top-down coordinate system to obtain a top-down sub-map corresponding to each fisheye image; Stitching the top-down sub-maps corresponding to the plurality of fisheye images to obtain a stitched top view; If the road segmentation detection result maps of the plurality of fisheye images further include object detection frames, converting the object detection frames to the stitched top view to obtain the third top view.
5. The method according to claim 4, characterized in that The road segmentation detection result map includes the class attribute of each pixel point. The stitching of the top-down sub-maps corresponding to the plurality of fisheye images to obtain a stitched top view includes: If a first area in a first top-down sub-map and a second area in a second top-down sub-map are overlapping areas, when the class attributes of two pixel points at the same position in the first area and the second area are different, determining a pixel point with the highest class attribute priority from the two pixel points; Taking the class attribute of the determined pixel point as the class attribute of the pixel point at the corresponding position in the stitched top view.
6. The method according to claim 1, characterized in that The road segmentation detection result map includes the class attribute of each pixel point. The fusing of the road segmentation detection result maps of the third top view and the first top view to obtain the fused detection result map includes: If the distance represented by each pixel point in the third top view is different from the distance represented by each pixel point in the road segmentation detection result map of the first top view, then convert the pixel points in the third top view so that the distance represented by each pixel point in the converted third top view is the same as the distance represented by each pixel point in the road segmentation detection result map of the first top view; For multiple third pixel points in the converted third top view that have corresponding pixel points in the road segmentation detection result map of the first top view, use the category attribute with the highest priority among the category attributes of each third pixel point and the corresponding pixel point as the category attribute of the corresponding pixel point; If the converted third top view further includes a target detection frame, then fuse the target detection frame in the converted third top view into the road segmentation detection result map of the first top view to obtain the fused detection result map.
7. The method according to claim 1, wherein The fused detection result map includes the category attribute of each pixel point and the target detection frame. Obtaining road information according to the fused detection result map includes: Identifying the road elements included in the fused detection result map according to the category attribute of each pixel point in the fused detection result map, where the road elements refer to objects having an association relationship with the road; Identifying the road scene category in which the target vehicle is currently located according to the category attribute of each pixel point and the target detection frame in the fused detection result map; Taking the identified road elements and the road scene category as the road information.
8. A road information recognition device, characterized in that, The device includes: A first acquisition module, configured to acquire a plurality of fisheye images collected by a target vehicle, and acquire a first top view corresponding to the plurality of fisheye images, where the plurality of fisheye images are images collected from multiple perspectives; The second acquisition module includes a second conversion sub-module and a fusion sub-module. The second conversion sub-module is configured to convert the road segmentation detection result maps of the plurality of fisheye images to the top view coordinate system through coordinate conversion to obtain a third top view; the fusion sub-module is configured to fuse the third top view and the road segmentation detection result map of the first top view to obtain a fused detection result map; A fusion module, configured to fuse the road segmentation detection result maps of the plurality of fisheye images and the road segmentation detection result map of the first top view to obtain a fused detection result map; A third acquisition module, configured to acquire road information according to the fused detection result map.
9. The device according to claim 8, wherein The second acquisition module includes: A segmentation sub-module, configured to perform image semantic segmentation on each of the plurality of fisheye images to obtain a segmentation result map of each fisheye image; A detection sub-module, configured to perform target detection on each fisheye image to obtain a target detection result of each fisheye image, where the target detection result is used to indicate whether a target detection frame is included in the corresponding image; A generation sub-module, configured to generate a road segmentation detection result map of the corresponding fisheye image according to the segmentation result map and the target detection result of each fisheye image.
10. The device according to claim 8, characterized in that, The first acquisition module includes: A first conversion sub-module, configured to convert the plurality of fisheye images to the top view coordinate system through coordinate conversion to obtain the first top view.
11. The device according to claim 8, wherein The second conversion sub-module is mainly used for: Converting the road segmentation detection result graph of each fisheye image to the overhead coordinate system to obtain an overhead sub-graph corresponding to each fisheye image; Stitching the overhead sub-graphs corresponding to the multiple fisheye images respectively to obtain a stitched overhead view; If the road segmentation detection result graphs of the multiple fisheye images further include target detection frames, converting the target detection frames to the stitched overhead view to obtain the third overhead view.
12. The device according to claim 11, characterized in that, The road segmentation detection result graph includes the class attribute of each pixel point. The second conversion sub-module is mainly used for: If the first area in the first overhead sub-graph and the second area in the second overhead sub-graph are overlapping areas, when the class attributes of two pixel points at the same position in the first area and the second area are different, determining a pixel point with the highest class attribute priority from the two pixel points; Taking the class attribute of the determined pixel point as the class attribute of the pixel point at the corresponding position in the stitched overhead view.
13. The device according to claim 8, characterized in that, The road segmentation detection result graph includes the class attribute of each pixel point. The fusion sub-module is mainly used for: If the distance represented by each pixel point in the third overhead view is different from the distance represented by each pixel point in the road segmentation detection result graph of the first overhead view, converting the pixel points in the third overhead view so that the distance represented by each pixel point in the converted third overhead view is the same as the distance represented by each pixel point in the road segmentation detection result graph of the first overhead view; For multiple third pixel points in the converted third overhead view that have corresponding pixel points in the road segmentation detection result graph of the first overhead view, taking the class attribute with the highest priority among the class attributes of each third pixel point and the corresponding pixel point as the class attribute of the corresponding pixel point; If the converted third overhead view further includes a target detection frame, fusing the target detection frame in the converted third overhead view into the road segmentation detection result graph of the first overhead view to obtain the fusion detection result graph.
14. The device according to claim 8, characterized in that, The fusion detection result graph includes the class attribute of each pixel point and a target detection frame. The third acquisition module includes: A first recognition sub-module, configured to recognize road elements included in the fusion detection result graph according to the class attribute of each pixel point in the fusion detection result graph, where the road elements refer to objects having an association relationship with the road; A second recognition sub-module, configured to recognize the road scene category in which the target vehicle is currently located according to the class attribute of each pixel point and the target detection frame in the fusion detection result graph; A third determination sub-module, configured to use the recognized road elements and the road scene category as the road information.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, it implements the road information recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Lane line detection method and device
CN112990099A