Map segmentation method and device based on bird's eye view features and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2026-08-11
AI Technical Summary
[0029] The present invention provides a map segmentation method, apparatus, and electronic device based on bird's-eye view features. By inputting the target detection result of the previous moment and the image features of the vehicle's direction of travel at the current moment into a pre-trained deep learning model for implicit viewpoint transformation, the bird's-eye view features of the vehicle's direction of travel at the current moment are obtained. The bird's-eye view features are then processed by a preset network model to determine a more accurate map segmentation result from the bird's-eye view perspective.
Smart Images

Figure CN116051568B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of map segmentation applications, and in particular to a map segmentation method, apparatus, and electronic device based on bird's-eye view features. Background Technology
[0002] The technological development of autonomous driving has shifted from rule-based design to data-driven approaches, which is reflected in various process modules such as control, planning, map fusion, and generalized perception. The generalized perception module, in particular, relies almost entirely on data-driven processes. This technological advancement has led to process integration and placed new demands on algorithms: in the perception module, the approach is no longer satisfied with sequentially executing a single inference task and then performing post-fusion on the results. Instead, the goal is to design end-to-end perception models capable of performing multiple task inferences in parallel and outputting information that can be directly utilized by subsequent modules.
[0003] Therefore, in map segmentation applications, there is an urgent need for a method that can meet current requirements while achieving accurate segmentation. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a map segmentation method based on bird's-eye view features, which alleviates the technical problem of map segmentation applications that cannot meet current needs while achieving accurate segmentation.
[0005] In a first aspect, an embodiment provides a map segmentation method based on bird's-eye view features, the method comprising:
[0006] Obtain the target detection results of the current vehicle at the previous moment and the image features collected in the direction of travel at the current moment;
[0007] The image features and the target detection results are processed based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective is transformed.
[0008] The bird's-eye view features are input into a preset network model to determine the map segmentation result from the bird's-eye view perspective.
[0009] In an optional implementation, before the step of processing the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective transformation, the method further includes:
[0010] Based on the location information of each acquisition module, the location code of each acquisition module is determined;
[0011] The intrinsic and extrinsic parameters of each acquisition module are encoded using the inverse perspective transformation (IPM) method to obtain the direction code of each acquisition module.
[0012] In an optional implementation, before the step of processing the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective transformation, the method further includes:
[0013] The target detection results are structured and feature aligned according to the preset categories of the target objects to determine the map topology prior information.
[0014] In an optional implementation, the step of processing the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective transformation includes:
[0015] The direction encoding, the location encoding, and the map topology prior information are input into a pre-trained deep learning model. The viewpoint is transformed based on the direction encoding and the location encoding to determine the bird's-eye view features.
[0016] In an optional implementation, the step of obtaining the target detection result of the current vehicle at the previous moment and the image features acquired in the direction of travel at the current moment includes:
[0017] Obtain the target detection result of the current vehicle in the direction of travel in the previous moment;
[0018] The acquisition module collects an image sequence of the vehicle in its current direction of travel at the current moment.
[0019] The image sequence is input into an image feature extractor to obtain the image features corresponding to the current vehicle's direction of travel.
[0020] In an optional implementation, the step of inputting the bird's-eye view features into a preset network model to determine the map segmentation result from the bird's-eye view perspective includes:
[0021] Based on a preset network model, the bird's-eye view features are processed according to the preset category and preset space occupied by the target object to determine the map segmentation result.
[0022] In an optional implementation, the pre-trained deep learning model includes a self-attention mechanism deep learning model.
[0023] Secondly, an embodiment provides a map segmentation device based on bird's-eye view features, the device comprising:
[0024] The acquisition module acquires the target detection results of the current vehicle at the previous moment and the image features collected in the direction of travel at the current moment.
[0025] The determination module processes the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective conversion;
[0026] The segmentation module inputs the bird's-eye view features into a preset network model to determine the map segmentation result from the bird's-eye view perspective.
[0027] Thirdly, an embodiment provides an electronic device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method described in any of the foregoing embodiments.
[0028] Fourthly, an embodiment provides a machine-readable storage medium storing machine-executable instructions, which, when invoked and executed by a processor, cause the processor to perform the steps of the method described in any of the foregoing embodiments.
[0029] The present invention provides a map segmentation method, apparatus, and electronic device based on bird's-eye view features. By inputting the target detection result of the previous moment and the image features of the vehicle's direction of travel at the current moment into a pre-trained deep learning model for implicit viewpoint transformation, the bird's-eye view features of the vehicle's direction of travel at the current moment are obtained. The bird's-eye view features are then processed by a preset network model to determine a more accurate map segmentation result from the bird's-eye view perspective.
[0030] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0031] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0032] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0033] Figure 1 A flowchart of a map segmentation method based on bird's-eye view features provided in an embodiment of the present invention;
[0034] Figure 2A flowchart of another map segmentation method based on bird's-eye view features provided in an embodiment of the present invention;
[0035] Figure 3 A functional module diagram of a map segmentation device based on bird's-eye view features provided in an embodiment of the present invention;
[0036] Figure 4 This is a schematic diagram of the hardware architecture of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] The inventors discovered that the newly developed Bird-Eye-View (BEV) paradigm can bring significant improvements to the system. Compared with traditional methods, it can make full use of information from various sensors and improve perception performance.
[0039] Based on this, the present invention provides a map segmentation method based on bird's-eye view features, which alleviates the technical problem of map segmentation applications failing to meet current needs while achieving accurate segmentation.
[0040] To facilitate understanding of this embodiment, a map segmentation method based on bird's-eye view features disclosed in this embodiment will first be described in detail. This method can be applied to intelligent control devices such as vehicle controllers, vehicle-mounted systems, host computers, and servers.
[0041] Figure 1 This is a flowchart of a map segmentation method based on bird's-eye view features provided in an embodiment of the present invention.
[0042] like Figure 1 As shown, the method includes the following steps:
[0043] Step S102: Obtain the target detection result of the current vehicle at the previous moment and the image features collected in the direction of travel at the current moment.
[0044] Here, the image feature can be understood as the road condition image features of the road ahead of the vehicle, which may include one or more target objects, such as pedestrians, road signs, vehicles, etc. The target detection result of the previous moment can be understood as the result of detecting target objects based on the image features collected in the vehicle's direction of travel at the previous moment.
[0045] Step S104: Based on the pre-trained deep learning model, the image features and target detection results are processed to determine the bird's-eye view features after the current vehicle's perspective is transformed.
[0046] Among them, the pre-trained deep learning model may include a self-attention mechanism deep learning model, such as a transformer. This model performs implicit viewpoint transformation processing by using the viewpoint information at the current moment and the depth information corresponding to the target detection result at the previous moment, and obtains the bird's-eye view features after the viewpoint transformation at the current moment.
[0047] Step S106: Input the bird's-eye view features into the preset network model to determine the map segmentation result from the bird's-eye view perspective.
[0048] Here, the preset network model may include a head network, which can perform more accurate map segmentation based on the bird's-eye view features after the perspective transformation.
[0049] In a preferred embodiment of practical application, the target detection result of the previous moment and the image features of the vehicle's direction of travel at the current moment are input into a pre-trained deep learning model for implicit viewpoint transformation to obtain the bird's-eye view features of the vehicle's direction of travel at the current moment. The bird's-eye view features are then processed by a preset network model to determine a more accurate map segmentation result from the bird's-eye view perspective.
[0050] The bird's-eye view (BEV) network architecture (pre-trained deep learning model) provided in this embodiment of the invention has several features: 1. It can directly output 3D spatial results; 2. It can process a set of perception tasks; 3. Compared with the post-fusion mode of rule design, it can fuse sensor information in a unified feature space; 4. It is a data-driven end-to-end learnable model.
[0051] In some embodiments, by acquiring the target detection results from the previous moment and the image features from the current moment, a more accurate map segmentation effect is achieved in the vehicle's direction of travel at the current moment; exemplarily, this step S102 includes:
[0052] Step 1.1) Obtain the target detection results of the current vehicle in the direction of travel in the previous moment.
[0053] This can be understood as the detection and recognition of target objects in road images collected during the vehicle's movement at the previous moment, such as the distribution of pedestrians, road signs, and various types of vehicles on the road at the previous moment.
[0054] Step 1.2) Collect the image sequence of the current vehicle in the direction of travel at the current moment through the acquisition module.
[0055] This can be achieved through multiple acquisition modules, such as multiple cameras, to acquire image sequences of the road in the direction of vehicle travel, and through other sensors (IMU, Odom, etc.) to acquire sensor data, which is then sent to the processing unit for subsequent processing steps.
[0056] Step 1.3) Input the image sequence into the image feature extractor to obtain the image features corresponding to the current vehicle's direction of travel.
[0057] The image sequence here can be understood as a 2D image input, and the image feature extractor has a mature network structure capable of extracting image features.
[0058] The embodiments of the present invention can be adapted to single BEV perception tasks (such as map segmentation) or multiple BEV perception tasks (map segmentation, 3D object detection). Figure 2 This is an example of an algorithm for a single BEV perception task. At time t, the camera captures n images with a resolution of H*W. The resolution of the input image sequence is then n x 3 x H*W. The image sequence with this resolution is input into a feature extractor to obtain image features.
[0059] In some embodiments, before inputting the image features and object detection results into the pre-trained deep learning model in step S104, the method further includes:
[0060] Step 2.1) Determine the location code of each acquisition module based on the location information of each acquisition module.
[0061] Step 2.2) Encode the intrinsic and extrinsic parameters of each acquisition module according to the inverse perspective transformation (IPM) method to obtain the direction code of each acquisition module.
[0062] Before utilizing the pre-trained deep learning model, i.e., the transformer structure, it is necessary to encode the positional information of the input image sequence. Here, the traditional IPM inverse transform is decomposed to encode the intrinsic and extrinsic parameters of the acquired acquisition module (camera). Here, K is the intrinsic parameter matrix, R is the extrinsic parameter matrix, and its direction encoding includes the product K of the inverse intrinsic and extrinsic parameter matrices. -1 R -1 The encoding for time t is the same as the encoding for location, which is the encoding for each camera.
[0063] In some embodiments, before inputting the image features and object detection results into the pre-trained deep learning model in step S104, the method further includes:
[0064] Step 3.1) The target detection results are structured and feature aligned according to the preset categories of the target objects to determine the map topology prior information.
[0065] This invention uses the output of the 3D object detection from the previous moment, which may contain nine dimensions of information (center coordinates, size parameters, yaw angle, and velocity) regarding other key categories such as vehicles and pedestrians. The center coordinates may include three coordinate axes (cx, cy, cz), the size parameters may include width, height, and length (w', h', l), the yaw angle theta, and the velocity includes velocity components on both the horizontal and vertical axes (vx, vy). The inventors have found that users, especially in partially occluded environments, always refer to previous information about other vehicles and pedestrians when adjusting their vehicle's motion. Therefore, this embodiment of the invention structures the target detection results from the previous moment, which are then utilized by the bird's-eye view query bev_query in the transformer.
[0066] Based on the aforementioned embodiments, step S104, which utilizes a pre-trained deep learning model, i.e., a transformer, to implement implicit viewpoint transformation, can also be achieved through the following steps:
[0067] Step 4.1) Input the orientation encoding, position encoding and map topology prior information into the pre-trained deep learning model, and perform viewpoint transformation on the map topology prior information according to the orientation encoding and position encoding to determine the bird's-eye view features.
[0068] In this method, the transformer is used to implicitly transform the viewpoint of the map topology prior information by applying camera intrinsic and extrinsic parameters.
[0069] like Figure 2 As shown, the image captured by the camera at time t is used for training and feature extraction. Then, the camera and its intrinsic and extrinsic parameters are encoded to obtain position and orientation codes. Simultaneously, the results of the 3D object detection and perception task at the previous time t-1 are obtained. The categories of target objects (vehicles, pedestrians) are structured and feature-aligned as map topology priors. Then, the Transformer structure is used to perform viewpoint transformation and learn the BEV features of the bird's-eye view.
[0070] This invention addresses map segmentation tasks based on bird's-eye view perspectives. It utilizes the transformer method to transform the view from the camera's perspective (PV) to the bird's-eye view (BEV). By implicitly encoding camera parameters, it replaces geometric transformations with perspective transformations from a deep learning network to achieve the view transformation. The 3D object detection results from the previous time step are structured and aligned, serving as a predefined prior for the current time step, ensuring a more accurate bird's-eye view. The camera's perspective (PV) can also be referred to as the front view.
[0071] In some embodiments, the more accurate bird's-eye view obtained using the foregoing embodiments can play an important role in the reliability of map segmentation applications; exemplarily, step S106 may include:
[0072] Step 5.1): Based on the preset network model, process the bird's-eye view features according to the preset category and preset space occupied by the target object to determine the map segmentation result.
[0073] In this process, the accurate bird's-eye view obtained from the aforementioned embodiment is used to obtain the task output from the preset network model (head network), which is a road structure segmentation map from the BEV perspective.
[0074] It should be noted that the head network is a structure used to further perform task-related tasks on the BEV features of the bird's-eye view, such as detecting road structure and target object categories. The output of the head network is a binary map of size m*h*w, where m is the category of interest, such as lane lines, pedestrian crossings, drivable areas, etc., and h*w is the predefined size of the BEV space, that is, the space occupied by the target object of each category.
[0075] This invention utilizes a Transformer to implicitly perform viewpoint transformation instead of traditional geometric transformation. Since the depth information in the results of 3D object detection tasks is relatively clear, the target objects are large and dependent on road structures, this prior can better serve map segmentation tasks.
[0076] like Figure 3 As shown, this embodiment of the invention also provides a map segmentation device 200 based on bird's-eye view features, the device comprising:
[0077] The acquisition module 201 acquires the target detection result of the current vehicle at the previous moment and the image features collected in the direction of travel at the current moment.
[0078] The determination module 202 processes the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective conversion;
[0079] The segmentation module 203 inputs the bird's-eye view features into a preset network model to determine the map segmentation result from the bird's-eye view perspective.
[0080] In some embodiments, before the step of processing the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective transformation, the determining module 202 is further specifically used to: determine the position code of each acquisition module based on the position information of each acquisition module; and encode the intrinsic and extrinsic parameters of each acquisition module according to the inverse perspective transformation (IPM) method to obtain the orientation code of each acquisition module.
[0081] In some embodiments, before the step of processing the image features and the target detection results based on a pre-trained deep learning model to determine the bird's-eye view features after the current vehicle perspective transformation, the determining module 202 is further specifically used to perform structuring processing and feature alignment operations on the target detection results according to the preset category of the target object to determine the map topology prior information.
[0082] In some embodiments, the determining module 202 is further specifically configured to input the direction encoding, the position encoding, and the map topology prior information into a pre-trained deep learning model, and perform viewpoint transformation on the map topology prior information based on the direction encoding and the position encoding to determine the bird's-eye view features.
[0083] In some embodiments, the acquisition module 201 is further specifically used to: acquire the target detection result of the current vehicle in the direction of travel at the previous moment; acquire the image sequence of the current vehicle in the direction of travel at the current moment through the acquisition module; and input the image sequence into the image feature extractor to obtain the image features corresponding to the direction of travel of the current vehicle.
[0084] In some embodiments, the segmentation module 203 is further configured to process the bird's-eye view features according to the preset category and preset space occupied by the target object based on a preset network model, and determine the map segmentation result.
[0085] In some embodiments, the pre-trained deep learning model includes a self-attention mechanism deep learning model.
[0086] Figure 4 This is a schematic diagram of the hardware architecture of the electronic device 300 provided in an embodiment of the present invention. See also... Figure 4 As shown, the electronic device 300 includes a machine-readable storage medium 301 and a processor 302, and may also include a non-volatile storage medium 303, a communication interface 304, and a bus 305; wherein the machine-readable storage medium 301, the processor 302, the non-volatile storage medium 303, and the communication interface 304 communicate with each other through the bus 305. The processor 302 can execute the map segmentation method based on bird's-eye view features described in the above embodiments by reading and executing machine-executable instructions for map segmentation based on bird's-eye view features in the machine-readable storage medium 301.
[0087] The machine-readable storage medium mentioned in this article can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0088] Non-volatile media can be non-volatile memory, flash memory, storage drives (such as hard disk drives), any type of storage disk (such as optical discs, DVDs, etc.), or similar non-volatile storage media, or combinations thereof.
[0089] It is understood that the specific operation methods of each functional module in this embodiment can be referred to the detailed description of the corresponding steps in the above method embodiment, and will not be repeated here.
[0090] The computer-readable storage medium provided in the embodiments of the present invention stores a computer program. When the computer program code is executed, it can implement the map segmentation method based on bird's-eye view features as described in any of the above embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0092] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0093] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0094] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention.
Claims
1. A map segmentation method based on bird's-eye view features, characterized in that, The method includes: The target detection result of the current vehicle at the previous moment and the image sequence collected in the direction of travel at the current moment are obtained. The image sequence is then input into the image feature extractor to obtain the image features corresponding to the direction of travel at the current moment. Based on the position information of each acquisition module, the position code of each acquisition module is determined; the intrinsic and extrinsic parameters of each acquisition module are encoded according to the inverse perspective transformation (IPM) method to obtain the orientation code of each acquisition module. The 3D target detection results from the previous moment are structured and feature aligned according to the preset category of the target object to determine the map topology prior information. The 3D target detection results include the center coordinates, size parameters, yaw angle and velocity information of the target object. The direction encoding, the position encoding, and the map topology prior information are input into a pre-trained self-attention mechanism deep learning model. The map topology prior information is implicitly transformed based on the direction encoding and the position encoding to determine the bird's-eye view features after the current vehicle's perspective transformation. The bird's-eye view features are then input into a preset network model to determine the map segmentation result under the bird's-eye view perspective.
2. The method according to claim 1, characterized in that, The steps of acquiring the target detection result of the current vehicle at the previous moment and the image sequence collected in the current direction of travel, and inputting the image sequence into an image feature extractor to obtain the image features corresponding to the current direction of travel include: Obtain the target detection result of the current vehicle in the direction of travel in the previous moment; The acquisition module collects an image sequence of the vehicle in its current direction of travel at the current moment. The image sequence is input into an image feature extractor to obtain the image features corresponding to the current vehicle's direction of travel.
3. The method according to claim 1, characterized in that, The steps of inputting the bird's-eye view features into a preset network model to determine the map segmentation results from the bird's-eye view perspective include: Based on a preset network model, the bird's-eye view features are processed according to the preset category and preset space occupied by the target object to determine the map segmentation result.
4. A map segmentation device based on bird's-eye view features, characterized in that, The device includes: The acquisition module acquires the target detection result of the current vehicle at the previous moment and the image sequence collected in the direction of travel at the current moment. The image sequence is then input into the image feature extractor to obtain the image features corresponding to the direction of travel at the current moment. The system comprises a determination module, which determines the position code of each acquisition module based on its position information; it encodes the intrinsic and extrinsic parameters of each acquisition module using the Inverse Perspective Transform (IPM) method to obtain the direction code of each acquisition module; it performs structured processing and feature alignment operations on the 3D target detection results of the previous time step according to the preset category of the target object to determine the map topology prior information, wherein the 3D target detection results include the center coordinates, size parameters, yaw angle, and velocity information of the target object; it inputs the direction code, the position code, and the map topology prior information into a pre-trained self-attention mechanism deep learning model, and performs implicit viewpoint transformation on the map topology prior information based on the direction code and the position code to determine the bird's-eye view features after the current vehicle's viewpoint transformation; and a segmentation module inputs the bird's-eye view features into a preset network model to determine the map segmentation results under the bird's-eye view perspective.
5. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 3.
6. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions that, when invoked and executed by a processor, cause the processor to perform the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Image processing method, device and equipment and computer readable storage medium
CN114723955A