Assisted driving method and apparatus
By displaying grid information generated through multimodal perception information and perception prediction networks, the safety issues of in-vehicle assisted driving in specific scenarios are solved, the driver's environmental perception ability is improved, collisions are avoided, and low-speed parking assistance function is realized.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-15
AI Technical Summary
Existing in-vehicle driver assistance functions are insufficient to provide drivers with adequate and safe driving assistance in scenarios such as parking in narrow spaces, driving in narrow alleys, making U-turns on narrow roads, passing other vehicles on narrow roads, and automatic emergency braking.
By acquiring multimodal perception information and using data collected from various sensors such as cameras, lidar, and millimeter-wave radar, combined with a perception prediction network, multiple occupancy grid information is generated and the target area is displayed on the central control screen with different grid granularities, thereby improving the driver's perception of the surrounding environment.
It improves the driver's panoramic perception of obstacles over a wide area and the refined perception of obstacles at close range, helping the driver avoid scrapes and providing low-speed parking assistance to enhance driving safety.
Smart Images

Figure CN2024121728_15052026_PF_FP_ABST
Abstract
Description
Assisted driving methods and devices Technical Field
[0001] This application relates to vehicle-mounted driver assistance technology, and more particularly to a driver assistance method and device. Background Technology
[0002] With advancements in engineering technology and improvements in intelligent sensing capabilities, in-vehicle driver assistance functions have become standard equipment on various vehicles, such as reversing guide lines and parking cameras. However, current in-vehicle driver assistance functions struggle to provide adequate and safe driving assistance in certain specific scenarios, such as parking in narrow spaces, driving in narrow alleys, making U-turns on narrow roads, passing other vehicles on narrow roads, automatic emergency braking (AEB), and rear automatic emergency braking (RAEB).
[0003] Summary of the Invention
[0004] This application provides a driving assistance method and device to enhance the driver's perception of the surrounding environment and avoid collisions.
[0005] In a first aspect, this application provides an assisted driving method, comprising: acquiring multimodal perception information, the multimodal perception information being collected by multiple sensors installed on a vehicle; acquiring multiple sets of display information, the display information including a target area to be displayed and a grid granularity; performing perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information, the occupied grid information being used to describe the distribution of obstacles around the vehicle; and displaying multiple occupied grid images on a human-machine interface based on the multiple occupied grid information.
[0006] In this embodiment, by displaying multiple occupied grid images with corresponding grid granularities on the vehicle's central control screen, the multiple occupied grid images display target areas of different ranges with different grid granularities. This can improve the driver's panoramic perception of obstacles in a large range, as well as the driver's refined perception of obstacles in a close range. The combination of the two can improve the driver's perception of the surrounding environment and avoid scratches.
[0007] In this embodiment, the driver can trigger the vehicle's low-speed parking assist function in several ways:
[0008] 1. The driver inputs the start command via vehicle buttons or the vehicle's central control screen.
[0009] In one possible implementation, the vehicle is equipped with an assisted driving mode button, which the driver can press to trigger the low-speed parking assist function. It should be understood that activating the target function (e.g., the low-speed parking assist function) via a button can be implemented in various ways, such as pressing the button corresponding to the target function, or repeatedly pressing the button briefly to poll for the target function, etc. This application does not specifically limit the implementation in this way.
[0010] In one possible implementation, the central control screen displays a main interface including an icon corresponding to the low-speed parking assist function. The driver clicks the icon to trigger the activation of the low-speed parking assist function. It should be understood that activating the target function (e.g., the low-speed parking assist function) via the central control screen (which has touchscreen capabilities) can be implemented in various ways. For example, directly clicking the target function's icon or selecting the target function's option from the vehicle's assisted driving mode menu, etc. This application embodiment does not specifically limit this approach.
[0011] 2. The driver can start the vehicle by voice input.
[0012] In one possible implementation, the driver speaks "Please activate low-speed parking assist function" into the vehicle's speaker. When the vehicle's system receives and recognizes the meaning of the voice, it activates the low-speed parking assist function.
[0013] It should be noted that, in addition to the two methods mentioned above, the embodiments of this application may also use other methods to trigger the vehicle to start the low-speed parking assist function, and no specific limitation is made thereto.
[0014] When the vehicle activates the low-speed parking assist function, the vehicle's infotainment system can first acquire multimodal perception information. The vehicle is equipped with various sensors, including camera sensors, lidar sensors, and millimeter-wave radar sensors. In this embodiment, the multimodal perception information can be obtained based on information collected from one or more of the aforementioned sensors.
[0015] A set of display information may include the target area to be displayed and the grid granularity. The target area can refer to the range of a real-world region displayed on the central control screen. For example, in the real world, the area within 3 meters around the vehicle (i.e., the vehicle itself), centered on the vehicle, can be a 6m × 6m square planar area extending 3m forward, 3m backward, 3m to the left, and 3m to the right, with a height of 2m added to form a 6m × 6m × 2m three-dimensional space; this three-dimensional space is the target area. The length, width, and height of the target area can be represented as W (length) × H (width) × D (height). When representing the target area, the vehicle itself can be used as the origin, and the range of the target area can be represented by the three-dimensional coordinates of its eight vertices. Alternatively, the range of the target area can be represented by the three-dimensional coordinates of a single vertex of the target area, along with the length, width, and height of the target area. This application does not specifically limit the representation of the target area's range in its embodiments.
[0016] The grid granularity is based on the principle of the occupied grid. An occupied grid refers to dividing the three-dimensional space into voxels, with each voxel represented by a binary value of 0 or 1 indicating whether it is occupied. Therefore, one voxel corresponds to one grid, and the grid granularity can be used to represent the granularity of dividing the three-dimensional space into voxels. For example, if a voxel is 1m × 1m × 1m, the aforementioned 6m × 6m × 2m three-dimensional space can be divided into 6 × 6 × 2 voxels, meaning this three-dimensional space corresponds to 6 × 6 × 2 grids.
[0017] Optionally, the larger the target area, the larger the selectable raster granularity; conversely, the smaller the target area, the smaller the selectable raster granularity. This allows for the display of more raster occupancy information within a larger target area, and more detailed raster occupancy information within a smaller target area.
[0018] In one possible implementation, the target area in a set of displayed information can refer to the area within 6m around the vehicle, that is, the length, width and height of the target area are 12m×12m×2m, and the grid granularity can be 20cm.
[0019] In one possible implementation, the target area in a set of displayed information can refer to the area within 3m around the vehicle, that is, the length, width and height of the target area are 6m×6m×2m, and the grid granularity can be 5cm.
[0020] In one possible implementation, in a set of displayed information, the target area can refer to the area within 1m around the vehicle, that is, the length, width and height of the target area are 2m×2m×2m, and the grid granularity can be 2cm.
[0021] In this embodiment, the target area and the grid granularity can be pre-established (which can be used as the default recommended configuration) or the correspondence between the target area and the grid granularity can be determined in real time, without any specific limitation.
[0022] In this embodiment of the application, the vehicle-mounted system can obtain display information through the following multiple methods:
[0023] 1. The displayed information is pre-set.
[0024] Multiple sets of display information can be pre-set and stored. When the display information is needed, the vehicle's infotainment system can directly read the memory to obtain these multiple sets of display information. For example, two sets of display information can be set. In one set, the target area refers to the area within 6 meters around the vehicle, that is, the length, width, and height of the target area are 12m × 12m × 2m, and the grid grain is 20cm. In the other set of display information, the target area refers to the area within 3 meters around the vehicle, that is, the length, width, and height of the target area are 6m × 6m × 2m, and the grid grain is 5cm.
[0025] 2. The displayed information is dynamically generated based on the driver's actions on the touchscreen.
[0026] The zooming operation described above can include the driver pressing two fingers on the touchscreen and sliding them outwards or inwards to change the range of the target area and the target resolution of obstacles; or, the zooming operation can include the driver pressing two fingers on the touchscreen and rotating, panning, or dragging to change the display orientation and range of the target area. The aforementioned operations can be referenced to the following operations when a driver uses an image application: to zoom in on image details, they can press the relevant area with their thumb and forefinger and then slide the two fingers outwards; or to zoom out to view the whole image, they can press the relevant area with their thumb and forefinger and then slide the two fingers inwards; or, to change the orientation of the image content, they can press the relevant area with two fingers and then drag the image to pan or rotate it.
[0027] In this embodiment, a human-machine interface is provided for the driver. The driver can zoom in / out on an initial target area using this interface to obtain the final target area and its corresponding grid granularity. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the aforementioned sets of display information.
[0028] In this embodiment, multimodal perception information and multiple sets of display information can be input into a pre-trained perception prediction network to output multiple occupancy grid information.
[0029] The aforementioned perception prediction network can refer to the relevant content of the neural network in this application, and its principles and specific implementation methods can also refer to that content. The perception prediction network can be pre-trained and downloaded to the vehicle's infotainment system. Then, when using the perception prediction network, the parameters in the perception prediction network can be updated based on real-time input and output to make the prediction results of the perception prediction network more accurate. This can obtain grid occupancy information that is more consistent with actual driving conditions, thereby assisting the driver in driving the vehicle more safely.
[0030] For a single set of display information (including the target area and grid granularity), the occupancy and semantic information of all grid positions within the target area constitute the occupancy grid information corresponding to that display information. For example, for a set of display information (the target area could be the area within 6m around the vehicle, i.e., the target area's length, width, and height are 12m × 12m × 2m, and the grid granularity could be 20cm), the occupancy and semantic information of all grid positions within the 6m area around the vehicle can be perceived and predicted, and the grid occupancy can be displayed with a 20cm grid granularity; or, for another set of display information (the target area could be the area within 3m around the vehicle, i.e., the target area's length, width, and height are 6m × 6m × 2m, and the grid granularity could be 5cm), the occupancy and semantic information of all grid positions within the 3m area around the vehicle can be perceived and predicted, and the grid occupancy can be displayed with a 5cm grid granularity.
[0031] In one possible implementation, the vehicle infotainment system can display a first occupied grid image on the central control screen, the first occupied grid image corresponding to a first target area and a first grid granularity; and display a second occupied grid image, the second occupied grid image corresponding to a second target area and a second grid granularity; wherein the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0032] Optionally, the first occupied grid screen includes n1 grids corresponding to obstacles, and the second occupied grid screen includes n2 grids corresponding to obstacles, where n1 and n2 are both positive integers.
[0033] The first occupied grid screen displays the image of the first target area, which displays the grid of n1 obstacles that are perceived and predicted within the first target area at the first grid granularity; the second occupied grid screen displays the image of the second target area, which displays the grid of n2 obstacles that are perceived and predicted within the second target area at the second grid granularity.
[0034] Optionally, the grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle refer to the same obstacle. n1 obstacles include the first obstacle and n2 obstacles include the second obstacle.
[0035] Both the first target area and the second target area are areas surrounding the vehicle. The difference between them lies in their extent. The first target area is larger than the second target area. Therefore, the first and second target areas may include the same obstacles. In other words, obstacles in the second target area may also exist in the first target area, and some obstacles in the first target area may not appear in the second target area.
[0036] To facilitate driver identification, in this embodiment, different obstacle grids can be displayed with different colors, different outlines, etc., or different obstacle grids can be marked with different information (text, numbers, etc.). This allows the driver to intuitively see the number, location, and type of obstacles in the target area, and thus take timely avoidance measures.
[0037] Based on this, in both the first and second occupied grid views, grids corresponding to the same obstacle can be identified with the same information. For example, in both views, the grids for the first and second obstacles can be displayed with the same color (e.g., yellow grids) and the same border (e.g., solid borders). Alternatively, in both views, the grids for the first and second obstacles (obstacles present in both the first and second target areas) can be marked with the same information (e.g., the same text, the same numbers, etc.). The aforementioned first and second obstacles refer to the same obstacle. This allows the driver to visually identify which obstacles are the same across different target areas, compare the distribution and distance of obstacles in different fields of view, and take timely evasive action.
[0038] Optionally, the grid of the third obstacle in the second occupied grid image is identified with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0039] To facilitate driver identification, in this embodiment, the grid of the nearest obstacle around the vehicle can be identified with specific information. For example, the grid of the nearest obstacle can be highlighted with a special color (red), or the grid of the nearest obstacle can be displayed in a flashing manner, or the nearest obstacle can be marked with text near the grid of the nearest obstacle as the nearest obstacle, or the distance between the nearest obstacle and the vehicle can be marked with numbers near the grid of the nearest obstacle. Furthermore, when the distance between the vehicle and the nearest obstacle is less than a set threshold (e.g., 0.5m), the vehicle's audio and lights will provide a synchronized warning.
[0040] It should be noted that in the embodiments of this application, other methods may also be used to display the grid of obstacles, and no specific limitation is made thereto.
[0041] In one possible implementation, the human-computer interaction interface described above can be displayed on the driver's terminal device (e.g., mobile phone, tablet, etc.) so that the driver can simultaneously view the distribution of obstacles around the vehicle and the vehicle's driving status.
[0042] In this embodiment, the driver can install a driver assistance application on the terminal device, which enables interconnection between the vehicle's infotainment system and the terminal device. The vehicle's infotainment system can transmit the images displayed on it to the terminal device in real time. The driver can then open the driver assistance application and view the images displayed on the vehicle's infotainment system.
[0043] In addition, the human-machine interface provided by the driver assistance application also allows users to operate the screen, such as zooming in / out, changing the screen orientation, etc.
[0044] In one possible implementation, the human-computer interaction interface of this application embodiment further includes a control for selecting a driving mode, which includes a driving mode and a parking mode. Therefore, the human-computer interaction interface is provided with two controls: "driving" and "parking".
[0045] In driving mode, the human-machine interface also includes controls for selecting auxiliary modes, which include automatic mode and manual mode. At this time, the human-machine interface has two controls: "automatic" and "manual".
[0046] In parking mode, the human-machine interface also includes controls for selecting assistance modes. These modes include automatic mode, manual mode, Automated Valet Parking (AVP) mode, and Auto Parking Assist (APA) mode. The human-machine interface displays four controls: "Automatic," "Manual," "AVP," and "APA." Automatic and manual modes can be selected when the driver is driving, while APA and AVP modes can be selected when the driver is in autonomous driving mode (the driver can be in or outside the vehicle and the vehicle will automatically complete the parking maneuver). APA mode activates the automatic parking function, and the display logic of the panoramic view area during parking follows the same pattern as in automatic mode.
[0047] Optionally, the vehicle can display the grid occupancy status and semantic information of obstacles around the vehicle on the central control screen at multiple grid granularities. Furthermore, during vehicle parking, the fineness of the human-machine interface can be adaptively adjusted according to real-time changes in the environment, allowing the driver to take preventative measures in advance and avoid scratches.
[0048] Optionally, the driver can manually define the target area, and the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities. This allows for a more refined display of areas of interest to the driver. Furthermore, the refinement of the human-machine interface can be adaptively adjusted in real time as the environment changes during vehicle parking, enabling the driver to take preventative measures in advance and avoid scratches.
[0049] Optionally, the entire garage / parking lot can be perceived in a panoramic view based on the vehicle's position and posture changes, thereby providing the driver with effective static environmental information for parking or automatic parking.
[0050] Secondly, this application provides an assisted driving device, comprising: a data acquisition module for acquiring multimodal perception information, the multimodal perception information being acquired by multiple sensors installed on a vehicle; an acquisition module for acquiring multiple sets of display information, the display information including a target area to be displayed and a grid granularity; a prediction module for performing perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information, the occupied grid information being used to describe the distribution of obstacles around the vehicle; and a display module for displaying multiple occupied grid images on a human-machine interface based on the multiple occupied grid information.
[0051] In one possible implementation, the display module is specifically used to display a first occupying grid image, the first occupying grid image corresponding to a first target area and a first grid granularity; and to display a second occupying grid image, the second occupying grid image corresponding to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0052] In one possible implementation, the first occupied grid screen includes n1 grids corresponding to obstacles, and the second occupied grid screen includes n2 grids corresponding to obstacles, where n1 and n2 are both positive integers.
[0053] In one possible implementation, the grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information, the first obstacle and the second obstacle are the same obstacle, the n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
[0054] In one possible implementation, the grid of the third obstacle in the second occupied grid image is identified with specific information, the third obstacle being the one closest to the vehicle among the n2 obstacles.
[0055] In one possible implementation, the display module is further configured to display a panoramic occupancy grid image on the human-machine interface based on panoramic occupancy grid information, wherein the panoramic occupancy grid information is obtained based on the historical occupancy grid information of the vehicle's driving environment and driving trajectory.
[0056] In one possible implementation, the displayed information is pre-set information.
[0057] In one possible implementation, the displayed information is dynamically generated based on the driver's actions on the touchscreen.
[0058] In one possible implementation, the operation includes at least one of the following operations: the driver presses the touch screen with two fingers and slides the two fingers outward or inward together; or, the driver presses the touch screen with two fingers and rotates, translates, or drags it.
[0059] In one possible implementation, the human-computer interaction interface further includes a control for selecting a driving mode, which includes one or more of the following modes: driving mode or parking mode.
[0060] In one possible implementation, in the driving mode, the human-machine interface further includes a control for selecting an assistance mode, which includes one or more of the following modes: automatic mode or manual mode.
[0061] In one possible implementation, in the parking mode, the human-machine interface further includes a control for selecting an assistance mode, which includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking (AVP) mode, or automatic parking assist (APA) mode.
[0062] In one possible implementation, the prediction module is specifically used to input the multimodal perception information and the multiple sets of display information into a pre-trained perception prediction network to output the multiple occupancy grid information.
[0063] In one possible implementation, the functions of the perception prediction network include: multimodal feature fusion based on multimodal perception information; temporal feature fusion; and perception prediction based on the display information.
[0064] In one possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein, the pinhole network is used to extract visual and spatial features of pinhole images acquired by a pinhole camera to obtain pinhole 3D image features, the fisheye network is used to extract visual and spatial features of fisheye images acquired by a fisheye camera to obtain fisheye 3D image features, the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features, and the multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
[0065] In one possible implementation, the perception prediction network further includes a temporal fusion network, which is used to align multi-frame 3D spatial features in a temporal order to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
[0066] In one possible implementation, the perceptual prediction network further includes a feature enhancement network, which performs multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and then fuses the image features again to achieve auxiliary feature enhancement.
[0067] In one possible implementation, the perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation. Then, based on the position features, it performs prediction of occupancy and semantic information to achieve prediction at arbitrary resolution.
[0068] In one possible implementation, the multiple sensors include pinhole cameras and fisheye cameras.
[0069] In one possible implementation, the multiple sensors also include lidar and / or millimeter-wave radar.
[0070] Thirdly, this application provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to perform the method as described in any one of the first aspects above.
[0071] Fourthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the first aspects above.
[0072] Fifthly, this application provides a computer program product comprising computer program code, which, when run on a computer, causes the computer to perform the method described in any one of the first aspects above. Attached Figure Description
[0073] Figure 1 is an exemplary functional block diagram of a vehicle 100 according to an embodiment of this application;
[0074] Figure 2 is an exemplary functional block diagram of an in-vehicle driver assistance system according to an embodiment of this application;
[0075] Figure 3 is a software structure block diagram of the vehicle system according to an embodiment of this application;
[0076] Figure 4 is a flowchart of process 400 of the assisted driving method provided in an embodiment of this application;
[0077] Figure 5 is a flowchart illustrating the perception prediction process according to an embodiment of this application;
[0078] Figure 6 is a schematic diagram of the human-computer interaction interface of the central control screen in an embodiment of this application;
[0079] Figure 7 is a schematic diagram of the human-computer interaction interface switching display perspective according to an embodiment of this application;
[0080] Figure 8 is a schematic diagram of the refined mode of an embodiment of this application;
[0081] Figure 9 is a schematic diagram of the human-computer interaction interface of the central control screen according to an embodiment of this application;
[0082] Figure 10 is a schematic diagram of a refined mode of an embodiment of this application;
[0083] Figures 11 and 12 are schematic diagrams of zooming operations in manual mode according to embodiments of this application;
[0084] Figure 13 is a schematic diagram of the human-computer interaction interface of the central control screen according to an embodiment of this application;
[0085] Figure 14 is a structural schematic diagram of the driver assistance device 1400 of this application. Detailed Implementation
[0086] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0087] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0088] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0089] Before describing the technical solutions of the embodiments of this application, the vehicle of the embodiments of this application will be described first with reference to the accompanying drawings.
[0090] Figure 1 is an exemplary functional block diagram of a vehicle 100 according to an embodiment of this application. As shown in Figure 1, components coupled to or included in the vehicle 100 may include a propulsion system 110, a sensor system 120, a control system 130, peripheral devices 140, a power supply 150, a computing device 160, and a driver interface 170. The components of the vehicle 100 may be configured to operate in a manner interconnected with each other and / or with other components coupled to the respective systems. For example, the power supply 150 may provide power to all components of the vehicle 100. The computing device 160 may be configured to receive data from and control the propulsion system 110, sensor system 120, control system 130, and peripheral devices 140. The computing device 160 may also be configured to generate an image display on the driver interface 170 and receive input from the driver interface 170.
[0091] It should be noted that in other examples, vehicle 100 may include more, fewer, or different systems, and each system may include more, fewer, or different components. Furthermore, the systems and components shown can be combined or divided in any manner, and this application does not impose any specific limitations on this.
[0092] The computing device 160 may include a processor 161, a transceiver 162, and a memory 163. The computing device 160 may be a controller of the vehicle 100 or part of a controller. The memory 163 may store instructions 1631 that run on the processor 161 to execute various functional applications and data processing of the vehicle 100, and may also store data created during the use of the vehicle 100 (e.g., map data 1632, etc.), and may also store an operating system (e.g., an embedded operating system such as Android, iOS, Windows, or Linux), and applications required for at least one function. The processor 161 included in the computing device 160 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., image processors, digital signal processors, etc.). Where the processor 161 includes more than one processor, such processors may operate individually or in combination. The computing device 160 can implement functions that control the vehicle 100 based on input received through the driver interface 170. The transceiver 162 is used for communication between the computing device 160 and various systems. Memory 163 may further include one or more volatile memory components and / or one or more non-volatile memory components, such as optical, magnetic, and / or organic storage devices, and memory 163 may be wholly or partially integrated with processor 161. Memory 163 may contain instructions 1631 (e.g., program logic) executable by processor 161 to perform various vehicle functions, including any of the functions or methods described herein.
[0093] The propulsion system 110 can provide powered movement for the vehicle 100. As shown in FIG1, the propulsion system 110 may include an engine / motor 114, an energy source 113, a transmission 112, and wheels / tires 111. In addition, the propulsion system 110 may additionally or alternatively include other components besides those shown in FIG1. This application does not specifically limit this.
[0094] Sensor system 120 may include several sensors for sensing information about the environment in which vehicle 100 is located. As shown in FIG1, the sensors of sensor system 120 include a Global Positioning System (GPS) 126, an Inertial Measurement Unit (IMU) 125, a lidar sensor 124, a camera sensor 123, a millimeter-wave radar sensor 122, and a brake 121 for modifying the position and / or orientation of the sensors. GPS 126 may be any sensor used to estimate the geographic location of vehicle 100. For this purpose, GPS 126 may include a transceiver to estimate the position of vehicle 100 relative to the Earth based on satellite positioning data. In an example, computing device 160 may be used to combine map data 1632 with GPS 126 to estimate the road on which vehicle 100 is traveling. IMU 125 may be used to sense changes in the position and orientation of vehicle 100 based on inertial acceleration and any combination thereof. In some examples, the combination of sensors in IMU 125 may include, for example, an accelerometer and a gyroscope. Other combinations of sensors in IMU 125 are also possible. The lidar sensor 124 can be viewed as an object detection system that uses light sensing to detect objects in the environment in which the vehicle 100 is located. Typically, the lidar sensor 124 can utilize optical remote sensing techniques to measure the distance to a target or other properties of the target by illuminating it with light. As an example, the lidar sensor 124 may include a laser source and / or laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For example, the lidar sensor 124 may include a laser rangefinder reflected by a rotating mirror and scans the laser around a digitized scene in one or two dimensions to acquire distance measurements at specified angular intervals. In the example, the lidar sensor 124 may include components such as a light (e.g., laser) source, scanner and optical system, light detector and receiver electronics, and a positioning and navigation system. By scanning the laser reflected back from an object, the lidar sensor 124 can determine the distance to the object, forming a 3D environmental map with accuracy up to the centimeter level. The camera sensor 123 may include any camera (e.g., pinhole camera, fisheye camera, still camera, video camera, etc.) for acquiring images of the environment in which the vehicle 100 is located. For this purpose, camera sensor 123 can be configured to detect visible light, or it can be configured to detect light from other parts of the spectrum, such as infrared or ultraviolet light. Other types of camera sensors 123 are also possible. Camera sensor 123 can be a two-dimensional detector, or it can have three-dimensional spatial range detection capabilities. In some examples, camera sensor 123 can be, for example, a distance detector configured to generate a two-dimensional image indicating the distance from camera sensor 123 to several points in the environment.For this purpose, camera sensor 123 can use one or more distance detection techniques. For example, camera sensor 123 can be configured to use structured light technology, in which vehicle 100 illuminates objects in the environment using a predetermined light pattern, such as a grid or checkerboard pattern, and uses camera sensor 123 to detect reflections from the predetermined light pattern on the objects. Based on the distortion in the reflected light pattern, vehicle 100 can be configured to detect the distance to points on the object. The predetermined light pattern can include infrared light or light of other wavelengths. Millimeter-wave radar sensor 122 typically refers to an object detection sensor with a wavelength of 1 to 10 mm and a frequency range of approximately 10 GHz to 200 GHz. The measurements of millimeter-wave radar sensor 122 contain depth information, which can provide the distance to the target; secondly, because millimeter-wave radar sensor 122 has a significant Doppler effect, it is very sensitive to velocity and can directly obtain the velocity of the target. The velocity of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream automotive millimeter-wave radar application frequency bands are 24GHz and 77GHz, respectively. The former has a wavelength of about 1.25cm and is mainly used for short-range perception, such as the vehicle's surrounding environment, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4mm and is used for medium and long-range measurement, such as automatic following, adaptive cruise control (ACC), emergency braking (AEB), etc.
[0095] Sensor system 120 may also include additional sensors, including, for example, sensors that monitor the internal systems of vehicle 100 (e.g., O2 monitor, fuel gauge, oil temperature, etc.). Sensor system 120 may also include other sensors. This application does not specifically limit this.
[0096] The control system 130 can be configured to control the operation of the vehicle 100 and its components. For this purpose, the control system 130 may include a steering unit 136, a throttle 135, a braking unit 134, a sensor fusion algorithm 133, a computer vision system 132, and a navigation / route control system 131. The control system 130 may additionally or alternatively include components other than those shown in FIG. 1. This application does not specifically limit its scope in this regard.
[0097] Peripheral device 140 can be configured to allow vehicle 100 to interact with external sensors, other vehicles, and / or the driver. For this purpose, peripheral device 140 may include, for example, a lighting system 145, a wireless communication system 144, a touchscreen 143, a microphone 142, and / or a speaker 141. Peripheral device 140 may additionally or alternatively include components other than those shown in FIG. 1. This application does not specifically limit its scope in this regard.
[0098] Power source 150 can be configured to provide power to some or all of the components of vehicle 100. For this purpose, power source 150 may include, for example, rechargeable lithium-ion or lead-acid batteries. In some examples, one or more battery packs may be configured to provide power. Other power materials and configurations are also possible. In some examples, power source 150 and energy source 113 may be implemented together, as in some fully electric vehicles.
[0099] The components of vehicle 100 can be configured to operate in a manner that interconnects with other components within and / or outside their respective systems. For this purpose, the components and systems of vehicle 100 can be communicatively linked together via system buses, networks, and / or other connection mechanisms.
[0100] Figure 2 is an exemplary functional block diagram of an in-vehicle driver assistance system according to an embodiment of this application. As shown in Figure 2, components coupled to or included in the in-vehicle driver assistance system may include a computing unit, sensors, a central control screen, a lighting system, and an audio system. The computing unit corresponds to the control system 130 in the embodiment shown in Figure 1, the sensors correspond to the sensor system 120 in the embodiment shown in Figure 1, mainly involving camera sensors 123 (including pinhole cameras and fisheye cameras), millimeter-wave radar sensors 122, and lidar sensors 124, the central control screen corresponds to the touch screen 143 in the embodiment shown in Figure 1, providing the driver with a human-machine interface, the lighting system corresponds to the lighting system 145 in the embodiment shown in Figure 1, and the audio system corresponds to the speaker 141 in the embodiment shown in Figure 1.
[0101] The vehicle-mounted driver assistance system in this application embodiment can also be referred to as vehicle infotainment system, vehicle central control system (hereinafter referred to as central control), etc., without specific limitation.
[0102] Figure 3 is a software structure block diagram of the vehicle system according to an embodiment of this application.
[0103] The layered architecture of the vehicle's infotainment system divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0104] The application layer can include a series of application packages.
[0105] As shown in Figure 3, the application package may include applications such as calling, maps, navigation, WLAN, Bluetooth, music, video, and driver assistance.
[0106] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0107] As shown in Figure 3, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0108] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0109] Content providers store and retrieve data, making that data accessible to applications. This data may include video, maps, audio, and recorded phone calls.
[0110] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A human-computer interface can consist of one or more views. For example, a human-computer interface that includes notification icons can include views for displaying text and views for displaying images.
[0111] The phone manager provides communication functions for the vehicle's infotainment system. This includes managing call status (including connection and disconnection).
[0112] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0113] The notification manager allows applications to display notifications in the status bar. These notifications can be used to convey informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, and flashing indicator lights.
[0114] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0115] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0116] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0117] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0118] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0119] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0120] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0121] A 2D graphics engine is a graphics engine for 2D drawing.
[0122] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0123] It is understood that the components included in the system framework layer, system library, and runtime layer shown in Figure 3 do not constitute a specific limitation on the vehicle infotainment system. In other embodiments of this application, the vehicle infotainment system may include more or fewer components than shown in the figure, or combine some components, or split some components, or have different component arrangements.
[0124] Since the embodiments of this application involve the application of neural networks, for ease of understanding, some nouns or terms used in the embodiments of this application will be explained below, and these nouns or terms are also part of the content of the invention.
[0125] (1) Neural Network
[0126] Neural Networks (NNs) are machine learning models. A neural network can be composed of neural units, which are computational units that take xs and an intercept of 1 as input. The output of such a computational unit can be:
[0127] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x s The weights are denoted by b, where b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.
[0128] (2) Deep Neural Networks
[0129] Deep neural networks (DNNs), also known as multilayer neural networks, can be understood as neural networks with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three layers based on their position: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs appear complex, the operation of each layer is actually quite simple, resembling a linear relationship as follows: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained through such a simple operation. Because DNNs have many layers, the coefficients W and the offset vector... The number of these parameters is therefore quite large. The definitions of these parameters in a DNN are as follows: Taking the coefficient W as an example: Assuming a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W resides, while the subscript corresponds to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the k-th neuron in layer L-1 to the j-th neuron in layer L are defined as follows: It's important to note that the input layer does not have a W parameter. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can perform more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrix of all layers in the trained deep neural network (a weight matrix formed by the vectors W from many layers).
[0130] (3) Convolutional Neural Network
[0131] A convolutional neural network (CNN) is a deep neural network with convolutional structures. It is a deep learning architecture, which refers to learning at multiple levels of abstraction using machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, where each neuron responds to an input image. A CNN contains a feature extractor consisting of convolutional layers and pooling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution with a trainable filter and an input image or a convolutional feature map.
[0132] A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution processing on the input signal. A convolutional layer can contain multiple convolution operators, also called kernels. In image processing, these operators act as filters, extracting specific information from the input image matrix. Essentially, a convolution operator can be a weight matrix, which is usually predefined. During the convolution operation, the weight matrix typically processes the input image pixel by pixel (or two pixels by two pixels, depending on the stride) along the horizontal direction, thus extracting specific features from the image. The size of the weight matrix should be related to the image size. It's important to note that the depth dimension of the weight matrix is the same as the depth dimension of the input image; during convolution, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a single-depth convolutional output. However, in most cases, multiple weight matrices of the same size (rows × columns) are used instead of a single weight matrix. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. This dimension can be understood as being determined by the "multiple" factors mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix can be used to extract edge information, another to extract specific colors, and yet another to blur unwanted noise. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these weight matrices also have the same size. These extracted feature maps are then merged to form the output of the convolution operation. The weight values in these weight matrices need to be obtained through extensive training in practical applications. The weight matrices formed by these trained weight values can be used to extract information from the input image, enabling the convolutional neural network to make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layers often extract more general features, which can also be called low-level features. As the depth of the convolutional neural network increases, the features extracted by later convolutional layers become increasingly complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem being solved.
[0133] Because it's often necessary to reduce the number of training parameters, pooling layers are frequently introduced periodically after convolutional layers. This can be a single convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. Average pooling calculates the average value of pixel values within a specific range as the result of average pooling. Max pooling takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.
[0134] After processing by convolutional / pooling layers, a convolutional neural network (CNN) is still insufficient to output the required information. As mentioned earlier, convolutional / pooling layers only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the CNN needs to utilize neural network layers to generate one or a set of desired class numbers of output. Therefore, the neural network can include multiple hidden layers, the parameters of which can be pre-trained based on training data relevant to a specific task type, such as image recognition, image classification, image super-resolution reconstruction, etc.
[0135] Optionally, after the multiple hidden layers in the neural network, there is also an output layer of the entire convolutional neural network. This output layer has a loss function similar to the classification cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the backpropagation will begin to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0136] (4) Recurrent Neural Network
[0137] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, the layers from the input layer to the hidden layer and then to the output layer are fully connected, but the nodes within each layer are unconnected. While this type of neural network has solved many difficult problems, it remains inadequate for many others. For example, predicting the next word in a sentence generally requires using the preceding words because words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is related to the outputs of previous sequences. Specifically, the network memorizes previous information and applies it to the calculation of the current output; that is, nodes within the same hidden layer are no longer unconnected but connected, and the input to a hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time step. Theoretically, RNNs can process sequential data of any length. Training an RNN is similar to training a traditional CNN or DNN. This algorithm also uses the backpropagation algorithm, but with one key difference: when an RNN is expanded, its parameters, such as W, are shared; however, this is not the case with traditional neural networks as illustrated above. Furthermore, in gradient descent, the output at each step depends not only on the network at the current step but also on the states of the network in previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).
[0138] Since we already have convolutional neural networks (CNNs), why do we need recurrent neural networks (RNNs)? The reason is simple. CNNs rely on the fundamental assumption that elements are independent of each other, and that input and output are also independent—like a cat and a dog. However, in the real world, many elements are interconnected. For example, stock prices fluctuate over time. Or, imagine someone saying, "I love traveling, and my favorite place is Yunnan. I definitely want to go there someday." Humans know the answer to this question is "Yunnan." Humans can infer from context. But how can machines do this? That's where RNNs come in. RNNs aim to give machines the ability to remember, just like humans. Therefore, the output of an RNN depends on both the current input information and historical memory information.
[0139] (5) Loss Function
[0140] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value. Based on the difference, we update the weight vector of each layer (usually pre-configuring parameters before the initial update). For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network predicts the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss.
[0141] (6) Backpropagation algorithm
[0142] Convolutional neural networks can employ backpropagation (BP) to correct the parameters in the initial super-resolution model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates an error loss; this error loss information is then propagated back to update the parameters in the initial super-resolution model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the super-resolution model, such as the weight matrix.
[0143] (7) Generative Adversarial Networks
[0144] Generative adversarial networks (GANs) are a type of deep learning model. This model comprises at least two modules: a generative model and a discriminative model. These two modules learn from each other through a game-like interaction, resulting in better outputs. Both the generative and discriminative models can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GANs is as follows: Taking an image-generating GAN as an example, suppose there are two networks, G (Generator) and D (Discriminator). G is a network that generates images by receiving random noise z and using this noise, denoted as G(z). D is a discriminative network used to determine whether an image is "real." Its input parameter is x, representing an image, and its output D(x) represents the probability that x is a real image. A value of 1 indicates that the image is 100% real, while a value of 0 indicates that the image is impossible to be real. During the training of this generative adversarial network (GAN), the goal of the generative network G is to generate realistic images to deceive the discriminator network D, while the goal of the discriminator network D is to distinguish the images generated by G from real images as much as possible. Thus, G and D constitute a dynamic "game," which is the "adversarial" aspect of the GAN. Ideally, the game will result in G generating images G(z) that are sufficiently realistic, while D struggles to determine whether the images generated by G are real or not, i.e., D(G(z)) = 0.5. This yields a superior generative model G that can be used to generate images.
[0145] In complex road conditions and driving processes, there are scenarios that demand high driving skills or present significant challenges, such as parking in narrow spaces, navigating narrow roads, making U-turns on narrow roads, passing oncoming vehicles on narrow roads, and automatic emergency braking (AEB) and rear automatic emergency braking (RAEB). Some vehicle assistance systems (such as 360-degree surround view systems) may not be able to effectively identify and display nearby obstacles (including objects, pedestrians, and other vehicles) in these scenarios, thus failing to provide the driver with information about nearby obstacles and hindering the provision of adequate and safe driving assistance. To address this issue, this application provides a driving assistance method and device.
[0146] Figure 4 is a flowchart of process 400 of the assisted driving method provided in an embodiment of this application. Process 400 can be executed by the vehicle 100 described above (especially the vehicle's infotainment system). Process 400 can be used for low-speed parking assistance functions in vehicle assisted driving. Process 400 is described as a series of steps or operations. It should be understood that process 400 can be executed in various orders and / or occur simultaneously, and is not limited to the execution order shown in Figure 4. Process 400 may include:
[0147] Step 401: Acquire multimodal perception information, which is collected by various sensors installed on the vehicle.
[0148] In this embodiment, the driver can trigger the vehicle's low-speed parking assist function in several ways:
[0149] 1. The driver inputs the start command via vehicle buttons or the vehicle's central control screen.
[0150] In one possible implementation, the vehicle is equipped with an assisted driving mode button, which the driver can press to trigger the low-speed parking assist function. It should be understood that activating the target function (e.g., the low-speed parking assist function) via a button can be implemented in various ways, such as pressing the button corresponding to the target function, or repeatedly pressing the button briefly to poll for the target function, etc. This application does not specifically limit the implementation in this way.
[0151] In one possible implementation, the central control screen displays a main interface including an icon corresponding to the low-speed parking assist function. The driver clicks the icon to trigger the activation of the low-speed parking assist function. It should be understood that activating the target function (e.g., the low-speed parking assist function) via the central control screen (which has touchscreen capabilities) can be implemented in various ways. For example, directly clicking the target function's icon or selecting the target function's option from the vehicle's assisted driving mode menu, etc. This application embodiment does not specifically limit this approach.
[0152] 2. The driver can start the vehicle by voice input.
[0153] In one possible implementation, the driver speaks "Please activate low-speed parking assist function" into the vehicle's speaker. When the vehicle's system receives and recognizes the meaning of the voice, it activates the low-speed parking assist function.
[0154] It should be noted that, in addition to the two methods mentioned above, the embodiments of this application may also use other methods to trigger the vehicle to start the low-speed parking assist function, and no specific limitation is made thereto.
[0155] When the vehicle activates the low-speed parking assist function, the vehicle's infotainment system can first acquire multimodal perception information.
[0156] Referring to the embodiment shown in Figure 1, the vehicle is equipped with various sensors, including a camera sensor (hereinafter referred to as camera), a lidar sensor (hereinafter referred to as lidar), and a millimeter-wave radar sensor (hereinafter referred to as millimeter-wave radar).
[0157] The camera can include any camera used to acquire images of the environment in which the vehicle is located (e.g., pinhole camera, fisheye camera, still camera, video camera, etc.). For this purpose, the camera can be configured to detect visible light, or it can be configured to detect light from other parts of the spectrum (such as infrared or ultraviolet light). Other types of cameras are also possible. The camera can be a two-dimensional detector, or it can have three-dimensional spatial range detection capabilities. In some examples, the camera can be, for example, a distance detector configured to generate a two-dimensional image indicating the distance from the camera to several points in the environment. For this purpose, the camera can use one or more distance detection techniques. For example, the camera can be configured to use structured light technology, where the vehicle illuminates objects in the environment using a predetermined light pattern, such as a grid, and uses the camera to detect reflections from the predetermined light pattern from the objects. Based on the distortion in the reflected light pattern, the vehicle can be configured to detect the distance to points on the objects. The predetermined light pattern can include infrared light or light of other wavelengths.
[0158] LiDAR can be viewed as an object detection system that uses light sensing to detect objects in the environment in which a vehicle is located. Typically, LiDAR is an optical remote sensing technique that measures the distance to a target or other properties of the target by using light to illuminate it. As an example, LiDAR may include a laser source and / or laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For instance, LiDAR may include a laser rangefinder reflected by a rotating mirror and scans the laser around a digitized scene in one or two dimensions to acquire distance measurements at specified angular intervals. In this example, LiDAR may include components such as a light (e.g., laser) source, scanner and optical system, photodetector and receiver electronics, and a positioning and navigation system. LiDAR determines the distance to an object by scanning the laser reflected back from it, creating a 3D environmental map with accuracy up to centimeters. Millimeter-wave radar typically refers to object detection sensors with wavelengths of 1–10 mm and frequencies generally ranging from 10 GHz to 200 GHz.
[0159] Millimeter-wave radar measurements contain depth information, providing the target's distance. Secondly, due to the significant Doppler effect, millimeter-wave radar is highly sensitive to velocity, allowing direct acquisition of target speed. The velocity can be extracted by detecting the Doppler frequency shift. Currently, the two main automotive millimeter-wave radar operating frequency bands are 24GHz and 77GHz. The former, with a wavelength of approximately 1.25cm, is primarily used for short-range sensing, such as the vehicle's surroundings, blind spots, parking assistance, and lane change assistance. The latter, with a wavelength of approximately 4mm, is used for medium- to long-range measurements, such as automatic following, adaptive cruise control (ACC), and automatic emergency braking (AEB).
[0160] In this embodiment of the application, the multimodal sensing information can be obtained based on information collected by one or more of the above-mentioned sensors.
[0161] Step 402: Obtain multiple sets of display information, including the target area to be displayed and the grid granularity.
[0162] A set of display information may include the target area to be displayed and the grid granularity. The target area can refer to the range of a real-world region displayed on the central control screen. For example, in the real world, the area within 3 meters around the vehicle (i.e., the vehicle itself), centered on the vehicle, can be a 6m × 6m square planar area extending 3m forward, 3m backward, 3m to the left, and 3m to the right, with a height of 2m added to form a 6m × 6m × 2m three-dimensional space; this three-dimensional space is the target area. The length, width, and height of the target area can be represented as W (length) × H (width) × D (height). When representing the target area, the vehicle itself can be used as the origin, and the range of the target area can be represented by the three-dimensional coordinates of its eight vertices. Alternatively, the range of the target area can be represented by the three-dimensional coordinates of a single vertex of the target area, along with the length, width, and height of the target area. This application does not specifically limit the representation of the target area's range in its embodiments.
[0163] The grid granularity is based on the principle of the occupied grid. An occupied grid refers to dividing the three-dimensional space into voxels, with each voxel represented by a binary value of 0 or 1 indicating whether it is occupied. Therefore, one voxel corresponds to one grid, and the grid granularity can be used to represent the granularity of dividing the three-dimensional space into voxels. For example, if a voxel is 1m × 1m × 1m, the aforementioned 6m × 6m × 2m three-dimensional space can be divided into 6 × 6 × 2 voxels, meaning this three-dimensional space corresponds to 6 × 6 × 2 grids.
[0164] Optionally, the larger the target area, the larger the selectable raster granularity; conversely, the smaller the target area, the smaller the selectable raster granularity. This allows for the display of more raster occupancy information within a larger target area, and more detailed raster occupancy information within a smaller target area.
[0165] In one possible implementation, the target area in a set of displayed information can refer to the area within 6m around the vehicle, that is, the length, width and height of the target area are 12m×12m×2m, and the grid granularity can be 20cm.
[0166] In one possible implementation, the target area in a set of displayed information can refer to the area within 3m around the vehicle, that is, the length, width and height of the target area are 6m×6m×2m, and the grid granularity can be 5cm.
[0167] In one possible implementation, in a set of displayed information, the target area can refer to the area within 1m around the vehicle, that is, the length, width and height of the target area are 2m×2m×2m, and the grid granularity can be 2cm.
[0168] In this embodiment, the target area and the grid granularity can be pre-established (which can be used as the default recommended configuration) or the correspondence between the target area and the grid granularity can be determined in real time, without any specific limitation.
[0169] In this embodiment of the application, the vehicle-mounted system can obtain display information through the following multiple methods:
[0170] 1. The displayed information is pre-set.
[0171] Multiple sets of display information can be pre-set and stored. When the display information is needed, the vehicle's infotainment system can directly read the memory to obtain these multiple sets of display information. For example, two sets of display information can be set. In one set, the target area refers to the area within 6 meters around the vehicle, that is, the length, width, and height of the target area are 12m × 12m × 2m, and the grid grain is 20cm. In the other set of display information, the target area refers to the area within 3 meters around the vehicle, that is, the length, width, and height of the target area are 6m × 6m × 2m, and the grid grain is 5cm.
[0172] 2. The displayed information is dynamically generated based on the driver's actions on the touchscreen.
[0173] The zooming operation described above can include the driver pressing two fingers on the touchscreen and sliding them outwards or inwards to change the range of the target area and the target resolution of obstacles; or, the zooming operation can include the driver pressing two fingers on the touchscreen and rotating, panning, or dragging to change the display orientation and range of the target area. The aforementioned operations can be referenced to the following operations when a driver uses an image application: to zoom in on image details, they can press the relevant area with their thumb and forefinger and then slide the two fingers outwards; or to zoom out to view the whole image, they can press the relevant area with their thumb and forefinger and then slide the two fingers inwards; or, to change the orientation of the image content, they can press the relevant area with two fingers and then drag the image to pan or rotate it.
[0174] In this embodiment, a human-machine interface is provided for the driver. The driver can zoom in / out on an initial target area using this interface to obtain the final target area and its corresponding grid granularity. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the aforementioned sets of display information.
[0175] Step 403: Perform perception prediction based on multimodal perception information and multiple sets of display information to obtain multiple occupied grid information.
[0176] In this embodiment, multimodal perception information and multiple sets of display information can be input into a pre-trained perception prediction network to output multiple occupancy grid information.
[0177] The aforementioned perception and prediction network can be referenced from the neural network section above, and its principles and specific implementation methods can also refer to that section. The perception and prediction network can be pre-trained and downloaded to the vehicle's infotainment system. When using the perception and prediction network, the parameters in the network can be updated based on real-time input and output to make the prediction results more accurate. This allows for obtaining grid occupancy information that better reflects actual driving conditions, thus assisting the driver in driving the vehicle more safely.
[0178] For example, Figure 5 is a schematic flowchart of the perception prediction process according to an embodiment of this application. As shown in Figure 5, the perception prediction process (corresponding to the perception prediction network) includes the following steps:
[0179] (1) The multimodal perception information acquired by the vehicle system includes M pinhole images (from pinhole cameras) and N fisheye images (from fisheye cameras). Feature extraction is performed on the M pinhole images and N fisheye images respectively, and the image features are transformed from 2D perspective view (PV) to 3D space according to the pose of the camera that acquired each image.
[0180] That is, the perception prediction network includes a pinhole network and a fisheye network based on multimodal perception information; wherein, the pinhole network is used to extract the visual and spatial features of the pinhole image acquired by the pinhole camera to obtain the pinhole 3D image features, and the fisheye network is used to extract the visual and spatial features of the fisheye image acquired by the fisheye camera to obtain the fisheye 3D image features.
[0181] a. Pinhole images and fisheye images can have different resolutions, therefore they can be separated and their features extracted separately. Among the M pinhole images, features can also be extracted separately for each pinhole image based on the location of the pinhole camera (e.g., in front of, behind, or to the side of the vehicle). Similarly, among the N fisheye images, features can also be extracted separately for each fisheye image based on the location of the fisheye camera (e.g., in front of, behind, or to the side of the vehicle).
[0182] The aforementioned feature extraction of images can be performed using methods based on convolutional neural networks (CNNs), such as ResNet or Feature Map Pyramid Network (FPN), or it can be performed using methods based on transformer structures, such as the Vision Transformer.
[0183] b. Based on feature extraction, a unified PV viewpoint feature map pv_feat is obtained, with dimensions C×H×W, where C is the number of feature channels, and H and W are the image dimensions.
[0184] c. Consistent with step a, a neural network model is used to perform monocular depth estimation on the depth structure information of the image, thereby obtaining the depth distribution map of the image with dimensions D×H×W, where D is the number of categories in the depth distribution.
[0185] d. Using the aforementioned depth distribution map, perform a weighted mapping (e.g., broadcast multiplication) on the aforementioned pv_feat to obtain a depth visual feature map with dimensions C×D×H×W. Then, based on camera intrinsic and extrinsic parameters (e.g., camera position, focal length, etc.), project the features from the depth visual feature map onto different locations in 3D space to obtain a 3D feature map with dimensions C1×X×Y×Z, where X, Y, and Z represent the length, width, and height in three-dimensional space.
[0186] Following steps ab above, M pinhole images and N fisheye images can be processed to obtain their respective 3D feature maps. It should be noted that the dimensions of the 3D feature maps obtained from different frames can be exactly the same, not exactly the same, or completely different; no specific limitation is imposed.
[0187] (2) The multimodal perception information acquired by the vehicle system includes laser point cloud (from lidar) and millimeter wave point cloud (from millimeter wave radar), and feature extraction is performed on the laser point cloud and the millimeter wave point cloud respectively.
[0188] That is, the perception prediction network includes a point cloud network based on multimodal perception information; the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features.
[0189] a. For laser point clouds, voxelization is performed using the correspondence between point clouds and grids (i.e., the correspondence between 3D points in the point cloud and X, Y, Z in three-dimensional space). Each voxel can include point cloud features such as the number of points, the maximum height of 3D points, the minimum height of 3D points, and the point cloud reflection intensity. After convolution, a lidar feature map is obtained.
[0190] b. For millimeter-wave point clouds, voxelization is performed using the correspondence between point clouds and grids (i.e., the correspondence between 3D points in the point cloud and X, Y, Z in three-dimensional space), and then convolution is performed to obtain millimeter-wave radar feature maps.
[0191] c. Perform feature fusion (e.g., concatenation, addition, etc.) on the lidar feature map and the radar feature map to obtain the radar feature map with dimensions C2×X×Y×Z.
[0192] (3) Perform multimodal feature fusion on 3D feature maps, lidar feature maps and radar feature maps in the same three-dimensional space.
[0193] That is, the perception prediction network includes a multimodal fusion network based on multimodal perception information; the multimodal fusion network is used to perform weighted fusion of pinhole 3D image features, fisheye 3D image features and point cloud voxel features.
[0194] a. Weighted fusion of the 3D feature maps corresponding to M pinhole images and N fisheye images to obtain the 3D feature map of the image.
[0195] b. The radar feature map and the image 3D feature map are then fused to obtain the fused 3D spatial feature map.
[0196] The above steps correspond to the time-based feature fusion function of the perceptual prediction network. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitations are made.
[0197] (4) Further temporal alignment and spatial feature extraction are performed on the features of the fused 3D spatial feature map.
[0198] That is, the perception prediction network also includes a temporal fusion network, which is used to align the time-series 3D spatial features of multiple frames to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
[0199] a. Combine the fused 3D spatial feature map of the current frame with the 3D spatial feature map of the historical T-1 frame, and perform coordinate transformation on the 3D features of the historical frame according to the vehicle motion parameters (ego motion) so that it is in the same coordinate system as the 3D features of the current moment, thus obtaining the 3D spatial feature map of the T frame aligned to the same coordinate system.
[0200] b. Combine the 3D spatial feature maps of frame T into a 3D spatial feature map of the current frame according to the temporal sequence.
[0201] c. Use a multi-stage feature extractor on the 3D spatial feature map of the current frame to extract spatial features again, and obtain an enhanced 3D spatial feature map.
[0202] The above steps correspond to the multimodal feature fusion function of the perception prediction network based on multimodal perception information. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitation is made.
[0203] (5) Upsample the enhanced 3D spatial feature map multiple times and then fuse the image features again to perform auxiliary feature enhancement;
[0204] That is, the perceptual prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and then fuse the image features again to achieve auxiliary feature enhancement.
[0205] a. The enhanced 3D spatial feature map is upsampled multiple times to obtain multiple resolutions corresponding to the above-mentioned multiple raster granularities;
[0206] b. Upsample the 3D feature map of the image multiple times to obtain multiple resolutions corresponding to the above multiple raster granularities;
[0207] c. At the same resolution, the features of the upsampled enhanced 3D spatial feature map and the upsampled image 3D feature map are weighted and fused to obtain the 3D spatial feature map at that resolution;
[0208] (6) Based on the target region and grid granularity, cropping, sampling and weighted fusion are performed on the 3D spatial feature map at multiple scale resolutions, and position features are obtained by interpolation to achieve arbitrary resolution.
[0209] That is, the perception prediction network also includes a dynamic resolution prediction network based on display information feedback. This dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation. Then, based on the position features, it predicts occupancy and semantic information to achieve prediction at any resolution.
[0210] a. Calculate the relative coordinates of the target region in the entire 3D feature space, and crop it on the 3D spatial feature map at multiple scale resolutions based on the relative coordinates.
[0211] b. Based on the raster granularity (i.e., its corresponding resolution), generate the coordinates of all positions that need to be predicted within the target area.
[0212] c. Based on the required coordinates of the predicted location, feature sampling and weighted fusion are performed on the cropped multi-scale resolution 3D spatial feature map. This feature sampling is obtained by trilinear interpolation of adjacent features on the feature map.
[0213] (7) Perform multilayer perceptron (MLP) prediction on all obtained location features, and then output the occupancy status and semantic information of the grid at the corresponding coordinate position. The occupancy status is a binary classification prediction, that is, there are only two categories: occupied / unoccupied. The semantic information is a multi-class prediction, including road surface / moving target / other obstacles, etc.
[0214] For a single set of display information (including the target area and grid granularity), the occupancy and semantic information of all grid positions within the target area constitute the occupancy grid information corresponding to that display information. For example, for a set of display information (the target area could be the area within 6m around the vehicle, i.e., the target area's length, width, and height are 12m × 12m × 2m, and the grid granularity could be 20cm), the occupancy and semantic information of all grid positions within the 6m area around the vehicle can be perceived and predicted, and the grid occupancy can be displayed with a 20cm grid granularity; or, for another set of display information (the target area could be the area within 3m around the vehicle, i.e., the target area's length, width, and height are 6m × 6m × 2m, and the grid granularity could be 5cm), the occupancy and semantic information of all grid positions within the 3m area around the vehicle can be perceived and predicted, and the grid occupancy can be displayed with a 5cm grid granularity.
[0215] The above steps correspond to the perceptual prediction function of the perceptual prediction network based on displayed information. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitations are made here.
[0216] It should be noted that the embodiment shown in Figure 5 only describes one method of perception prediction. Other methods can also be used for perception prediction in this application embodiment, and no specific limitation is made thereto.
[0217] Step 404: Display multiple occupied grid screens on the human-computer interaction interface based on the multiple occupied grid information.
[0218] In one possible implementation, the vehicle infotainment system can display a first occupied grid image on the central control screen, the first occupied grid image corresponding to a first target area and a first grid granularity; and display a second occupied grid image, the second occupied grid image corresponding to a second target area and a second grid granularity; wherein the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0219] Optionally, the first occupied grid screen includes n1 grids corresponding to obstacles, and the second occupied grid screen includes n2 grids corresponding to obstacles, where n1 and n2 are both positive integers.
[0220] The first occupied grid screen displays the image of the first target area, which displays the grid of n1 obstacles that are perceived and predicted within the first target area at the first grid granularity; the second occupied grid screen displays the image of the second target area, which displays the grid of n2 obstacles that are perceived and predicted within the second target area at the second grid granularity.
[0221] Optionally, the grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle refer to the same obstacle. n1 obstacles include the first obstacle, and n2 obstacles include the second obstacle.
[0222] Both the first target area and the second target area are areas surrounding the vehicle. The difference between them lies in their extent. The first target area is larger than the second target area. Therefore, the first and second target areas may include the same obstacles. In other words, obstacles in the second target area may also exist in the first target area, and some obstacles in the first target area may not appear in the second target area.
[0223] To facilitate driver identification, in this embodiment, different obstacle grids can be displayed with different colors, different outlines, etc., or different obstacle grids can be marked with different information (text, numbers, etc.). This allows the driver to intuitively see the number, location, and type of obstacles in the target area, and thus take timely avoidance measures.
[0224] Based on this, in both the first and second occupied grid views, grids corresponding to the same obstacle can be identified with the same information. For example, in both views, the grids for the first and second obstacles can be displayed with the same color (e.g., yellow grids) and the same border (e.g., solid borders). Alternatively, in both views, the grids for the first and second obstacles (obstacles present in both the first and second target areas) can be marked with the same information (e.g., the same text, the same numbers, etc.). The aforementioned first and second obstacles refer to the same obstacle. This allows the driver to visually identify which obstacles are the same across different target areas, compare the distribution and distance of obstacles in different fields of view, and take timely evasive action.
[0225] Optionally, the grid of the third obstacle in the second occupied grid image is identified with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0226] To facilitate driver identification, in this embodiment, the grid of the nearest obstacle around the vehicle can be identified with specific information. For example, the grid of the nearest obstacle can be highlighted with a special color (red), or the grid of the nearest obstacle can be displayed in a flashing manner, or the nearest obstacle can be marked with text near the grid of the nearest obstacle as the nearest obstacle, or the distance between the nearest obstacle and the vehicle can be marked with numbers near the grid of the nearest obstacle. Furthermore, when the distance between the vehicle and the nearest obstacle is less than a set threshold (e.g., 0.5m), the vehicle's audio and lights will provide a synchronized warning.
[0227] It should be noted that in the embodiments of this application, other methods may also be used to display the grid of obstacles, and no specific limitation is made thereto.
[0228] For example, Figure 6 is a schematic diagram of the human-computer interaction interface of the central control screen according to an embodiment of this application. As shown in Figure 6, the human-computer interaction interface consists of a surround view image display area and a panoramic occupancy display area, wherein...
[0229] The surround view image display area mainly displays image information captured by pinhole cameras and fisheye cameras around the vehicle.
[0230] The panoramic occupancy display area consists of two regions, left and right. The left region displays a wider target area and coarse-grained occupancy grid perception information, which can provide the driver with a more complete picture of the surrounding obstacle distribution. The right region displays a smaller target area and fine-grained occupancy grid perception information, which can provide the driver with a detailed outline of nearby obstacles, thus intuitively showing the distance between the vehicle and nearby obstacles.
[0231] The aforementioned left-hand area corresponds to the aforementioned first occupied grid image, and the aforementioned right-hand area corresponds to the aforementioned second occupied grid image.
[0232] The two grid-based images mentioned above display target areas in two different ranges with different grid granularities. The wider target area can improve the driver's panoramic perception of obstacles in a large area, while the smaller target area can improve the driver's fine perception of obstacles in a close range. The combination of the two can improve the driver's perception of the surrounding environment and avoid collisions.
[0233] It should be noted that the above embodiments are described using two occupying grid screens as an example. The embodiments of this application do not specifically limit the number of occupying grid screens displayed on the central control screen. The number can be equal to or greater than two. Each occupying grid screen corresponds to one display information, that is, a target area is displayed with one grid granularity.
[0234] In one possible implementation, the human-computer interaction interface described above can be displayed on the driver's terminal device (e.g., mobile phone, tablet, etc.) so that the driver can simultaneously view the distribution of obstacles around the vehicle and the vehicle's driving status.
[0235] In this embodiment, the driver can install an assisted driving application on the terminal device, which enables interconnection between the vehicle-mounted system and the terminal device. When the vehicle-mounted system executes step 404, it can transmit the screen displayed on the vehicle-mounted system to the terminal device in real time. The driver can open the assisted driving application and see the screen displayed on the vehicle-mounted system, as shown in the embodiment in Figure 6.
[0236] In addition, the human-machine interface provided by the driver assistance application also allows users to operate the screen, such as zooming in / out, changing the screen orientation, etc.
[0237] In this embodiment, by displaying multiple occupied grid images with corresponding grid granularities on the vehicle's central control screen, the multiple occupied grid images display target areas of different ranges with different grid granularities. This can improve the driver's panoramic perception of obstacles in a large range, as well as the driver's refined perception of obstacles in a close range. The combination of the two can improve the driver's perception of the surrounding environment and avoid scratches.
[0238] In one possible implementation, the human-computer interaction interface of this application embodiment, in addition to the screen shown in the embodiment of FIG6, also includes a control for selecting a driving mode, which includes a driving mode and a parking mode. Therefore, the human-computer interaction interface is provided with two controls, "driving" and "parking".
[0239] In driving mode, the human-machine interface also includes controls for selecting auxiliary modes, which include automatic mode and manual mode. At this time, the human-machine interface has two controls: "automatic" and "manual".
[0240] In parking mode, the human-machine interface also includes controls for selecting assistance modes. These modes include automatic mode, manual mode, Automated Valet Parking (AVP) mode, and Auto Parking Assist (APA) mode. The human-machine interface displays four controls: "Automatic," "Manual," "AVP," and "APA." Automatic and manual modes can be selected in a human-driven scenario (i.e., the driver is driving the vehicle), while APA and AVP modes can be selected in an intelligent driving scenario (i.e., the driver activates autonomous driving, and the vehicle completes the parking maneuver automatically; the driver can be inside or outside the vehicle). APA mode activates the automatic parking function. During the parking process, the display logic of the panoramic view occupying the display area can refer to the automatic mode in the human-driven scenario of the embodiment shown in Figure 9.
[0241] In one embodiment, the implementation process of the assisted driving method in a scenario of low-speed passage through narrow roads is as follows:
[0242] 1. As shown in Figure 6, the driver can activate the low-speed parking assist function by pressing a button on the vehicle or by clicking an icon on the vehicle's central control screen, and then select the "driving" mode in the human-machine interface displayed on the central control screen.
[0243] 2. The vehicle processes the data collected by sensors around the vehicle (including pinhole cameras, fisheye cameras, lidar, and millimeter-wave radar).
[0244] 3. Based on the sensor data collected by the vehicle, perform real-time perception and prediction of the environment surrounding the vehicle. This embodiment of the application implements this step through a perception prediction network. The inputs to the perception prediction network include: 1) multimodal perception information, including pinhole images, fisheye images, laser point clouds, millimeter-wave point clouds, etc.; 2) multiple display information (including target area and grid granularity); the output of the perception prediction network includes multiple grid occupancy information (including the occupancy status and semantic information of all grids at all coordinate positions within the target area).
[0245] The implementation process of the perception prediction network can be referred to the embodiment shown in Figure 5, and will not be repeated here.
[0246] 4. The vehicle displays the information from multiple occupied grids on the central control screen in real time.
[0247] As shown in Figure 6, the human-machine interface consists of a surround-view image display area and a panoramic occupancy display area. The surround-view image display area mainly displays image information captured by pinhole cameras and fisheye cameras around the vehicle. The panoramic occupancy display area includes left and right regions. The left region displays a wider target area and coarse-grained occupancy grid perception information, providing the driver with a more comprehensive view of the surrounding obstacles. The right region displays a smaller target area and fine-grained occupancy grid perception information, providing the driver with detailed outlines of nearby obstacles, thus intuitively showing the distance between the vehicle and nearby obstacles.
[0248] Optionally, the low-speed parking assist function in this embodiment of the application can support the driver in adjusting different grid display viewing angles, as shown in Figure 7 (Figure 7 is a schematic diagram of switching display angles in the human-machine interface of this embodiment of the application). The driver can perform drag and rotate operations on the central control screen to adjust to a bird's-eye view (BEV) perspective. This allows the driver to experience a multi-angle view of the surrounding obstacle distribution.
[0249] 5. The low-speed parking assist function of this application embodiment can provide two ways to finely display a specific area, namely "automatic" and "manual" modes, as shown in Figure 8 (Figure 8 is a schematic diagram of the fine-scale mode of this application embodiment). The two modes can be switched in the upper right corner of the human-machine interface. The driver can click the gesture button to switch to "automatic" mode.
[0250] Optionally, the low-speed parking assist function can be set to "automatic" mode by default. "Automatic" mode corresponds to the situation in step 402 above where "the displayed information is the preset information".
[0251] In this embodiment, in "driving" and "automatic" modes, the human-machine interface of the central control screen displays the grid information of the area within 6m around the vehicle (i.e., the target area) with a coarse grid granularity of 20cm on the left side of the panoramic occupancy display area; and displays the grid information of the area within 3m around the vehicle (i.e., the target area) with a fine grid granularity of 5cm on the right side of the panoramic occupancy display area.
[0252] By viewing the obstacle occupancy grid information displayed on the central control screen, drivers can know which direction there is an obstacle in their vehicle, the distance between their vehicle and each obstacle, and where their vehicle is closest to an obstacle, etc., so as to take preventive measures in advance and avoid scratches.
[0253] 6. As shown in Figure 8, the central control screen can also display the distance between the nearest obstacle and the vehicle. This allows the driver to see more intuitively how far away the nearest obstacle is, so as to take corresponding driving measures to avoid the obstacle and avoid collision.
[0254] In this embodiment, the vehicle can display the grid occupancy status and semantic information of obstacles around the vehicle on the central control screen at multiple grid granularities. Furthermore, as the environment changes in real time during vehicle driving and parking, the fineness of the human-machine interface is adaptively adjusted, allowing the driver to take preventative measures in advance and avoid collisions.
[0255] In another embodiment, in the scenario of parking in a narrow space, the implementation process of the assisted driving method is as follows:
[0256] 1. As shown in Figure 9 (Figure 9 is a schematic diagram of the human-machine interface of the central control screen in this application embodiment), the driver activates the low-speed parking assist function by pressing the vehicle button or the icon on the vehicle's central control screen, and then selects the "parking" mode in the human-machine interface displayed on the central control screen.
[0257] Steps 2-4 of this embodiment can be referred to as steps 2-4 of the embodiment shown in Figure 6, and will not be repeated here.
[0258] 5. In the case of human driving, the low-speed parking assistance function of this application embodiment can provide two ways to finely display a specific area: "automatic" and "manual" modes, as shown in Figure 10 (Figure 10 is a schematic diagram of the fine-display mode of this application embodiment). The two modes can be switched in the upper right corner of the human-machine interface. The driver can click the gesture button to switch to "manual" mode. The "manual" mode corresponds to the situation in step 402 above where "the displayed information is dynamically generated based on the driver's operation on the touch screen".
[0259] In this embodiment, under "Parking" and "Manual" modes, as shown in Figure 10, in the initial state, the human-machine interface of the central control screen displays the occupancy grid information of the area within 6m around the vehicle (i.e., the target area) in a coarse grid granularity of 20cm in the left and right areas of the panoramic occupancy display area. The right area can be used by the driver to perform zoom operations on it.
[0260] As shown in Figures 11 and 12 (Figures 11 and 12 are schematic diagrams of zooming operations in manual mode according to an embodiment of this application), the driver can manually move and zoom the right side of the panoramic display area on the human-machine interface. The vehicle system extracts the zoomed area and the zoom factor based on the driver's operation, and uses the zoomed area as the target area. This zoom factor can be used as the basis for obtaining the grid granularity. The driver positions the target area within 4 meters of the vehicle (i.e., the target area) through the zoom operation. At this time, the human-machine interface of the central control screen displays the grid information of the area within 4 meters of the vehicle (i.e., the target area) with a coarse grid granularity of 8cm on the right side of the panoramic display area.
[0261] In this embodiment, the driver can manually determine the target area, and the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities. This allows for a more refined display of the areas of interest to the driver. Furthermore, the refinement of the human-machine interface is adaptively adjusted as the environment changes in real time during vehicle parking, allowing the driver to take preventative measures in advance to avoid scratches.
[0262] In another embodiment, the vehicle's infotainment system can display a panoramic occupancy grid image on the human-machine interface based on panoramic occupancy grid information, which is obtained based on the vehicle's historical occupancy grid information and driving trajectory.
[0263] In the AVP garage / parking lot parking scenario, the implementation process of the assisted driving method is as follows:
[0264] 1. As shown in Figure 13 (Figure 13 is a schematic diagram of the human-machine interface of the central control screen in this application embodiment), after the vehicle enters the garage / parking lot, the driver activates the low-speed parking assist function by pressing the vehicle button or the icon on the vehicle's central control screen, and then selects the "parking" mode and AVP mode in the human-machine interface displayed on the central control screen.
[0265] Steps 2-3 in this embodiment can be referred to as steps 2-3 in the embodiment shown in Figure 6, and will not be repeated here.
[0266] 4. As the vehicle moves through the garage / parking lot, multi-frame results are fused based on the predicted grid occupancy and semantic information to achieve static structure reconstruction of the panorama. This includes the following steps:
[0267] (1) Remove dynamic targets based on the predicted semantic information and retain only static structural obstacles;
[0268] (2) Based on the vehicle's pose, convert the remaining occupied grids to a unified world coordinate system;
[0269] (3) Accumulate the occupation grid information in the world coordinate system of N historical frames, perform probability weighted fusion, and obtain the fused 3D occupation grid information (corresponding to panoramic occupation grid information);
[0270] (4) Convert the 3D occupancy grid information to the current vehicle coordinate system;
[0271] (5) Transform the vehicle's historical trajectory points to the current vehicle coordinate system;
[0272] 5. Real-time display and output based on the 3D occupancy grid information in the vehicle coordinate system. Referring to the embodiments shown in Figures 7 and 8, the driver can drag and zoom the human-machine interface on the central control screen.
[0273] 6. In AVP mode, drivers can view the underground parking garage reconstruction results, vehicle driving and parking trajectory through the driver assistance application on the terminal device (such as mobile phone, tablet, etc.), and the screen can be the same as the screen on the vehicle's infotainment system.
[0274] This embodiment can perform panoramic perception of the entire garage / parking lot based on the vehicle's position and posture changes, thereby providing effective static environmental information for the driver to find a parking spot or for automatic parking.
[0275] Figure 14 is a structural schematic diagram of the driver assistance device 1400 of this application. As shown in Figure 14, the driver assistance device 1400 of this embodiment can be applied to the terminal device mentioned above. The driver assistance device 1400 may include: a data acquisition module 1401, an acquisition module 1402, a prediction module 1403, and a display module 1404.
[0276] The acquisition module 1401 is used to acquire multimodal perception information, which is acquired by various sensors installed on the vehicle; the acquisition module 1402 is used to acquire multiple sets of display information, which includes the target area to be displayed and the grid granularity; the prediction module 1403 is used to perform perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information, which describes the distribution of obstacles around the vehicle; the display module 1404 is used to display multiple occupied grid images on the human-machine interface based on the multiple occupied grid information.
[0277] In one possible implementation, the display module 1404 is specifically used to display a first occupying grid screen, the first occupying grid screen corresponding to a first target area and a first grid granularity; and to display a second occupying grid screen, the second occupying grid screen corresponding to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0278] In one possible implementation, the first occupied grid screen includes n1 grids corresponding to obstacles, and the second occupied grid screen includes n2 grids corresponding to obstacles, where n1 and n2 are both positive integers.
[0279] In one possible implementation, the grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information, the first obstacle and the second obstacle are the same obstacle, the n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
[0280] In one possible implementation, the grid of the third obstacle in the second occupied grid image is identified with specific information, the third obstacle being the one closest to the vehicle among the n2 obstacles.
[0281] In one possible implementation, the display module 1404 is further configured to display a panoramic occupancy grid image on the human-machine interface based on panoramic occupancy grid information, wherein the panoramic occupancy grid information is obtained based on the historical occupancy grid information of the vehicle's driving environment and driving trajectory.
[0282] In one possible implementation, the displayed information is pre-set information.
[0283] In one possible implementation, the displayed information is dynamically generated based on the driver's actions on the touchscreen.
[0284] In one possible implementation, the operation includes at least one of the following operations: the driver presses the touch screen with two fingers and slides the two fingers outward or inward together; or, the driver presses the touch screen with two fingers and rotates, translates, or drags it.
[0285] In one possible implementation, the human-computer interaction interface further includes a control for selecting a driving mode, which includes one or more of the following modes: driving mode or parking mode.
[0286] In one possible implementation, in the driving mode, the human-machine interface further includes a control for selecting an assistance mode, which includes one or more of the following modes: automatic mode or manual mode.
[0287] In one possible implementation, in the parking mode, the human-machine interface further includes a control for selecting an assistance mode, which includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking (AVP) mode, or automatic parking assist (APA) mode.
[0288] In one possible implementation, the prediction module 1403 is specifically used to input the multimodal perception information and the multiple sets of display information into a pre-trained perception prediction network to output the multiple occupied grid information.
[0289] In one possible implementation, the functions of the perception prediction network include: multimodal feature fusion based on multimodal perception information; temporal feature fusion; and perception prediction based on the display information.
[0290] In one possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein, the pinhole network is used to extract visual and spatial features of pinhole images acquired by a pinhole camera to obtain pinhole 3D image features, the fisheye network is used to extract visual and spatial features of fisheye images acquired by a fisheye camera to obtain fisheye 3D image features, the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features, and the multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
[0291] In one possible implementation, the perception prediction network further includes a temporal fusion network, which is used to align multi-frame 3D spatial features in a temporal order to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
[0292] In one possible implementation, the perceptual prediction network further includes a feature enhancement network, which performs multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and then fuses the image features again to achieve auxiliary feature enhancement.
[0293] In one possible implementation, the perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation. Then, based on the position features, it performs prediction of occupancy and semantic information to achieve prediction at arbitrary resolution.
[0294] In one possible implementation, the multiple sensors include pinhole cameras and fisheye cameras.
[0295] In one possible implementation, the multiple sensors also include lidar and / or millimeter-wave radar.
[0296] The apparatus in this embodiment can be used to execute the technical solution of the method embodiment shown in FIG4. Its implementation principle and technical effect are similar, and will not be described again here.
[0297] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0298] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0299] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0300] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0301] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0302] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0303] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0304] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0305] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A driving assistance method, characterized in that, include: Acquire multimodal perception information, which is collected by various sensors installed on the vehicle; Acquire multiple sets of display information, including the target area to be displayed and the grid granularity; Perception prediction is performed based on the multimodal perception information and the multiple sets of display information to obtain multiple occupancy grid information, which is used to describe the distribution of obstacles around the vehicle; Based on the multiple occupied grid information, multiple occupied grid screens are displayed on the human-computer interaction interface.
2. The method according to claim 1, characterized in that, The step of displaying multiple occupied grid images based on the multiple occupied grid information includes: Display the first occupied grid image, which corresponds to the first target area and the first grid granularity; Display a second occupied grid image, which corresponds to a second target area and a second grid granularity; The range of the first target region is larger than the range of the second target region, and the first grid granularity is larger than the second grid granularity.
3. The method according to claim 2, characterized in that, The first occupied grid screen includes grids corresponding to n1 obstacles, and the second occupied grid screen includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
4. The method according to claim 3, characterized in that, The grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
5. The method according to any one of claims 2-4, characterized in that, The grid of the third obstacle in the second occupied grid image is identified with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
6. The method according to any one of claims 1-5, characterized in that, Also includes: The panoramic occupancy grid is displayed on the human-machine interface based on the panoramic occupancy grid information, which is obtained based on the vehicle's historical occupancy grid information and driving trajectory.
7. The method according to any one of claims 1-6, characterized in that, The displayed information is pre-set.
8. The method according to any one of claims 1-6, characterized in that, The displayed information is dynamically generated based on the driver's actions on the touchscreen.
9. The method according to claim 8, characterized in that, The operation includes at least one of the following operations: The driver can press the touchscreen with two fingers and slide them outwards or inwards; or... The driver presses two fingers on the touchscreen and performs rotation, translation, or drag operations.
10. The method according to any one of claims 1-9, characterized in that, The human-computer interaction interface also includes controls for selecting driving modes, which include one or more of the following modes: driving mode or parking mode.
11. The method according to claim 10, characterized in that, In the driving mode, the human-machine interface also includes a control for selecting an auxiliary mode, which includes one or more of the following modes: automatic mode or manual mode.
12. The method according to claim 10, characterized in that, In the parking mode, the human-machine interface also includes a control for selecting an auxiliary mode, which includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking (AVP) mode, or automatic parking assist (APA) mode.
13. The method according to any one of claims 1-12, characterized in that, The step of performing perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information includes: The multimodal perception information and the multiple sets of display information are input into a pre-trained perception prediction network to output the multiple occupied grid information.
14. The method according to claim 13, characterized in that, The perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein... The pinhole network is used to extract the visual and spatial features of pinhole images acquired by a pinhole camera to obtain pinhole 3D image features. The fisheye network is used to extract the visual and spatial features of fisheye images acquired by a fisheye camera to obtain fisheye 3D image features. The point cloud network is used to extract the 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features. The multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
15. The method according to claim 13 or 14, characterized in that, The perception prediction network also includes a temporal fusion network, which is used to align multi-frame 3D spatial features in time sequence to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
16. The method according to any one of claims 13-15, characterized in that, The perception prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and then fuse the image features again to achieve auxiliary feature enhancement.
17. The method according to any one of claims 13-16, characterized in that, The perception prediction network also includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and to obtain position features through interpolation. Then, based on the position features, it performs prediction of occupancy and semantic information to achieve prediction at any resolution.
18. The method according to any one of claims 1-17, characterized in that, The various sensors include pinhole cameras and fisheye cameras.
19. The method according to claim 18, characterized in that, The various sensors also include lidar and / or millimeter-wave radar.
20. A driver assistance device, characterized in that, include: The acquisition module is used to acquire multimodal perception information, which is acquired by various sensors installed on the vehicle. The acquisition module is used to acquire multiple sets of display information, including the target area to be displayed and the grid granularity. The prediction module is used to perform perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupancy grid information, which is used to describe the distribution of obstacles around the vehicle. The display module is used to display multiple occupied grid images on the human-computer interaction interface based on the multiple occupied grid information.
21. The apparatus according to claim 20, characterized in that, The display module is specifically used to display a first occupied grid image, which corresponds to a first target area and a first grid granularity; and to display a second occupied grid image, which corresponds to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
22. The apparatus according to claim 21, characterized in that, The first occupied grid screen includes grids corresponding to n1 obstacles, and the second occupied grid screen includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
23. The apparatus according to claim 22, characterized in that, The grid of the first obstacle in the first occupied grid image and the grid of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
24. The apparatus according to any one of claims 21-23, characterized in that, The grid of the third obstacle in the second occupied grid image is identified with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
25. The apparatus according to any one of claims 20-24, characterized in that, The display module is also used to display a panoramic occupancy grid image on the human-machine interface based on the panoramic occupancy grid information, wherein the panoramic occupancy grid information is obtained based on the historical occupancy grid information of the vehicle's driving environment and driving trajectory.
26. The apparatus according to any one of claims 20-25, characterized in that, The displayed information is pre-set.
27. The apparatus according to any one of claims 20-25, characterized in that, The displayed information is dynamically generated based on the driver's actions on the touchscreen.
28. The apparatus according to claim 27, characterized in that, The operation includes at least one of the following operations: The driver can press the touchscreen with two fingers and slide them outwards or inwards; or... The driver presses two fingers on the touchscreen and performs rotation, translation, or drag operations.
29. The apparatus according to any one of claims 20-28, characterized in that, The human-computer interaction interface also includes controls for selecting driving modes, which include one or more of the following modes: driving mode or parking mode.
30. The apparatus according to claim 29, characterized in that, In the driving mode, the human-machine interface also includes a control for selecting an auxiliary mode, which includes one or more of the following modes: automatic mode or manual mode.
31. The apparatus according to claim 29, characterized in that, In the parking mode, the human-machine interface also includes a control for selecting an auxiliary mode, which includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking (AVP) mode, or automatic parking assist (APA) mode.
32. The apparatus according to any one of claims 20-31, characterized in that, The prediction module is specifically used to input the multimodal perception information and the multiple sets of display information into a pre-trained perception prediction network to output the multiple occupied grid information.
33. The apparatus according to claim 32, characterized in that, The perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein... The pinhole network is used to extract the visual and spatial features of pinhole images acquired by a pinhole camera to obtain pinhole 3D image features. The fisheye network is used to extract the visual and spatial features of fisheye images acquired by a fisheye camera to obtain fisheye 3D image features. The point cloud network is used to extract the 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features. The multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
34. The apparatus according to claim 32 or 33, characterized in that, The perception prediction network also includes a temporal fusion network, which is used to align multi-frame 3D spatial features in time sequence to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
35. The apparatus according to any one of claims 32-34, characterized in that, The perception prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and then fuse the image features again to achieve auxiliary feature enhancement.
36. The apparatus according to any one of claims 32-35, characterized in that, The perception prediction network also includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and to obtain position features through interpolation. Then, based on the position features, it performs prediction of occupancy and semantic information to achieve prediction at any resolution.
37. The apparatus according to any one of claims 20-36, characterized in that, The various sensors include pinhole cameras and fisheye cameras.
38. The apparatus according to claim 37, characterized in that, The various sensors also include lidar and / or millimeter-wave radar.
39. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-19.
40. A computer-readable storage medium, characterized in that, Includes a computer program, which, when executed on a computer, causes the computer to perform the method of any one of claims 1-19.
41. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to perform the method of any one of claims 1-19.