Assisted driving method and apparatus
By integrating multiple sensors and perception prediction networks on the vehicle, the obstacle display of different grid particle sizes is generated, which solves the problem that existing vehicle-assisted driving systems are difficult to identify close-range obstacles in complex scenarios, and improves the driver's perception ability and safety.
Patent Information
- Application Number
- PCT/CN2024/121728
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2024-09-27
- Publication Date
- 2025-07-17
AI Technical Summary
The existing on-board assisted driving functions are difficult to effectively identify and display close-range obstacles in scenarios such as narrow parking spaces, narrow road traffic, narrow road turnover, narrow road car meetings, etc., resulting in insufficient driver safety assistance capabilities.
By integrating multiple sensors on the vehicle, such as cameras, lidars and millimeter wave radars, multimodal perception information is obtained, combined with a perception prediction network, it generates occupancy grid pictures of different grid particle sizes, displays the distribution of obstacles around the vehicle, and provides panoramic and refined perception.
It improves the driver's perception of obstacles at large range and close range, helps the driver avoid scratches, and improves the safety and operation convenience of the vehicle in complex scenarios.
Smart Images

Figure CN2024121728_17072025_PF_FP_ABST
Abstract
Description
Assisted driving method and device Technical Field
[0001] The present application relates to vehicle-mounted assisted driving technology, and in particular to an assisted driving method and device. Background Art
[0002] With the advancement of engineering technology and the improvement of intelligent perception capabilities, in-vehicle assisted driving features, such as reverse lane assist and parking cameras, have become standard features on all types of vehicles. However, current in-vehicle assisted driving features struggle to provide drivers with effective and safe driving assistance in certain scenarios, such as parking in tight spaces, driving in narrow alleys, making U-turns on narrow roads, meeting other vehicles on narrow roads, and automatic emergency braking (AEB) and rear automatic emergency braking (RAEB).
[0003] Summary of the Invention
[0004] The present application provides a driving assistance method and device to enhance the driver's perception of the surrounding environment and avoid scratches.
[0005] In a first aspect, the present application provides an assisted driving method, comprising: obtaining multimodal perception information, wherein the multimodal perception information is collected by multiple sensors arranged on a vehicle; obtaining multiple sets of display information, wherein the display information includes a target area to be displayed and a grid granularity; performing perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information, wherein the occupied grid information is used to describe the distribution of obstacles around the vehicle; and displaying multiple occupied grid screens on a human-computer interaction interface based on the multiple occupied grid information.
[0006] In an embodiment of the present application, multiple occupied grid screens corresponding to multiple grid granularities are displayed on the vehicle's central control screen. The multiple occupied grid screens display target areas of different ranges with different grid granularities. This can not only enhance the driver's panoramic perception of obstacles within a large range, but also enhance the driver's refined perception of obstacles within a close range. The combination of the two can enhance the driver's perception of the surrounding environment and avoid scratches.
[0007] In the embodiment of the present application, the driver can trigger the vehicle to start the low-speed parking assist function in the following ways:
[0008] 1. The driver inputs the start command through the vehicle button or the vehicle's central control screen.
[0009] In one possible implementation, the vehicle is provided with an assisted driving mode button that the driver can press to activate the low-speed parking assist function. It should be understood that activating a target function (e.g., the low-speed parking assist function) via a button can be implemented in a variety of ways, such as by pressing a button corresponding to the target function, or by repeatedly pressing a button to cycle through the target functions, and the present embodiment does not impose any specific limitations on this.
[0010] In one possible implementation, the central control screen displays a main interface that includes an icon corresponding to the low-speed parking assist function. The driver clicks the icon to trigger the low-speed parking assist function. It should be understood that activating a target function (e.g., the low-speed parking assist function) via the central control screen (which has touchscreen capabilities) can be implemented in a variety of ways, such as directly clicking the target function icon or selecting the target function option from a menu in the vehicle's assisted driving mode, etc., and this embodiment of the present application does not specifically limit this.
[0011] 2. The driver starts the vehicle by voice input.
[0012] In one possible implementation, the driver speaks to the speaker of the vehicle computer, "Please start the low-speed parking assist function." When the vehicle computer receives and recognizes the semantics of the voice, the low-speed parking assist function is started.
[0013] It should be noted that, in addition to the above two methods, the embodiment of the present application may also use other methods to trigger the vehicle to start the low-speed parking assist function, and there is no specific limitation on this.
[0014] When the vehicle activates the low-speed parking assist function, the vehicle computer can first obtain multimodal perception information. The vehicle is equipped with multiple sensors, including a camera sensor (referred to as camera), a lidar sensor (referred to as lidar), and a millimeter-wave radar sensor (referred to as millimeter-wave radar). In the embodiment of the present application, the multimodal perception information can be obtained based on information collected by one or more of the aforementioned sensors.
[0015] A set of display information may include a target area to be displayed and a grid granularity. Among them, the target area may refer to the range of the area in the real world displayed on the central control screen. For example, in the real world, the area within 3m around the vehicle (i.e., the vehicle itself), with the vehicle (i.e., the vehicle itself) as the center, a 6m×6m square plane area 3m forward, 3m backward, 3m left and 3m right, and then superimposed with a height of 2m to form a 6m×6m×2m three-dimensional space, which is the target area. The length, width and height of the target area can be expressed as W (length) × H (width) × D (height). When representing the target area, the vehicle itself can be used as the origin, and the three-dimensional coordinates of the 8 vertices of the target area can be used to represent the range of the target area. Alternatively, the three-dimensional coordinates of a vertex of the target area and the length, width and height of the target area can be used to represent the range of the target area. The embodiment of the present application does not specifically limit the representation form of the range of the target area.
[0016] The grid granularity is based on the principle of occupancy grid. An occupancy grid divides a three-dimensional space into voxels, with each voxel represented by a binary value of 0 or 1 to indicate whether it is occupied or not. Thus, one voxel corresponds to one grid, and the grid granularity can be used to represent the granularity of the voxels in the three-dimensional space. For example, a voxel has a size of 1m×1m×1m. The above three-dimensional space of 6m×6m×2m can be divided into 6×6×2 voxels, meaning that the three-dimensional space corresponds to 6×6×2 grids.
[0017] Optionally, the larger the target area, the larger the grid granularity that can be selected, and the smaller the target area, the smaller the grid granularity that can be selected. This allows more occupied grid information to be displayed in larger target areas, while more detailed occupied grid information can be displayed in smaller target areas.
[0018] In a possible implementation, in a set of display information, the target area may refer to an area within 6 meters around the vehicle, that is, the length, width and height of the target area are 12m×12m×2m, and the grid granularity may be 20cm.
[0019] In a possible implementation, in a set of display information, the target area may refer to an area within 3 meters around the vehicle, that is, the length, width and height of the target area are 6m×6m×2m, and the grid granularity may be 5cm.
[0020] In a possible implementation, in a set of display information, the target area may refer to an area within 1 meter around the vehicle, that is, the length, width and height of the target area are 2m×2m×2m, and the grid granularity may be 2cm.
[0021] In the embodiment of the present application, the correspondence between the target area and the grid granularity can be pre-established (which can be used as a default recommended configuration), or the correspondence between the target area and the grid granularity can be determined in real time, and there is no specific limitation on this.
[0022] In the embodiment of the present application, the vehicle computer can obtain display information in the following ways:
[0023] 1. The displayed information is pre-set information.
[0024] Multiple sets of display information can be pre-set and stored. When the display information is needed, the vehicle computer can directly read the memory to obtain the multiple sets of display information. For example, two sets of display information can be set. In one set of display information, the target area is the area within 6 meters around the vehicle, that is, the length, width and height of the target area are 12m × 12m × 2m, and the grid granularity is 20cm. In the other set of display information, the target area is the area within 3 meters around the vehicle, that is, the length, width and height of the target area are 6m × 6m × 2m, and the grid granularity is 5cm.
[0025] 2. The displayed information is dynamically generated based on the driver's operations on the touch screen.
[0026] The zoom operation may include the driver pressing the touch screen with two fingers and sliding the two fingers outward or sliding the two fingers together inward, so as to change the range of the target area and the target resolution of the obstacle; or the zoom operation may include the driver pressing the touch screen with two fingers and rotating, translating or dragging the two fingers to change the display orientation and range of the target area. The aforementioned operation may refer to the driver using a picture application, in order to magnify the details of the picture, pressing the relevant position with the thumb and index finger and then sliding the two fingers outward, or zooming out the picture to view the entire picture, pressing the relevant position with the thumb and index finger and then sliding the two fingers together inward; or, in order to change the direction of viewing the picture content, pressing the relevant position with two fingers and then dragging the picture to translate or rotate the picture direction.
[0027] In this embodiment, a human-computer interaction interface is provided for the driver. The driver can use this interface to autonomously zoom in or out of the initial target area, thereby obtaining the final target area and the grid granularity corresponding to the target area. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the aforementioned multiple sets of display information.
[0028] In an embodiment of the present application, multimodal perception information and multiple sets of display information can be input into a pre-trained perception prediction network to output multiple occupancy grid information.
[0029] The aforementioned perception prediction network can refer to the relevant content of the neural network in this application, and its principles and specific implementation methods can also refer to the content therein. The perception prediction network can be pre-trained and downloaded to the vehicle computer. Then, when the perception prediction network is used, the parameters in the perception prediction network can be updated based on real-time input and output to make the perception prediction network's prediction results more accurate. This can obtain occupancy grid information that is more consistent with actual driving conditions, thereby assisting the driver in driving the vehicle more safely.
[0030] For a single set of display information (including a target area and grid granularity), the grid occupancy and semantic information for all coordinate positions within the target area constitute the grid occupancy information corresponding to that display information. For example, for one set of display information (the target area can be the area within 6 meters around the ego vehicle, i.e., the target area has a length, width, and height of 12m × 12m × 2m, and the grid granularity can be 20cm), the grid occupancy and semantic information for all coordinate positions within 6 meters around the ego vehicle can be sensed and predicted, and the grid occupancy information can be displayed at a grid granularity of 20cm. Alternatively, for another set of display information (the target area can be the area within 3 meters around the ego vehicle, i.e., the target area has a length, width, and height of 6m × 6m × 2m, and the grid granularity can be 5cm), the grid occupancy and semantic information for all coordinate positions within 3 meters around the ego vehicle can be sensed and predicted, and the grid occupancy information can be displayed at a grid granularity of 5cm.
[0031] In one possible implementation, the vehicle computer can display a first occupied grid screen on the central control screen, which corresponds to the first target area and the first grid granularity; and display a second occupied grid screen, which corresponds to the second target area and the second grid granularity; wherein the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0032] Optionally, the first occupied grid image includes grids corresponding to n1 obstacles, and the second occupied grid image includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
[0033] The first occupied grid screen displays the screen of the first target area, which displays the grid of n1 obstacles perceived and predicted within the first target area at a first grid granularity; the second occupied grid screen displays the screen of the second target area, which displays the grid of n2 obstacles perceived and predicted within the second target area at a second grid granularity.
[0034] Optionally, the grid of the first obstacle in the first occupied grid screen and the grid of the second obstacle in the second occupied grid screen are identified by the same information, the first obstacle and the second obstacle refer to the same obstacle, n1 obstacles include the first obstacle and n2 obstacles include the second obstacle.
[0035] The first target area and the second target area are both areas around the vehicle. The difference between the two is their range. The range of the first target area is larger than that of the second target area. Therefore, the first target area and the second target area may include the same obstacles. In other words, obstacles in the second target area may also exist in the first target area, and some obstacles in the first target area may not appear in the second target area.
[0036] To facilitate driver identification, in the embodiment of the present application, different obstacle grids can be displayed in different colors, different wireframes, etc., or different obstacle grids can be marked with different information (text, numbers, etc.). In this way, the driver can intuitively see the number, location, type, etc. of obstacles included in the target area, and then take timely avoidance measures.
[0037] Based on this, in the first occupied grid screen and the second occupied grid screen, the grids corresponding to the same obstacle can be identified with the same information. For example, in the first occupied grid screen and the second occupied grid screen, the grids of the first obstacle and the second obstacle are respectively displayed as the same color (for example, yellow grid) and the same wireframe (for example, solid wireframe), or, in the first occupied grid screen and the second occupied grid screen, the grids of the first obstacle and the second obstacle (obstacles existing in both the first target area and the second target area) are respectively marked with the same information (for example, the same text, the same number, etc.). The aforementioned first obstacle and the second obstacle refer to the same obstacle. In this way, the driver can intuitively see which are the same obstacles in the target areas of different ranges, and compare the distribution and distance of the obstacles under different fields of view, and then take avoidance measures in time.
[0038] Optionally, the grid of the third obstacle in the second occupied grid image is identified by specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0039] To facilitate driver identification, in an embodiment of the present application, the nearest obstacle grid around the vehicle can be marked with specific information. For example, the nearest obstacle grid can be highlighted with a special color (red), or the nearest obstacle grid can be displayed in a flashing manner, or the nearest obstacle grid can be marked with text near the nearest obstacle grid as the nearest obstacle, or the nearest obstacle grid can be marked with numbers near the nearest obstacle grid to indicate its distance from the vehicle. Furthermore, when the distance between the vehicle and the nearest obstacle is less than a set threshold (e.g., 0.5m), the vehicle's audio and lighting will synchronize warnings.
[0040] It should be noted that in the embodiment of the present application, other methods can also be used to display the grid of obstacles, and there is no specific limitation on this.
[0041] In one possible implementation, the screen of the above human-computer interaction interface can be displayed on the driver's terminal device (for example, a mobile phone, tablet, etc.) so that the driver can simultaneously view the distribution of obstacles around the vehicle and the driving conditions of the vehicle.
[0042] In an embodiment of the present application, the driver can install a driving assistance application on a terminal device, which can be used to connect the vehicle computer and the terminal device. The vehicle computer can transmit the image displayed on the vehicle computer to the terminal device in real time. When the driver opens the driving assistance application, he can see the image displayed on the vehicle computer.
[0043] In addition, the human-computer interaction interface provided by the assisted driving application also allows users to operate the screen, such as zooming in / out, changing the screen direction, etc.
[0044] In one possible implementation, the human-computer interaction interface of the embodiment of the present application also includes a control for selecting a driving mode, which includes a driving mode and a parking mode. Therefore, two controls, "driving" and "parking", are provided on the human-computer interaction interface.
[0045] Among them, in driving mode, the human-computer interaction interface also includes a control for selecting an auxiliary mode, which includes an automatic mode and a manual mode. At this time, two controls, "automatic" and "manual", are set on the human-computer interaction interface.
[0046] In parking mode, the human-computer interaction interface also includes controls for selecting auxiliary modes, which include automatic mode, manual mode, Automated Valet Parking (AVP) mode and Auto Parking Asist (APA) mode. At this time, four controls, "Automatic", "Manual", "AVP" and "APA", are set on the human-computer interaction interface. Automatic mode and manual mode can be selected under human driving conditions (that is, the driver drives the vehicle), and APA mode and AVP mode can be selected under intelligent driving conditions (that is, the driver turns on automatic driving, and the vehicle completes the parking instructions by itself. At this time, the driver can be in the car or outside the car). The APA mode can start the automatic parking function. During parking, the display logic of the panoramic display area can refer to the automatic mode under human driving conditions.
[0047] Optionally, the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities. As the environment changes in real time during parking, the vehicle can adaptively adjust the level of sophistication of the human-computer interaction interface, allowing the driver to take preventive measures in advance to avoid scratches.
[0048] Optionally, the driver can manually determine the target area, and the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities, allowing the driver to display the area of interest in detail. In addition, as the environment changes in real time during parking, the vehicle can adaptively adjust the level of detail of the human-computer interaction interface, allowing the driver to take preventive measures in advance to avoid scratches.
[0049] Optionally, the entire garage / parking lot can be perceived panoramically based on the vehicle's posture changes, thereby providing the driver with effective full-scene static environment information for parking or automatic parking.
[0050] In the second aspect, the present application provides an assisted driving device, including: an acquisition module for acquiring multimodal perception information, wherein the multimodal perception information is acquired by multiple sensors arranged on the vehicle; an acquisition module for acquiring multiple sets of display information, wherein the display information includes a target area to be displayed and a grid granularity; a prediction module for performing perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupied grid information, wherein the occupied grid information is used to describe the distribution of obstacles around the vehicle; a display module for displaying multiple occupied grid screens on a human-computer interaction interface based on the multiple occupied grid information.
[0051] In one possible implementation, the display module is specifically used to display a first occupied grid screen, wherein the first occupied grid screen corresponds to a first target area and a first grid granularity; and display a second occupied grid screen, wherein the second occupied grid screen corresponds to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0052] In a possible implementation, the first occupied grid picture includes grids corresponding to n1 obstacles, and the second occupied grid picture includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
[0053] In one possible implementation, the grid of the first obstacle in the first occupied grid screen and the grid of the second obstacle in the second occupied grid screen are identified by the same information, the first obstacle and the second obstacle are the same obstacle, the n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
[0054] In a possible implementation, the grid of the third obstacle in the second occupied grid image is identified by specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0055] In one possible implementation, the display module is also used to display a panoramic occupancy grid screen on the human-computer interaction interface based on the panoramic occupancy grid information, and the panoramic occupancy grid information is obtained based on the historical occupancy grid information and driving trajectory of the vehicle's driving environment.
[0056] In a possible implementation manner, the display information is preset information.
[0057] In a possible implementation, the display information is dynamically generated based on the driver's operation on the touch screen.
[0058] In one possible implementation, the operation includes at least one of the following operations: the driver presses the touch screen with two fingers and slides the two fingers apart outward or slides the two fingers together inward; or the driver presses the touch screen with two fingers and rotates, translates or drags.
[0059] In a possible implementation, the human-computer interaction interface further includes a control for selecting a driving mode, where the driving mode includes one or more of the following modes: a driving mode or a parking mode.
[0060] In a possible implementation, in the driving mode, the human-computer interaction interface further includes a control for selecting an auxiliary mode, where the auxiliary mode includes one or more of the following modes: an automatic mode or a manual mode.
[0061] In one possible implementation, in the parking mode, the human-computer interaction interface also includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode, or automatic parking assist APA mode.
[0062] In a possible implementation, the prediction module is specifically configured to input the multimodal perception information and the multiple sets of display information into a pre-trained perception prediction network to output the multiple occupancy grid information.
[0063] In a possible implementation, the functions implemented by the perception prediction network include: multimodal feature fusion based on multimodal perception information; feature fusion based on time series; and perception prediction based on the display information.
[0064] In one possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein the pinhole network is used to extract visual and spatial features of the pinhole image captured by the pinhole camera to obtain pinhole 3D image features, the fisheye network is used to extract visual and spatial features of the fisheye image captured by the fisheye camera to obtain fisheye 3D image features, the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features, and the multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
[0065] In one possible implementation, the perception prediction network also includes a temporal fusion network, which is used to align the 3D spatial features of multiple frames in time sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.
[0066] In one possible implementation, the perception prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and to fuse image features again to achieve auxiliary feature enhancement.
[0067] In one possible implementation, the perception prediction network also includes a dynamic resolution prediction network based on the display information feedback, and the dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction of arbitrary resolution.
[0068] In one possible implementation, the multiple sensors include a pinhole camera and a fisheye camera.
[0069] In one possible implementation, the multiple sensors further include a laser radar and / or a millimeter-wave radar.
[0070] In a third aspect, the present application provides an electronic device comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement a method as described in any one of the above-mentioned first aspects.
[0071] In a fourth aspect, the present application provides a computer-readable storage medium comprising a computer program, wherein when the computer program is executed on a computer, the computer is enabled to perform any of the methods described in the first aspect.
[0072] In a fifth aspect, the present application provides a computer program product, which includes computer program code. When the computer program code is run on a computer, the computer executes any one of the methods in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] FIG1 is an exemplary functional block diagram of a vehicle 100 according to an embodiment of the present application;
[0074] FIG2 is an exemplary functional block diagram of a vehicle-mounted assisted driving system according to an embodiment of the present application;
[0075] FIG3 is a software structure block diagram of the vehicle computer according to an embodiment of the present application;
[0076] FIG4 is a flow chart of a process 400 of a driving assistance method according to an embodiment of the present application;
[0077] FIG5 is a schematic diagram of a flow chart of a perception prediction process according to an embodiment of the present application;
[0078] FIG6 is a schematic diagram of a human-computer interaction interface of a central control screen according to an embodiment of the present application;
[0079] FIG7 is a schematic diagram of switching display viewing angles of a human-computer interaction interface according to an embodiment of the present application;
[0080] FIG8 is a schematic diagram of a refinement mode according to an embodiment of the present application;
[0081] FIG9 is a schematic diagram of a human-computer interaction interface of a central control screen according to an embodiment of the present application;
[0082] FIG10 is a schematic diagram of a refinement mode according to an embodiment of the present application;
[0083] 11 and 12 are schematic diagrams of zoom operations in manual mode according to an embodiment of the present application;
[0084] FIG13 is a schematic diagram of a human-computer interaction interface of a central control screen according to an embodiment of the present application;
[0085] FIG14 is a schematic structural diagram of a driving assistance device 1400 of the present application. DETAILED DESCRIPTION
[0086] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0087] The terms "first," "second," and the like in the description, embodiments, claims, and drawings of this application are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or order. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0088] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0089] Before describing the technical solution of the embodiment of the present application, the vehicle of the embodiment of the present application will first be described with reference to the accompanying drawings.
[0090] FIG1 is an exemplary functional block diagram of a vehicle 100 according to an embodiment of the present application. As shown in FIG1 , components coupled to or included in the vehicle 100 may include a propulsion system 110, a sensor system 120, a control system 130, a peripheral device 140, a power supply 150, a computing device 160, and a driver interface 170. The components of the vehicle 100 may be configured to operate in a manner interconnected with each other and / or with other components coupled to each system. For example, the power supply 150 may provide power to all components of the vehicle 100. The computing device 160 may be configured to receive data from the propulsion system 110, the sensor system 120, the control system 130, and the peripheral device 140 and to control them. The computing device 160 may also be configured to generate an image display on the driver interface 170 and to receive input from the driver interface 170.
[0091] It should be noted that in other examples, the vehicle 100 may include more, fewer, or different systems, and each system may include more, fewer, or different components. In addition, the systems and components shown may be combined or divided in any manner, and this application does not specifically limit this.
[0092] Computing device 160 may include a processor 161, a transceiver 162, and a memory 163. Computing device 160 may be a controller or part of a controller of vehicle 100. Memory 163 may store instructions 1631 executed by processor 161 to execute various functional applications and data processing of vehicle 100. It may also store data generated during use of vehicle 100 (e.g., map data 1632), an operating system (e.g., an embedded operating system such as Android, Apple Mobile Platform (iOS), Microsoft Windows, or a UNIX-like operating system (Linux)), and applications required for at least one function. Processor 161 included in computing device 160 may include one or more general-purpose processors and / or one or more specialized processors (e.g., an image processor, a digital signal processor, etc.). To the extent processor 161 includes more than one processor, such processors may operate individually or in combination. Computing device 160 may implement functions for controlling vehicle 100 based on input received through driver interface 170. Transceiver 162 facilitates communication between computing device 160 and various systems. The memory 163, in turn, may include one or more volatile storage components and / or one or more non-volatile storage components, such as optical, magnetic, and / or organic storage devices, and the memory 163 may be fully or partially integrated with the processor 161. The memory 163 may contain instructions 1631 (e.g., program logic) executable by the processor 161 to perform various vehicle functions, including any of the functions or methods described herein.
[0093] The propulsion system 110 can provide power for the vehicle 100. As shown in FIG1 , the propulsion system 110 may include an engine 114, an energy source 113, a transmission 112, and wheels / tires 111. Furthermore, the propulsion system 110 may additionally or alternatively include other components than those shown in FIG1 . This application does not impose any specific limitations on this.
[0094] Sensor system 120 may include several sensors for sensing information about the environment in which vehicle 100 is located. As shown in FIG1 , the sensors of sensor system 120 include a global positioning system (GPS) 126, an inertial measurement unit (IMU) 125, a lidar sensor 124, a camera sensor 123, a millimeter-wave radar sensor 122, and an actuator 121 for modifying the position and / or orientation of the sensors. GPS 126 may be any sensor used to estimate the geographic location of vehicle 100. To this end, GPS 126 may include a transceiver that estimates the position of vehicle 100 relative to the Earth based on satellite positioning data. In an example, computing device 160 may be configured to use GPS 126 in conjunction with map data 1632 to estimate the path traveled by vehicle 100. IMU 125 may be configured to sense changes in the position and orientation of vehicle 100 based on inertial acceleration or any combination thereof. In some examples, the combination of sensors in IMU 125 may include, for example, an accelerometer and a gyroscope. Other combinations of sensors in IMU 125 are also possible. The lidar sensor 124 can be considered an object detection system that uses light sensing to detect objects in the environment in which the vehicle 100 is located. The lidar sensor 124 is typically an optical remote sensing technology that can measure the distance to a target or other properties of the target by illuminating the target with light. By way of example, the lidar sensor 124 can include a laser source and / or laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For example, the lidar sensor 124 can include a laser rangefinder reflected by a rotating mirror and scan the laser in one or two dimensions around a digitized scene, thereby collecting distance measurements at specified angular intervals. In an example, the lidar sensor 124 can include components such as a light (e.g., laser) source, a scanner and optical system, light detectors and receiver electronics, and a positioning and navigation system. The lidar sensor 124 determines the distance to an object by scanning the laser light reflected from the object, and can form a 3D image of the environment with up to centimeter-level accuracy. The camera sensor 123 can include any camera (e.g., a pinhole camera, a fisheye camera, a still camera, a video camera, etc.) for capturing images of the environment in which the vehicle 100 is located. To this end, the camera sensor 123 can be configured to detect visible light, or can be configured to detect light from other parts of the spectrum (such as infrared light or ultraviolet light). Other types of camera sensors 123 are also possible. The camera sensor 123 can be a two-dimensional detector, or can have three-dimensional spatial range detection capabilities. In some examples, the camera sensor 123 can be, for example, a distance detector that is configured to generate a two-dimensional image indicating the distance from the camera sensor 123 to several points in the environment.To this end, the camera sensor 123 can use one or more distance detection technologies. For example, the camera sensor 123 can be configured to use structured light technology, in which the vehicle 100 uses a predetermined light pattern, such as a grid or checkerboard pattern, to illuminate objects in the environment, and uses the camera sensor 123 to detect the reflection of the predetermined light pattern from the object. Based on the distortion in the reflected light pattern, the vehicle 100 can be configured to detect the distance of a point on the object. The predetermined light pattern may include infrared light or light of other wavelengths. Millimeter-Wave Radar sensor 122 generally refers to an object detection sensor with a wavelength of 1 to 10 mm and a frequency range of approximately 10 GHz to 200 GHz. The measurement value of the millimeter-wave radar sensor 122 has depth information and can provide the distance to the target; secondly, because the millimeter-wave radar sensor 122 has a significant Doppler effect and is very sensitive to speed, the speed of the target can be directly obtained, and the speed of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream automotive millimeter-wave radar application frequency bands are 24GHz and 77GHz. The former has a wavelength of about 1.25cm and is mainly used for short-range perception, such as the surrounding environment of the vehicle body, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4mm and is used for medium and long-range measurement, such as automatic following, adaptive cruise control (ACC), and automatic emergency braking (AEB).
[0095] The sensor system 120 may also include additional sensors, including, for example, sensors that monitor internal systems of the vehicle 100 (e.g., an O2 monitor, a fuel gauge, an oil temperature, etc.). The sensor system 120 may also include other sensors. This application does not make specific limitations on this.
[0096] The control system 130 may be configured to control the operation of the vehicle 100 and its components. To this end, the control system 130 may include a steering unit 136, a throttle 135, a brake unit 134, a sensor fusion algorithm 133, a computer vision system 132, and a navigation / route control system 131. The control system 130 may additionally or alternatively include other components in addition to those shown in FIG. This application does not impose any specific limitations on this.
[0097] Peripheral devices 140 may be configured to allow vehicle 100 to interact with external sensors, other vehicles, and / or the driver. To this end, peripheral devices 140 may include, for example, a lighting system 145, a wireless communication system 144, a touch screen 143, a microphone 142, and / or a speaker 141. Peripheral devices 140 may additionally or alternatively include other components in addition to those shown in FIG. This application does not impose any specific limitations on this.
[0098] Power source 150 can be configured to provide power to some or all components of vehicle 100. To this end, power source 150 can include, for example, a rechargeable lithium-ion or lead-acid battery. In some examples, one or more battery packs can be configured to provide power. Other power source materials and configurations are also possible. In some examples, power source 150 and energy source 113 can be implemented together, as in some all-electric vehicles.
[0099] The components of the vehicle 100 can be configured to work in an interconnected manner with other components within and / or outside of their respective systems. To this end, the components and systems of the vehicle 100 can be communicatively linked together through a system bus, a network, and / or other connection mechanisms.
[0100] Figure 2 is an exemplary functional block diagram of the vehicle-mounted assisted driving system of an embodiment of the present application. As shown in Figure 2, components coupled to or included in the vehicle-mounted assisted driving system may include a computing unit, a sensor, a central control screen, a lighting system, and an audio system. Among them, the computing unit corresponds to the control system 130 in the embodiment shown in Figure 1, the sensor corresponds to the sensor system 120 in the embodiment shown in Figure 1, mainly involving camera sensors 123 (including pinhole cameras, fisheye cameras), millimeter-wave radar sensors 122, and lidar sensors 124, the central control screen corresponds to the touch screen 143 in the embodiment shown in Figure 1, providing the driver with an interface / interface for human-computer interaction, the lighting system corresponds to the lighting system 145 in the embodiment shown in Figure 1, and the audio system corresponds to the speaker 141 in the embodiment shown in Figure 1.
[0101] The vehicle-mounted assisted driving system in the embodiment of the present application may also be referred to as a vehicle computer, a vehicle central control system (abbreviated as central control), etc., without any specific limitation.
[0102] FIG3 is a block diagram of the software structure of the vehicle computer according to an embodiment of the present application.
[0103] The layered architecture of the vehicle computer divides the software into several layers, each with clear roles and divisions of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0104] The application layer can include a series of application packages.
[0105] As shown in FIG3 , the application package may include applications such as calls, maps, navigation, WLAN, Bluetooth, music, video, and assisted driving.
[0106] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0107] As shown in FIG3 , the application framework layer may include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, and the like.
[0108] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0109] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, maps, audio, and calls made and received.
[0110] The view system includes visual controls, such as those that display text and images. The view system is used to build applications. A user interface can consist of one or more views. For example, a user interface for a notification icon might include a view that displays text and a view that displays an image.
[0111] The phone manager is used to provide communication functions for the vehicle computer, such as the management of call status (including answering, hanging up, etc.).
[0112] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0113] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without driver interaction. For example, the Notification Manager can be used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the system's top status bar as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, audible alerts, and flashing indicators.
[0114] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0115] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0116] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0117] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0118] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0119] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0120] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0121] A 2D graphics engine is a drawing engine for 2D drawings.
[0122] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0123] It is understood that the components included in the system framework layer, system library, and runtime layer shown in Figure 3 do not constitute a specific limitation on the vehicle computer. In other embodiments of the present application, the vehicle computer may include more or fewer components than shown, or may combine or split some components, or arrange the components differently.
[0124] Since the embodiments of the present application involve the application of neural networks, in order to facilitate understanding, some nouns or terms used in the embodiments of the present application are explained below, and these nouns or terms are also considered part of the content of the invention.
[0125] (1) Neural Network
[0126] A neural network (NN) is a machine learning model. A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:
[0127] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s is x s The weight of the neural unit, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0128] (2) Deep Neural Networks
[0129] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. The "many" here does not have a specific metric. Based on the position of different layers in a DNN, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript corresponds to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).
[0130] (3) Convolutional Neural Network
[0131] A convolutional neural network (CNN) is a deep neural network with a convolutional structure and a deep learning architecture. A deep learning architecture involves multiple levels of learning at different levels of abstraction using machine learning algorithms. As a deep learning architecture, a CNN is a feed-forward artificial neural network in which individual neurons respond to input images. A CNN consists of a feature extractor consisting of convolutional and pooling layers. The feature extractor can be thought of as a filter, and the convolution process can be thought of as convolving an input image or feature map with a trainable filter.
[0132] A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution on the input signal. A convolutional layer can include multiple convolution operators, also known as kernels. In image processing, a convolution operator acts as a filter that extracts specific information from the input image matrix. A convolution operator is essentially a weight matrix, which is usually predefined. During the convolution operation, the weight matrix is typically applied horizontally to the input image, pixel by pixel (or two pixels by two pixels, depending on the stride), to extract specific features from the image. The size of the weight matrix should be proportional to the image size. It is important to note that the depth dimension of the weight matrix is the same as the depth dimension of the input image; during the convolution operation, the weight matrix extends across the entire depth of the input image. Therefore, convolution with a single weight matrix produces a convolved output with a single depth dimension. However, in most cases, multiple weight matrices of the same size (row × column) are applied instead. This is known as multiple homogeneous matrices. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. The dimension here can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors in the image, and yet another weight matrix is used to blur unwanted noise in the image. The multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these multiple weight matrices of the same size are also the same size. The extracted feature maps of the same size are then merged to form the output of the convolution operation. In practical applications, the weight values in these weight matrices require extensive training. The weight matrices formed by the trained weight values can be used to extract information from the input image, allowing the convolutional neural network to make accurate predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layers often extract more general features, which can also be called low-level features. As the depth of the convolutional neural network increases, the features extracted by subsequent convolutional layers become increasingly complex, such as high-level semantic features. Features with higher semantics are more applicable to the problem being solved.
[0133] Because it's often necessary to reduce the number of trainable parameters, pooling layers are often periodically introduced after convolutional layers. This can be done in a single convolutional layer followed by a pooling layer, or in a multi-layered system followed by one or more pooling layers. In image processing, the sole purpose of a pooling layer is to reduce the spatial size of an image. Pooling layers can include average pooling and / or max pooling operators, which are used to downsample the input image to produce a smaller image. The average pooling operator calculates the average value of pixel values within a specific range, producing the average pooling result. The max pooling operator takes the pixel with the largest value within a specific range as the max pooling result. Furthermore, just as the size of the weight matrix used in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The output image size after processing by a pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding subregion of the image input to the pooling layer.
[0134] After being processed by the convolution layer / pooling layer, the convolutional neural network is still not sufficient to output the required output information. Because as mentioned above, the convolution layer / pooling layer only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolutional neural network needs to use the neural network layer to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer may include multiple hidden layers, and the parameters contained in the multiple hidden layers can be pre-trained based on relevant training data of a specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0135] Optionally, after the multiple hidden layers in the neural network layer, an output layer of the entire convolutional neural network is also included. The output layer has a loss function similar to the classification cross entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the backpropagation will begin to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0136] (4) Recurrent Neural Network
[0137] Recurrent neural networks (RNNs) are designed to process sequential data. In traditional neural network models, layers are fully connected, from the input layer to the hidden layer to the output layer, while nodes within each layer are disconnected. While these conventional neural networks have solved many difficult problems, they are still inadequate for many others. For example, to predict the next word in a sentence, you generally need to use the previous words, as the previous and next words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is dependent on the previous output. Specifically, the network memorizes previous information and applies it to the calculation of the current output. This means that nodes within the hidden layer are no longer disconnected but connected, and the input to a hidden layer includes not only the output of the input layer but also the output of the previous hidden layer. In theory, RNNs can process sequence data of any length. Training an RNN is similar to training a traditional CNN or DNN. This approach also uses the backpropagation algorithm, but with one key difference: if the RNN is expanded, its parameters, such as W, are shared; this is not the case with traditional neural networks, as in the example above. Furthermore, when using gradient descent, the output of each step depends not only on the state of the network at the current step but also on the state of the network at several previous steps. This learning algorithm is called backpropagation through time (BPTT).
[0138] Given the existence of convolutional neural networks, why do we still need recurrent neural networks? The reason is simple. Convolutional neural networks assume that elements are independent of each other, and that inputs and outputs are also independent, such as cats and dogs. However, in the real world, many elements are interconnected, such as the changes in stock prices over time. Or, for example, someone says, "I love traveling, and my favorite place is Yunnan. I must visit it someday." Humans should know to fill in the blank with "Yunnan." This is because humans make inferences based on context, but how can machines do this? RNNs were invented. RNNs are designed to give machines the ability to remember, like humans. Therefore, the output of an RNN depends on both the current input and historical memory.
[0139] (5) Loss function
[0140] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, before the first update, there is usually an initialization process, which pre-configures the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss.
[0141] (6) Backpropagation algorithm
[0142] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.
[0143] (7) Generative Adversarial Networks
[0144] Generative adversarial networks (GANs) are a type of deep learning model. They consist of at least two modules: a generative model and a discriminative model. These two modules learn from each other through interaction to produce better outputs. Both the generative and discriminative models can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of a GAN is as follows: For example, consider a GAN that generates images. Suppose there are two networks, G (Generator) and D (Discriminator). G is the image generator network, which receives random noise z and generates an image from it, denoted as G(z). D is the discriminator network, which determines whether an image is "real." Its input parameter is x, representing an image. Its output, D(x), represents the probability that x is real. A value of 1 indicates a 100% probability of authenticity, while a value of 0 indicates a high probability of non-authenticity. During the training of this generative adversarial network, the goal of the generative network G is to generate realistic images as much as possible to deceive the discriminative network D, while the goal of the discriminative network D is to distinguish the images generated by G from real images as much as possible. This creates a dynamic "game" between G and D, which is the "game" in "generative adversarial network." Ultimately, under ideal conditions, G can generate images G(z) that are sufficiently realistic, while D has difficulty determining whether the images generated by G are real, i.e., D(G(z)) = 0.5. This results in an excellent generative model G that can be used to generate images.
[0145] In complex road conditions and driving processes, there are some scenarios that require high driving skills or are difficult to drive, such as parking in narrow parking spaces, passing narrow roads, turning around on narrow roads, meeting other vehicles on narrow roads, automatic emergency braking (AEB), rear automatic emergency braking (RAEB), etc. For the aforementioned scenarios, the auxiliary systems on some vehicles (such as 360-degree panoramic views) may not be able to effectively identify and display close obstacles (including objects, pedestrians, other vehicles, etc.), and thus cannot provide the driver with close obstacle information, making it difficult to provide the driver with good and safe driving assistance capabilities. In order to solve this problem, the present application provides an assisted driving method and device.
[0146] FIG4 is a flow chart of process 400 of the assisted driving method provided in an embodiment of the present application. Process 400 may be executed by the vehicle 100 described above (particularly the vehicle computer in the vehicle). Process 400 may be used for the low-speed parking assistance function in vehicle assisted driving. Process 400 is described as a series of steps or operations. It should be understood that process 400 may be executed in various orders and / or occur simultaneously, and is not limited to the execution order shown in FIG4. Process 400 may include:
[0147] Step 401: Acquire multimodal sensing information, where the multimodal sensing information is collected by a variety of sensors installed on the vehicle.
[0148] In the embodiment of the present application, the driver can trigger the vehicle to start the low-speed parking assist function in the following ways:
[0149] 1. The driver inputs the start command through the vehicle button or the vehicle's central control screen.
[0150] In one possible implementation, the vehicle is provided with an assisted driving mode button that the driver can press to activate the low-speed parking assist function. It should be understood that activating a target function (e.g., the low-speed parking assist function) via a button can be implemented in a variety of ways, such as by pressing a button corresponding to the target function, or by repeatedly pressing a button to cycle through the target functions, and the present embodiment does not impose any specific limitations on this.
[0151] In one possible implementation, the central control screen displays a main interface that includes an icon corresponding to the low-speed parking assist function. The driver clicks the icon to trigger the low-speed parking assist function. It should be understood that activating a target function (e.g., the low-speed parking assist function) via the central control screen (which has touchscreen capabilities) can be implemented in a variety of ways, such as directly clicking the target function icon or selecting the target function option from a menu in the vehicle's assisted driving mode, etc., and this embodiment of the present application does not specifically limit this.
[0152] 2. The driver starts the vehicle by voice input.
[0153] In one possible implementation, the driver speaks to the speaker of the vehicle computer, "Please start the low-speed parking assist function." When the vehicle computer receives and recognizes the semantics of the voice, the low-speed parking assist function is started.
[0154] It should be noted that, in addition to the above two methods, the embodiment of the present application may also use other methods to trigger the vehicle to start the low-speed parking assist function, and there is no specific limitation on this.
[0155] When the vehicle starts the low-speed parking assist function, the vehicle computer can first obtain multimodal perception information.
[0156] 1 , a variety of sensors are provided in the vehicle, including a camera sensor (referred to as camera), a laser radar sensor (referred to as laser radar), and a millimeter wave radar sensor (referred to as millimeter wave radar), wherein:
[0157] The camera may include any camera for acquiring images of the environment in which the vehicle is located (e.g., a pinhole camera, a fisheye camera, a still camera, a video camera, etc.). To this end, the camera may be configured to detect visible light, or may be configured to detect light from other parts of the spectrum (such as infrared light or ultraviolet light). Other types of cameras are also possible. The camera may be a two-dimensional detector, or may have a three-dimensional spatial range detection capability. In some examples, the camera may be, for example, a distance detector that is configured to generate a two-dimensional image indicating the distance from the camera to several points in the environment. To this end, the camera may use one or more distance detection technologies. For example, the camera may be configured to use structured light technology, in which the vehicle illuminates an object in the environment using a predetermined light pattern, such as a grid, and uses the camera to detect reflections of the predetermined light pattern from the object. Based on the distortion in the reflected light pattern, the vehicle may be configured to detect the distance to a point on the object. The predetermined light pattern may include infrared light or light of other wavelengths.
[0158] LiDAR can be considered an object detection system that uses light sensing to detect objects in the environment in which the vehicle is located. LiDAR is an optical remote sensing technology that can measure the distance to a target or other properties of a target by illuminating the target with light. As an example, a LiDAR may include a laser source and / or a laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For example, a LiDAR may include a laser rangefinder that is reflected by a rotating mirror and scans the laser in one or two dimensions around a digitized scene, thereby collecting distance measurements at specified angular intervals. In an example, a LiDAR may include components such as a light (e.g., laser) source, a scanner and optical system, a light detector and receiver electronics, and a positioning and navigation system. LiDAR determines the distance to an object by scanning the laser reflected from an object, and can form a 3D environment map with an accuracy of up to centimeters. Millimeter wave radar generally refers to an object detection sensor with a wavelength of 1 to 10 mm and a frequency range of approximately 10 GHz to 200 GHz.
[0159] Millimeter-wave radar measurements provide depth information and can provide target distance. Furthermore, due to its pronounced Doppler effect, millimeter-wave radar is highly sensitive to velocity and can directly determine the target's velocity. This velocity can be extracted by detecting the Doppler shift. Currently, the two mainstream automotive millimeter-wave radar frequency bands are 24 GHz and 77 GHz. The former, with a wavelength of approximately 1.25 cm, is primarily used for short-range sensing, such as vehicle surroundings, blind spots, parking assistance, and lane change assistance. The latter, with a wavelength of approximately 4 mm, is used for medium- and long-range measurements, such as automatic following, adaptive cruise control (ACC), and automatic emergency braking (AEB).
[0160] In an embodiment of the present application, multimodal perception information can be obtained based on information collected by one or more of the above-mentioned sensors.
[0161] Step 402: Acquire multiple sets of display information, where the display information includes a target area to be displayed and a grid granularity.
[0162] A set of display information may include a target area to be displayed and a grid granularity. Among them, the target area may refer to the range of the area in the real world displayed on the central control screen. For example, in the real world, the area within 3m around the vehicle (i.e., the vehicle itself), with the vehicle (i.e., the vehicle itself) as the center, a 6m×6m square plane area 3m forward, 3m backward, 3m left and 3m right, and then superimposed with a height of 2m to form a 6m×6m×2m three-dimensional space, which is the target area. The length, width and height of the target area can be expressed as W (length) × H (width) × D (height). When representing the target area, the vehicle itself can be used as the origin, and the three-dimensional coordinates of the 8 vertices of the target area can be used to represent the range of the target area. Alternatively, the three-dimensional coordinates of a vertex of the target area and the length, width and height of the target area can be used to represent the range of the target area. The embodiment of the present application does not specifically limit the representation form of the range of the target area.
[0163] The grid granularity is based on the principle of occupancy grid. An occupancy grid divides a three-dimensional space into voxels, with each voxel represented by a binary value of 0 or 1 to indicate whether it is occupied or not. Thus, one voxel corresponds to one grid, and the grid granularity can be used to represent the granularity of the voxels in the three-dimensional space. For example, a voxel has a size of 1m×1m×1m. The above three-dimensional space of 6m×6m×2m can be divided into 6×6×2 voxels, meaning that the three-dimensional space corresponds to 6×6×2 grids.
[0164] Optionally, the larger the target area, the larger the grid granularity that can be selected, and the smaller the target area, the smaller the grid granularity that can be selected. This allows more occupied grid information to be displayed in larger target areas, while more detailed occupied grid information can be displayed in smaller target areas.
[0165] In a possible implementation, in a set of display information, the target area may refer to an area within 6 meters around the vehicle, that is, the length, width and height of the target area are 12m×12m×2m, and the grid granularity may be 20cm.
[0166] In a possible implementation, in a set of display information, the target area may refer to an area within 3 meters around the vehicle, that is, the length, width and height of the target area are 6m×6m×2m, and the grid granularity may be 5cm.
[0167] In a possible implementation, in a set of display information, the target area may refer to an area within 1 meter around the vehicle, that is, the length, width and height of the target area are 2m×2m×2m, and the grid granularity may be 2cm.
[0168] In the embodiment of the present application, the correspondence between the target area and the grid granularity can be pre-established (which can be used as a default recommended configuration), or the correspondence between the target area and the grid granularity can be determined in real time, and there is no specific limitation on this.
[0169] In the embodiment of the present application, the vehicle computer can obtain display information in the following ways:
[0170] 1. The displayed information is pre-set information.
[0171] Multiple sets of display information can be pre-set and stored. When the display information is needed, the vehicle computer can directly read the memory to obtain the multiple sets of display information. For example, two sets of display information can be set. In one set of display information, the target area is the area within 6 meters around the vehicle, that is, the length, width and height of the target area are 12m × 12m × 2m, and the grid granularity is 20cm. In the other set of display information, the target area is the area within 3 meters around the vehicle, that is, the length, width and height of the target area are 6m × 6m × 2m, and the grid granularity is 5cm.
[0172] 2. The displayed information is dynamically generated based on the driver's operations on the touch screen.
[0173] The zoom operation may include the driver pressing the touch screen with two fingers and sliding the two fingers outward or sliding the two fingers together inward, so as to change the range of the target area and the target resolution of the obstacle; or the zoom operation may include the driver pressing the touch screen with two fingers and rotating, translating or dragging the two fingers to change the display orientation and range of the target area. The aforementioned operation may refer to the driver using a picture application, in order to magnify the details of the picture, pressing the relevant position with the thumb and index finger and then sliding the two fingers outward, or zooming out the picture to view the entire picture, pressing the relevant position with the thumb and index finger and then sliding the two fingers together inward; or, in order to change the direction of viewing the picture content, pressing the relevant position with two fingers and then dragging the picture to translate or rotate the picture direction.
[0174] In this embodiment, a human-computer interaction interface is provided for the driver. The driver can use this interface to autonomously zoom in or out of the initial target area, thereby obtaining the final target area and the grid granularity corresponding to the target area. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the aforementioned multiple sets of display information.
[0175] Step 403: Perform perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple occupancy grid information.
[0176] In an embodiment of the present application, multimodal perception information and multiple sets of display information can be input into a pre-trained perception prediction network to output multiple occupancy grid information.
[0177] The aforementioned perception prediction network can be referenced above in the context of the neural network, including its principles and specific implementations. The perception prediction network can be pre-trained and downloaded to the vehicle computer. Subsequently, when the perception prediction network is used, its parameters can be updated based on real-time input and output to improve the accuracy of the network's predictions. This allows for occupancy grid information that better reflects actual driving conditions, assisting the driver in driving the vehicle more safely.
[0178] For example, FIG5 is a flow chart of the perception prediction process according to an embodiment of the present application. As shown in FIG5 , the perception prediction process (corresponding to the perception prediction network) includes the following steps:
[0179] (1) The multimodal perception information acquired by the vehicle computer includes M pinhole images (from a pinhole camera) and N fisheye images (from a fisheye camera). Feature extraction is performed on the M pinhole images and N fisheye images respectively, and the image features are transformed from the 2D perspective view (PV) to the 3D space according to the calibrated pose of the camera that collected each image.
[0180] That is, the perception prediction network includes a pinhole network and a fisheye network based on multimodal perception information; wherein, the pinhole network is used to extract the visual and spatial features of the pinhole image captured by the pinhole camera to obtain the pinhole 3D image features, and the fisheye network is used to extract the visual and spatial features of the fisheye image captured by the fisheye camera to obtain the fisheye 3D image features.
[0181] a. The resolutions of pinhole images and fisheye images can differ. Therefore, separate pinhole and fisheye images are used for feature extraction. For each of the M pinhole images, feature extraction can be performed separately based on the position of the pinhole camera (e.g., in front of, behind, or to the side of the vehicle). For each of the N fisheye images, feature extraction can also be performed separately based on the position of the fisheye camera (e.g., in front of, behind, or to the side of the vehicle).
[0182] The above-mentioned feature extraction of the image can adopt a method based on a convolutional neural network (CNN), such as a residual network (Resnet) or a feature map pyramid network (FPN), or a method based on a transformer (Transformer) structure, such as a vision transformer (Vision Transformer).
[0183] b. Based on feature extraction, the feature map pv_feat of the unified PV perspective is obtained, with dimensions C×H×W, where C is the number of feature channels, and H and W are the image sizes.
[0184] c. As in step a, use a neural network model to perform monocular depth estimation on the depth structure information of the image, thereby obtaining a depth distribution map of the image with dimensions D×H×W, where D is the number of depth distribution categories.
[0185] d. Use the depth profile to perform weighted mapping (e.g., broadcast multiplication) on the pv_feat to obtain a depth feature map with dimensions C × D × H × W. Based on camera intrinsic and extrinsic parameters (e.g., camera position, focal length, etc.), project the features in the depth feature map onto different locations in 3D space to obtain a 3D feature map with dimensions C1 × X × Y × Z, where X, Y, and Z represent the length, width, and height in the three-dimensional space.
[0186] Referring to steps ab above, the M pinhole images and the N fisheye images can be processed to obtain their respective 3D feature maps. It should be noted that the dimensions of the 3D feature maps obtained from different frames of images can be completely identical, not completely identical, or completely different, and this is not specifically limited.
[0187] (2) The multimodal perception information obtained by the vehicle computer includes laser point cloud (from laser radar) and millimeter wave point cloud (from millimeter wave radar), and feature extraction is performed on the laser point cloud and millimeter wave point cloud respectively.
[0188] That is, the perception prediction network includes a point cloud network based on multimodal perception information; the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter wave radar data to obtain point cloud voxel features.
[0189] a. For laser point clouds, voxel processing is performed using the correspondence between point clouds and grids (that is, the correspondence between 3D points in the point cloud and X, Y, and Z in three-dimensional space). Each voxel can include point cloud features such as the number of points, the maximum height of the 3D point, the minimum height of the 3D point, and the point cloud reflection intensity. After a convolution operation, a lidar feature map is obtained.
[0190] b. For the millimeter-wave point cloud, voxel processing is performed using the correspondence between the point cloud and the grid (that is, the correspondence between the 3D points in the point cloud and the X, Y, and Z coordinates in three-dimensional space). Then, a convolution operation is performed to obtain the millimeter-wave radar feature map.
[0191] c. Perform feature fusion (e.g., concatenation, addition, etc.) on the lidar feature map and the radar feature map to obtain a radar feature map with dimensions of C2×X×Y×Z.
[0192] (3) Perform multimodal feature fusion on 3D feature maps, lidar feature maps and radar feature maps in the same three-dimensional space.
[0193] That is, the perception prediction network includes a multimodal fusion network based on multimodal perception information; the multimodal fusion network is used to perform weighted fusion of pinhole 3D image features, fisheye 3D image features and point cloud voxel features.
[0194] a. Perform weighted fusion of the 3D feature maps corresponding to the M pinhole images and the N fisheye images to obtain the image 3D feature map.
[0195] b. Fuse the radar feature map with the image 3D feature map to obtain a fused 3D spatial feature map.
[0196] The above steps correspond to the time-series-based feature fusion function of the perception prediction network. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and there is no specific limitation on this.
[0197] (4) Further temporal alignment and spatial feature extraction are performed on the features of the fused 3D spatial feature map.
[0198] That is, the perception prediction network also includes a temporal fusion network, which is used to align the 3D spatial features of multiple frames in time sequence to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.
[0199] a. Combine the fused 3D spatial feature map of the current frame and the 3D spatial feature map of the historical frame T-1, and transform the 3D features of the historical frame according to the vehicle motion parameters (ego motion) so that they are in the same coordinate system as the 3D features at the current moment, obtaining the 3D spatial feature map of the T frame aligned to the same coordinate system.
[0200] b. The 3D spatial feature maps of T frames are fused into a 3D spatial feature map of the current frame according to the time sequence.
[0201] c. Use a multi-stage feature extractor to extract spatial features from the 3D spatial feature map of the current frame to obtain an enhanced 3D spatial feature map.
[0202] The above steps correspond to the function of multimodal feature fusion based on multimodal perception information of the perception prediction network. It should be noted that in addition to the above steps, this function can also be implemented in other ways, and there is no specific limitation on this.
[0203] (5) Perform multiple upsampling on the enhanced 3D spatial feature map and fuse the image features again for auxiliary feature enhancement;
[0204] That is, the perception prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.
[0205] a. The enhanced 3D spatial feature map is upsampled multiple times to obtain multiple resolutions corresponding to the above-mentioned multiple grid granularities;
[0206] b. Upsampling the image 3D feature map multiple times to obtain multiple resolutions corresponding to the multiple grid granularities;
[0207] c. At the same resolution, perform weighted fusion on the features of the upsampled enhanced 3D spatial feature map and the upsampled image 3D feature map to obtain a 3D spatial feature map at that resolution;
[0208] (6) Cropping sampling and weighted fusion are performed on the 3D spatial feature map at multiple scales of resolution according to the target area and grid granularity, and position features are obtained by interpolation to achieve arbitrary resolution.
[0209] That is, the perception prediction network also includes a dynamic resolution prediction network based on display information feedback, which is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction of arbitrary resolution.
[0210] a. Calculate the relative coordinates of the target area in the entire 3D feature space and crop it on the 3D spatial feature map at multiple scales based on the relative coordinates.
[0211] b. Based on the grid granularity (i.e., its corresponding resolution), generate the coordinates of all the positions that need to be predicted in the target area.
[0212] c. According to the coordinates of the occupied position predicted as needed, feature sampling and weighted fusion are performed on the cropped multi-scale resolution 3D spatial feature map. The feature sampling is obtained by trilinear interpolation of adjacent features on the feature map.
[0213] (7) Perform multilayer perceptron (MLP) prediction on all obtained position features, and then output the occupancy status and semantic information of the grid at the corresponding coordinate position. Among them, the occupancy status is a binary classification prediction, that is, there are only two categories: occupied / unoccupied, and the semantic information is a multi-class prediction, including road surface / moving target / other obstacles, etc.
[0214] For a single set of display information (including a target area and grid granularity), the grid occupancy and semantic information for all coordinate positions within the target area constitute the grid occupancy information corresponding to that display information. For example, for one set of display information (the target area can be the area within 6 meters around the ego vehicle, i.e., the target area has a length, width, and height of 12m × 12m × 2m, and the grid granularity can be 20cm), the grid occupancy and semantic information for all coordinate positions within 6 meters around the ego vehicle can be sensed and predicted, and the grid occupancy information can be displayed at a grid granularity of 20cm. Alternatively, for another set of display information (the target area can be the area within 3 meters around the ego vehicle, i.e., the target area has a length, width, and height of 6m × 6m × 2m, and the grid granularity can be 5cm), the grid occupancy and semantic information for all coordinate positions within 3 meters around the ego vehicle can be sensed and predicted, and the grid occupancy information can be displayed at a grid granularity of 5cm.
[0215] The above steps correspond to the perception prediction function of the perception prediction network based on display information. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and there is no specific limitation on this.
[0216] It should be noted that the embodiment shown in FIG5 only describes a method for implementing perception prediction. The embodiment of the present application may also adopt other methods for perception prediction, and no specific limitation is made to this.
[0217] Step 404: Display multiple occupied grid images on the human-computer interaction interface according to the multiple occupied grid information.
[0218] In one possible implementation, the vehicle computer can display a first occupied grid screen on the central control screen, which corresponds to the first target area and the first grid granularity; and display a second occupied grid screen, which corresponds to the second target area and the second grid granularity; wherein the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0219] Optionally, the first occupied grid image includes grids corresponding to n1 obstacles, and the second occupied grid image includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
[0220] The first occupied grid screen displays the screen of the first target area, which displays the grid of n1 obstacles perceived and predicted within the first target area at a first grid granularity; the second occupied grid screen displays the screen of the second target area, which displays the grid of n2 obstacles perceived and predicted within the second target area at a second grid granularity.
[0221] Optionally, the grid of the first obstacle in the first occupied grid screen and the grid of the second obstacle in the second occupied grid screen are identified by the same information, the first obstacle and the second obstacle refer to the same obstacle, n1 obstacles include the first obstacle, and n2 obstacles include the second obstacle.
[0222] The first target area and the second target area are both areas around the vehicle. The difference between the two is their range. The range of the first target area is larger than that of the second target area. Therefore, the first target area and the second target area may include the same obstacles. In other words, obstacles in the second target area may also exist in the first target area, and some obstacles in the first target area may not appear in the second target area.
[0223] To facilitate driver identification, in the embodiment of the present application, different obstacle grids can be displayed in different colors, different wireframes, etc., or different obstacle grids can be marked with different information (text, numbers, etc.). In this way, the driver can intuitively see the number, location, type, etc. of obstacles included in the target area, and then take timely avoidance measures.
[0224] Based on this, in the first occupied grid screen and the second occupied grid screen, the grids corresponding to the same obstacle can be identified with the same information. For example, in the first occupied grid screen and the second occupied grid screen, the grids of the first obstacle and the second obstacle are respectively displayed as the same color (for example, yellow grid) and the same wireframe (for example, solid wireframe), or, in the first occupied grid screen and the second occupied grid screen, the grids of the first obstacle and the second obstacle (obstacles existing in both the first target area and the second target area) are respectively marked with the same information (for example, the same text, the same number, etc.). The aforementioned first obstacle and the second obstacle refer to the same obstacle. In this way, the driver can intuitively see which are the same obstacles in the target areas of different ranges, and compare the distribution and distance of the obstacles under different fields of view, and then take avoidance measures in time.
[0225] Optionally, the grid of the third obstacle in the second occupied grid image is identified by specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0226] To facilitate driver identification, in an embodiment of the present application, the nearest obstacle grid around the vehicle can be marked with specific information. For example, the nearest obstacle grid can be highlighted with a special color (red), or the nearest obstacle grid can be displayed in a flashing manner, or the nearest obstacle grid can be marked with text near the nearest obstacle grid as the nearest obstacle, or the nearest obstacle grid can be marked with numbers near the nearest obstacle grid to indicate its distance from the vehicle. Furthermore, when the distance between the vehicle and the nearest obstacle is less than a set threshold (e.g., 0.5m), the vehicle's audio and lighting will synchronize warnings.
[0227] It should be noted that in the embodiment of the present application, other methods can also be used to display the grid of obstacles, and there is no specific limitation on this.
[0228] For example, FIG6 is a schematic diagram of the human-computer interaction interface of the central control screen of the embodiment of the present application. As shown in FIG6, the human-computer interaction interface is composed of a surround image display area and a panoramic occupied display area, wherein:
[0229] The surround view image display area mainly displays image information collected by the pinhole camera and fisheye camera around the vehicle.
[0230] The panoramic occupancy display area consists of two areas, the left area showing a wider target area and coarse-grained occupancy grid perception information, which can provide the driver with a more complete distribution of surrounding obstacles. The right area shows a smaller target area and fine-grained occupancy grid perception information, which can provide the driver with a detailed outline of close-range obstacles, thereby intuitively showing the distance between the vehicle and the close-range obstacles.
[0231] The left area corresponds to the first occupied grid screen, and the right area corresponds to the second occupied grid screen.
[0232] The above two grid screens display target areas in two ranges with different grid granularity. The wider target area can enhance the driver's panoramic perception of obstacles within a large range, and the smaller target area can enhance the driver's refined perception of obstacles within a close range. The combination of the two can enhance the driver's perception of the surrounding environment and avoid scratches.
[0233] It should be noted that the above embodiment is described using two occupied grid screens as an example. The embodiment of the present application does not specifically limit the number of occupied grid screens displayed in the central control screen, which can be equal to or greater than two. Each occupied grid screen corresponds to one display information, that is, one target area is displayed with a grid granularity.
[0234] In one possible implementation, the screen of the above human-computer interaction interface can be displayed on the driver's terminal device (for example, a mobile phone, tablet, etc.) so that the driver can simultaneously view the distribution of obstacles around the vehicle and the driving conditions of the vehicle.
[0235] In this embodiment of the present application, the driver can install a driving assistance application on a terminal device, which enables interconnection between the vehicle computer and the terminal device. When the vehicle computer executes step 404, it can transmit the image displayed on the vehicle computer to the terminal device in real time. When the driver opens the driving assistance application, he can see the image displayed on the vehicle computer, as shown in the embodiment of Figure 6.
[0236] In addition, the human-computer interaction interface provided by the assisted driving application also allows users to operate the screen, such as zooming in / out, changing the screen direction, etc.
[0237] In an embodiment of the present application, multiple occupied grid screens corresponding to multiple grid granularities are displayed on the vehicle's central control screen. The multiple occupied grid screens display target areas of different ranges with different grid granularities. This can not only enhance the driver's panoramic perception of obstacles within a large range, but also enhance the driver's refined perception of obstacles within a close range. The combination of the two can enhance the driver's perception of the surrounding environment and avoid scratches.
[0238] In one possible implementation, the human-computer interaction interface of the embodiment of the present application includes, in addition to the screen of the embodiment shown in Figure 6, a control for selecting a driving mode, which includes a driving mode and a parking mode. Therefore, two controls, "driving" and "parking", are set on the human-computer interaction interface.
[0239] Among them, in driving mode, the human-computer interaction interface also includes a control for selecting an auxiliary mode, which includes an automatic mode and a manual mode. At this time, two controls, "automatic" and "manual", are set on the human-computer interaction interface.
[0240] In parking mode, the human-computer interaction interface also includes controls for selecting auxiliary modes, which include automatic mode, manual mode, Automated Valet Parking (AVP) mode and Auto Parking Asist (APA) mode. At this time, four controls, "Automatic", "Manual", "AVP" and "APA", are set on the human-computer interaction interface. Automatic mode and manual mode can be selected under human driving conditions (that is, the driver drives the vehicle), and APA mode and AVP mode can be selected under intelligent driving conditions (that is, the driver turns on automatic driving, and the vehicle completes the parking instructions by itself. At this time, the driver can be in the car or outside the car). The APA mode can start the automatic parking function. During parking, the display logic of the panoramic display area can refer to the automatic mode under human driving conditions in the embodiment shown in Figure 9.
[0241] In one embodiment, in a scenario of low-speed driving on a narrow road, the implementation process of the assisted driving method is as follows:
[0242] 1. As shown in Figure 6, the driver activates the low-speed parking assist function through the vehicle button or the icon on the vehicle's central control screen, and then selects "Driving" mode in the human-computer interaction interface displayed on the central control screen.
[0243] 2. The vehicle processes the data collected by sensors around the vehicle (including pinhole cameras, fisheye cameras, lidar, and millimeter-wave radar).
[0244] 3. Based on the sensor data collected by the vehicle, perform real-time perception prediction of the environment around the vehicle. This embodiment of the present application implements this step through a perception prediction network. The input of the perception prediction network includes: 1) multimodal perception information, including pinhole images, fisheye images, laser point clouds, millimeter wave point clouds, etc.; 2) multiple display information (including target areas and grid granularity); the output of the perception prediction network includes multiple occupancy grid information (including occupancy status and semantic information of grids at all coordinate positions within the target area).
[0245] The implementation process of the perception prediction network can refer to the embodiment shown in Figure 5, which will not be repeated here.
[0246] 4. The vehicle displays information of multiple occupied grids in real time on the central control screen.
[0247] As shown in Figure 6, the human-computer interaction interface consists of a surround view image display area and a panoramic occupancy display area. The surround view image display area primarily displays image information collected by the pinhole and fisheye cameras surrounding the ego vehicle. The panoramic occupancy display area consists of two areas: the left area displays a wider target area and coarse-grained occupancy grid perception information, providing the driver with a more comprehensive distribution of surrounding obstacles. The right area displays a smaller target area and fine-grained occupancy grid perception information, providing the driver with a detailed outline of nearby obstacles, thereby intuitively displaying the distance between the ego vehicle and the nearby obstacles.
[0248] Optionally, the low-speed parking assist function in this embodiment of the present application allows the driver to adjust the viewing angles of different grid displays. As shown in Figure 7 (Figure 7 is a schematic diagram of the human-computer interaction interface switching display angles in this embodiment of the present application), the driver can drag and rotate the central control screen to adjust to a bird's-eye view (BEV). This allows the driver to experience a multi-perspective driving experience of viewing the distribution of surrounding obstacles.
[0249] 5. The low-speed parking assist function of this embodiment of the present application can provide two modes for fine-tuning the display of specific areas: "Automatic" and "Manual" modes. As shown in Figure 8 (Figure 8 is a schematic diagram of the fine-tuning mode of this embodiment of the present application), the two modes can be switched in the upper right corner of the human-computer interaction interface. The driver can switch to "Automatic" mode by clicking a gesture button.
[0250] Optionally, the low-speed parking assist function may be set to an "automatic" mode by default. The "automatic" mode may correspond to the situation in step 402 above where "the displayed information is pre-set information."
[0251] In the "driving" and "automatic" modes of this embodiment, in the human-computer interaction interface of the central control screen, the occupied grid information of the area within 6 meters around the vehicle (i.e., the target area) is displayed in the left area of the panoramic occupied display area with a coarse grid granularity of 20 cm; the occupied grid information of the area within 3 meters around the vehicle (i.e., the target area) is displayed in the right area of the panoramic occupied display area with a fine grid granularity of 5 cm.
[0252] The driver can use the obstacle occupancy grid information displayed on the central control screen to know which direction of the vehicle there is an obstacle, the distance between the vehicle and each obstacle, and which position of the vehicle is close to the obstacle, etc., so as to take preventive measures in advance and avoid scratches.
[0253] 6. As shown in Figure 8, the central control screen can also display the distance between the nearest obstacle and the vehicle. This allows the driver to more intuitively see how far the nearest obstacle is and take appropriate driving measures to avoid the obstacle and prevent scratches.
[0254] In this embodiment, the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities. As the environment changes in real time during parking, the vehicle can adaptively adjust the level of sophistication of the human-computer interaction interface, allowing the driver to take preventive measures in advance to avoid scratches.
[0255] In another embodiment, in a scenario where the vehicle is parking in a narrow parking space, the assisted driving method is implemented as follows:
[0256] 1. As shown in Figure 9 (Figure 9 is a schematic diagram of the human-computer interaction interface of the central control screen in an embodiment of the present application), the driver activates the low-speed parking assist function through the vehicle button or the icon on the vehicle's central control screen, and then selects "Parking" mode in the human-computer interaction interface displayed on the central control screen.
[0257] Steps 2-4 of this embodiment may refer to steps 2-4 of the embodiment shown in FIG6 , and are not described again here.
[0258] 5. When the driver is driving, the low-speed parking assist function of the embodiment of the present application can provide two ways to refine the display of specific areas: "automatic" and "manual" modes. As shown in Figure 10 (Figure 10 is a schematic diagram of the refined modes of the embodiment of the present application), the two modes can be switched in the upper right corner of the human-computer interaction interface. The driver can click a gesture button to switch to "manual" mode. "Manual" mode corresponds to the situation in step 402 above where "the displayed information is dynamically generated based on the driver's operations on the touch screen."
[0259] In the "Parking" and "Manual" modes of this embodiment, as shown in FIG10 , the central control screen's human-machine interface initially displays occupation grid information for an area within 6 meters around the vehicle (i.e., the target area) in a coarse 20 cm grid granularity in the left and right areas of the panoramic occupation display area. The right area is available for the driver to perform zoom operations.
[0260] As shown in Figures 11 and 12 (Figures 11 and 12 are schematic diagrams of the zoom operation in the manual mode of an embodiment of the present application), the driver can manually move and zoom the right area of the panoramic occupied display area on the human-computer interaction interface. The vehicle computer extracts the zoomed area and multiple based on the driver's operation, and uses the zoomed area as the target area. The multiple can be used as a basis for obtaining the grid granularity. The driver locates the target area to the area within 4m around the vehicle (i.e., the target area) through the zoom operation. At this time, in the human-computer interaction interface of the central control screen, the occupied grid information of the area within 4m around the vehicle (i.e., the target area) is displayed in the right area of the panoramic occupied display area with a coarse grid granularity of 8cm.
[0261] In this embodiment, the driver can manually determine the target area, and the vehicle can display the grid occupancy and semantic information of obstacles around the vehicle on the central control screen at various grid granularities, allowing the driver to display the area of interest in a refined manner. In addition, as the environment changes in real time during parking, the refinement of the human-computer interaction interface can be adaptively adjusted, allowing the driver to take preventive measures in advance to avoid scratches.
[0262] In another embodiment, the vehicle computer may display a panoramic occupancy grid image on the human-computer interaction interface according to the panoramic occupancy grid information, where the panoramic occupancy grid information is obtained based on the historical occupancy grid information and driving trajectory of the vehicle's driving environment.
[0263] In the AVP garage / parking lot parking scenario, the implementation process of the assisted driving method is as follows:
[0264] 1. As shown in Figure 13 (Figure 13 is a schematic diagram of the human-computer interaction interface of the central control screen in an embodiment of the present application), after the vehicle enters the garage / parking lot, the driver activates the low-speed parking assist function using the vehicle button or the icon on the vehicle's central control screen, then selects "Parking" mode in the human-computer interaction interface displayed on the central control screen, and selects AVP mode.
[0265] Steps 2-3 of this embodiment may refer to steps 2-3 of the embodiment shown in FIG6 , and are not described again here.
[0266] 4. As the vehicle drives through the garage / parking lot, the predicted occupancy grid and semantic information are fused across multiple frames to reconstruct the static structure of the entire scene. This involves the following steps:
[0267] (1) Remove dynamic objects based on the predicted semantic information and only retain static structural obstacles;
[0268] (2) Convert the remaining occupied grids to a unified world coordinate system based on the vehicle’s position;
[0269] (3) Accumulate the occupancy grid information in the world coordinate system of the historical N frames, perform probability weighted fusion, and obtain the fused 3D occupancy grid information (corresponding to the panoramic occupancy grid information);
[0270] (4) Convert the 3D occupancy grid information to the current vehicle coordinate system;
[0271] (5) The historical trajectory points of the vehicle are also converted to the current vehicle coordinate system;
[0272] 5. Real-time display and output based on the 3D occupancy grid information in the vehicle coordinate system. Referring to the embodiments shown in Figures 7 and 8, the driver can perform drag and zoom operations on the human-computer interaction interface on the central control screen.
[0273] 6. In AVP mode, the driver can view the basement reconstruction results, vehicle driving and parking trajectories through the assisted driving application on the terminal device (such as mobile phone, tablet, etc.), and the picture can be the same as the picture of the car computer.
[0274] This embodiment can perform panoramic perception of the entire garage / parking lot based on the changes in the vehicle's posture, thereby providing the driver with effective full-scene static environment information for parking or automatic parking.
[0275] FIG14 is a schematic diagram of the structure of the assisted driving device 1400 of the present application. As shown in FIG14 , the assisted driving device 1400 of this embodiment can be applied to the terminal device mentioned above. The assisted driving device 1400 may include: a collection module 1401, an acquisition module 1402, a prediction module 1403 and a display module 1404.
[0276] An acquisition module 1401 is used to acquire multimodal perception information, where the multimodal perception information is acquired by a variety of sensors installed on the vehicle. An acquisition module 1402 is used to acquire multiple sets of display information, where the display information includes a target area to be displayed and a grid granularity. A prediction module 1403 is used to perform perception prediction based on the multimodal perception information and the multiple sets of display information to obtain multiple pieces of occupied grid information, where the occupied grid information is used to describe the distribution of obstacles around the vehicle. A display module 1404 is used to display multiple occupied grid images on a human-computer interaction interface based on the multiple pieces of occupied grid information.
[0277] In one possible implementation, the display module 1404 is specifically used to display a first occupied grid screen, where the first occupied grid screen corresponds to a first target area and a first grid granularity; and to display a second occupied grid screen, where the second occupied grid screen corresponds to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
[0278] In a possible implementation, the first occupied grid picture includes grids corresponding to n1 obstacles, and the second occupied grid picture includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
[0279] In one possible implementation, the grid of the first obstacle in the first occupied grid screen and the grid of the second obstacle in the second occupied grid screen are identified by the same information, the first obstacle and the second obstacle are the same obstacle, the n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
[0280] In a possible implementation, the grid of the third obstacle in the second occupied grid image is identified by specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.
[0281] In one possible implementation, the display module 1404 is also used to display a panoramic occupancy grid screen on the human-computer interaction interface based on the panoramic occupancy grid information, and the panoramic occupancy grid information is obtained based on the historical occupancy grid information and driving trajectory of the vehicle's driving environment.
[0282] In a possible implementation manner, the display information is preset information.
[0283] In a possible implementation, the display information is dynamically generated based on the driver's operation on the touch screen.
[0284] In one possible implementation, the operation includes at least one of the following operations: the driver presses the touch screen with two fingers and slides the two fingers apart outward or slides the two fingers together inward; or the driver presses the touch screen with two fingers and rotates, translates or drags.
[0285] In a possible implementation, the human-computer interaction interface further includes a control for selecting a driving mode, where the driving mode includes one or more of the following modes: a driving mode or a parking mode.
[0286] In a possible implementation, in the driving mode, the human-computer interaction interface further includes a control for selecting an auxiliary mode, where the auxiliary mode includes one or more of the following modes: an automatic mode or a manual mode.
[0287] In one possible implementation, in the parking mode, the human-computer interaction interface also includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode, or automatic parking assist APA mode.
[0288] In a possible implementation, the prediction module 1403 is specifically configured to input the multimodal perception information and the multiple sets of display information into a pre-trained perception prediction network to output the multiple occupancy grid information.
[0289] In a possible implementation, the functions implemented by the perception prediction network include: multimodal feature fusion based on multimodal perception information; feature fusion based on time series; and perception prediction based on the display information.
[0290] In one possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multimodal fusion network based on the multimodal perception information; wherein the pinhole network is used to extract visual and spatial features of the pinhole image captured by the pinhole camera to obtain pinhole 3D image features, the fisheye network is used to extract visual and spatial features of the fisheye image captured by the fisheye camera to obtain fisheye 3D image features, the point cloud network is used to extract 3D voxel features corresponding to lidar and / or millimeter-wave radar data to obtain point cloud voxel features, and the multimodal fusion network is used to perform weighted fusion of the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
[0291] In one possible implementation, the perception prediction network also includes a temporal fusion network, which is used to align the 3D spatial features of multiple frames in time sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.
[0292] In one possible implementation, the perception prediction network also includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsampling on the 3D spatial feature map, and to fuse image features again to achieve auxiliary feature enhancement.
[0293] In one possible implementation, the perception prediction network also includes a dynamic resolution prediction network based on the display information feedback, and the dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area and grid granularity selected on the human-computer interaction interface, and obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction of arbitrary resolution.
[0294] In one possible implementation, the multiple sensors include a pinhole camera and a fisheye camera.
[0295] In one possible implementation, the multiple sensors further include a laser radar and / or a millimeter-wave radar.
[0296] The device of this embodiment can be used to execute the technical solution of the method embodiment shown in Figure 4. Its implementation principle and technical effects are similar and will not be repeated here.
[0297] During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware coding processor, or can be executed by a combination of hardware and software modules in the coding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0298] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0299] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0300] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0301] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0302] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0303] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0304] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0305] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An assisted driving method, characterized in that, Including: Obtaining multi-modal perception information, which is collected by a variety of sensors arranged on the vehicle; Obtaining multiple sets of display information, where the display information includes a target area to be displayed and a grid granularity; Performing perception prediction based on the multi-modal perception information and the multiple sets of display information to obtain multiple occupancy grid information, which is used to describe the distribution of obstacles around the vehicle; Displaying multiple occupancy grid pictures on the human-machine interaction interface according to the multiple occupancy grid information.
2. The method according to claim 1, wherein The displaying multiple occupancy grid pictures according to the multiple occupancy grid information includes: Displaying a first occupancy grid picture, which corresponds to a first target area and a first grid granularity; Displaying a second occupancy grid picture, which corresponds to a second target area and a second grid granularity; The range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
3. The method according to claim 2, characterized in that The first occupancy grid picture includes grids corresponding to n1 obstacles, and the second occupancy grid picture includes grids corresponding to n2 obstacles, where both n1 and n2 are positive integers.
4. The method according to claim 3, wherein The grids of the first obstacle in the first occupancy grid picture and the grids of the second obstacle in the second occupancy grid picture are marked with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
5. The method according to any one of claims 2-4, characterized in that The grids of the third obstacle in the second occupancy grid picture are marked with specific information. The third obstacle is the one closest to the vehicle among the n2 obstacles.
6. The method according to any one of claims 1-5, characterized in that It also includes: Displaying a panoramic occupancy grid picture on the human-machine interaction interface according to the panoramic occupancy grid information, where the panoramic occupancy grid information is obtained based on the historical occupancy grid information of the driving environment of the vehicle and the driving trajectory.
7. The method according to any one of claims 1-6, characterized in that, The display information is pre-set information.
8. The method according to any one of claims 1-6, characterized in that, The display information is dynamically generated based on the operations of the driver on the touch screen.
9. The method according to claim 8, wherein The operations include at least one of the following operations: The operation that the driver presses the touch screen with two fingers and slides them outward or inward; or, The operation that the driver presses the touch screen with two fingers and performs rotation, translation or dragging.
10. The method according to any one of claims 1-9, characterized in that The human-machine interaction interface also includes a control for selecting a driving mode, and the driving mode includes one or more of the following modes: driving mode or parking mode.
11. The method according to claim 10, characterized in that, In the driving mode, the human-machine interaction interface also includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode or manual mode.
12. The method according to claim 10, wherein In the parking mode, the human-machine interaction interface also includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode or automatic parking assist APA mode.
13. The method according to any one of claims 1 to 12, characterized in that, The performing perception prediction based on the multi-modal perception information and the multiple sets of display information to obtain multiple occupancy grid information includes: Input the multi-modal perception information and the multiple groups of display information into a pre-trained perception prediction network to output the multiple occupancy grid information.
14. The method according to claim 13, wherein The perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multi-modal fusion network based on the multi-modal perception information; where the pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain pinhole 3D image features, the fisheye network is used to extract the visual and spatial features of the fisheye image collected by the fisheye camera to obtain fisheye 3D image features, the point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter wave radar data to obtain point cloud voxel features, and the multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.
15. The method according to claim 13 or 14, characterized in that, The perception prediction network further includes a temporal fusion network, which is used to align multi-frame 3D spatial features in chronological order to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.
16. The method according to any one of claims 13 - 15, characterized in that, The perception prediction network further includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.
17. The method according to any one of claims 13-16, characterized in that, The perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area selected on the human-computer interaction interface and the grid granularity, and obtain position features through interpolation, and then predict occupancy and semantic information according to the position features to achieve prediction at any resolution.
18. The method according to any one of claims 1 to 17, characterized in that The multiple sensors include a pinhole camera and a fisheye camera.
19. The method according to claim 18, wherein The multiple sensors further include a lidar and / or a millimeter wave radar.
20. An assisted driving device, characterized in that, Comprising: An acquisition module, configured to acquire multi-modal perception information, where the multi-modal perception information is acquired by multiple sensors disposed on the vehicle; An acquisition module, configured to acquire multiple groups of display information, where the display information includes the target area to be displayed and the grid granularity; A prediction module, configured to perform perception prediction according to the multi-modal perception information and the multiple groups of display information to obtain multiple occupancy grid information, where the occupancy grid information is used to describe the distribution of obstacles around the vehicle; A display module, configured to display multiple occupancy grid pictures on the human-computer interaction interface according to the multiple occupancy grid information.
21. The device according to claim 20, wherein The display module is specifically configured to display a first occupancy grid picture corresponding to a first target area and a first grid granularity; display a second occupancy grid picture corresponding to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.
22. The device according to claim 21, wherein, The first occupancy grid picture includes grids corresponding to n1 obstacles, and the second occupancy grid picture includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.
23. The device according to claim 22, characterized in that, The grids of the first obstacle in the first occupied grid image and the grids of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.
24. The device according to any one of claims 21 to 23, characterized in that, The grids of the third obstacle in the second occupied grid image are identified with specific information. The third obstacle is the one closest to the vehicle among the n2 obstacles.
25. The device according to any one of claims 20-24, characterized in that The display module is further configured to display an omnidirectional occupied grid image on the human-machine interaction interface according to the omnidirectional occupied grid information, where the omnidirectional occupied grid information is obtained based on the historical occupied grid information and the driving trajectory of the vehicle's driving environment.
26. The device according to any one of claims 20-25, characterized in that, The display information is pre-set information.
27. The device according to any one of claims 20-25, characterized in that, The display information is dynamically generated based on the operations of the driver on the touch screen.
28. The device according to claim 27, wherein The operations include at least one of the following operations: The operation that the driver presses the touch screen with two fingers and slides the two fingers outward or inward; or, The operation that the driver presses the touch screen with two fingers and performs rotation, translation, or dragging.
29. The device according to any one of claims 20-28, characterized in that, The human-machine interaction interface further includes a control for selecting a driving mode, and the driving mode includes one or more of the following modes: driving mode or parking mode.
30. The device according to claim 29, wherein, In the driving mode, the human-machine interaction interface further includes a control for selecting an assist mode, and the assist mode includes one or more of the following modes: automatic mode or manual mode.
31. The device according to claim 29, characterized in that, In the parking mode, the human-machine interaction interface further includes a control for selecting an assist mode, and the assist mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode, or automatic parking assist APA mode.
32. The device according to any one of claims 20 - 31, characterized in that, The prediction module is specifically configured to input the multi-modal perception information and the multiple groups of display information into a pre-trained perception prediction network to output the multiple occupied grid information.
33. The device according to claim 32, characterized in that, The perception prediction network includes one or more of the following networks: a pinhole network, a fish-eye network, a point cloud network, or a multi-modal fusion network based on the multi-modal perception information; where, The pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain the pinhole 3D image features. The fish-eye network is used to extract the visual and spatial features of the fish-eye image collected by the fish-eye camera to obtain the fish-eye 3D image features. The point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter-wave radar data to obtain the point cloud voxel features. The multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fish-eye 3D image features, and the point cloud voxel features.
34. The device according to claim 32 or 33, characterized in that, The perception prediction network further includes a temporal fusion network, and the temporal fusion network is used to align the multi-frame 3D spatial features in sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.
35. The device according to any one of claims 32-34, characterized in that, The perception prediction network further includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.
36. The device according to any one of claims 32 - 35, characterized in that, The perception prediction network further includes a dynamic resolution prediction network based on the display information feedback. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area selected on the human-computer interaction interface and the grid granularity, obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction at any resolution.
37. The device according to any one of claims 20-36, characterized in that, The multiple sensors include a pinhole camera and a fish-eye camera.
38. The device according to claim 37, characterized in that, The multiple sensors further include a lidar and / or a millimeter-wave radar.
39. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-19.
40. A computer-readable storage medium, characterized in that, Including a computer program, when the computer program is executed on a computer, the computer executes the method according to any one of claims 1-19.
41. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on a computer, the computer executes the method according to any one of claims 1-19.
Citation Information
Patent Citations
Vehicle surroundings monitoring device
CN102652327A
Display control device
CN110895443A
Display system in a vehicle
CN111163968A
Method for reproducing the surroundings of a vehicle
CN111201558A
Display method, electronic device and computer readable storage medium
CN113829996A
Cited By
Method, device and system for sensing driver state and vehicle
CN121133717A
Environment perception method and system, computer equipment and storage medium
CN121144791A
Visual monitoring method and system for hidden engineering of airport
CN121305476A
Trajectory prediction planning method, trajectory prediction planning model training method, trajectory prediction planning model training device, medium and equipment
CN121438254A
Trajectory prediction planning and model training method and device, medium and equipment
CN121438254B