Auxiliary driving method and device

Through multimodal perception information processing and grid display technology, the problem of insufficient close obstacle recognition in specific scenarios of the vehicle-mounted driving system is solved, panoramic and refined environmental perception is achieved, and driving safety is improved.

CN120270238APending Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410029970.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing on-board assisted driving functions are difficult to provide drivers with good safe driving assistance in scenarios such as narrow parking spaces, narrow road traffic, narrow road turnover, narrow road traffic, automatic emergency braking, etc., especially in the identification and display of close obstacles.

Method used

Through the acquisition and processing of multimodal perceptual information, combined with multiple sets of display information, the perceptual prediction network is used to generate multiple occupancy grid information, and the distribution of obstacles in different ranges is displayed on the central control screen at different grid granularities, providing panoramic and refined environmental perception capabilities.

Benefits of technology

It improves the driver's perception of the surrounding environment, can accurately identify obstacles within different ranges, avoid scratches, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120270238A_ABST
    Figure CN120270238A_ABST
Patent Text Reader

Abstract

The invention provides an auxiliary driving method and device. The auxiliary driving method comprises the steps that multi-mode sensing information is acquired, and the multi-mode sensing information is acquired by multiple sensors arranged on a vehicle; obtaining multiple groups of display information, wherein the display information comprises a to-be-displayed target area and grid granularity; according to the multi-mode sensing information and the multiple groups of display information, sensing prediction is carried out to obtain multiple pieces of occupied grid information, and the occupied grid information is used for describing the distribution condition of obstacles around the vehicle; and displaying a plurality of occupied grid pictures on a human-computer interaction interface according to the plurality of pieces of occupied grid information. According to the invention, the panoramic perception ability of the driver to the obstacles in a large range can be improved, the refined perception ability of the driver to the obstacles in a close range can be improved, the perception ability of the driver to the surrounding environment can be improved by combining the panoramic perception ability and the refined perception ability, and scratching is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to vehicle assisted driving technology, and in particular to an assisted driving method and device. Background Art

[0002] With the development of engineering technology and the improvement of intelligent perception ability, vehicle assisted driving functions have become standard equipment for various vehicles. For example, reverse assist lines, parking images, etc. However, current vehicle assisted driving functions are difficult to provide good and safe driving assistance capabilities for drivers in some specific scenarios, such as narrow parking spaces, narrow lane driving, narrow road turning around, narrow road meeting, Automatic Emergency Braking (AEB), Rear Automatic Emergency Braking (RAEB), etc. Summary of the Invention

[0003] This application provides an assisted driving method and device to improve the driver's perception ability of the surrounding environment and avoid scratches.

[0004] In a first aspect, this application provides an assisted driving method, including: obtaining multi-modal perception information, where the multi-modal perception information is collected by a variety of sensors disposed on the vehicle; obtaining multiple sets of display information, where the display information includes a target area to be displayed and a grid granularity; performing perception prediction based on the multi-modal perception information and the multiple sets of display information to obtain multiple occupancy grid information, where the occupancy grid information is used to describe the distribution of obstacles around the vehicle; and displaying multiple occupancy grid pictures on a human-machine interaction interface according to the multiple occupancy grid information.

[0005] In an embodiment of this application, by displaying multiple occupancy grid pictures corresponding to multiple grid granularities on the center control screen of the vehicle, the multiple occupancy grid pictures display different ranges of target areas with different grid granularities, which can not only improve the driver's panoramic perception ability of obstacles in a large range, but also improve the driver's refined perception ability of obstacles in a short distance range. The combination of the two can improve the driver's perception ability of the surrounding environment and avoid scratches.

[0006] In an embodiment of this application, the driver can trigger the vehicle to start the low-speed driving and parking assistance function in the following multiple ways:

[0007] 1. The driver inputs a start command through a vehicle button or the vehicle center control screen.

[0008] In a possible implementation, an assisted driving mode button is provided in the vehicle, and the driver can press the button to trigger the start of the low-speed parking and driving assistance function. It should be understood that there are various specific implementation manners for starting the target function (such as the low-speed parking and driving assistance function) through a button. For example, pressing the button corresponding to the target function, or pressing the button multiple times briefly to poll to the target function, etc. The embodiments of the present application do not make specific limitations on this.

[0009] In a possible implementation, the main interface is displayed on the central control screen, and the main interface includes an icon corresponding to the low-speed parking and driving assistance function. The driver clicks the icon to trigger the start of the low-speed parking and driving assistance function. It should be understood that there are various specific implementation manners for starting the target function (such as the low-speed parking and driving assistance function) through the central control screen (with the ability of a touch screen). For example, directly clicking the icon of the target function, or selecting the option of the target function from the menu of the vehicle assisted driving mode, etc. The embodiments of the present application do not make specific limitations on this.

[0010] 2. The driver starts the instruction through voice input.

[0011] In a possible implementation, the driver says "Please start the low-speed parking and driving assistance function" to the speaker of the vehicle head unit. When the vehicle head unit receives and recognizes the semantics of the voice, the low-speed parking and driving assistance function is started.

[0012] It should be noted that in addition to the above two methods, the embodiments of the present application can also use other methods to trigger the vehicle to start the low-speed parking and driving assistance function, and no specific limitations are made on this.

[0013] After the vehicle starts the low-speed parking and driving assistance function, the vehicle head unit can first obtain multi-modal perception information. A variety of sensors are provided in the vehicle, including a camera sensor (abbreviated as camera), a lidar sensor (abbreviated as lidar), and a millimeter wave radar sensor (abbreviated as millimeter wave radar). In the embodiments of the present application, the multi-modal perception information can be obtained based on the information collected by one or more of the foregoing various sensors.

[0014] A set of display information may include a target area to be displayed and a grid granularity. Among them, the target area may refer to the range of the area of the real world displayed on the center control screen. For example, in the real world, the area within 3m around the vehicle (i.e., the host vehicle), centered on the vehicle (i.e., the host vehicle), a 6m×6m square plane area 3m forward, 3m backward, 3m left, and 3m right, and then superimposed with a height of 2m to form a 6m×6m×2m three-dimensional space, and this three-dimensional space is the target area. The length, width, and height of the target area can be expressed as W (length)×H (width)×D (height). When representing the target area, the host vehicle can be used as the origin, and the three-dimensional coordinates of the 8 vertices of the target area can be used to represent the range of the target area. Alternatively, the range of the target area can also be represented by the three-dimensional coordinates of a certain vertex of the target area and the length, width, and height of the target area. The embodiments of the present application do not make specific limitations on the representation form of the range of the target area.

[0015] The grid granularity refers to the principle of occupancy grid. Occupancy grid means dividing the three-dimensional space in the form of voxels, and each voxel is characterized by binary values 0 and 1 to indicate whether the voxel is occupied. It can be seen that a voxel can correspond to a grid, and the grid granularity can be used to represent the granularity of dividing the three-dimensional space into voxels. For example, the size of a voxel is 1m×1m×1m, and the above 6m×6m×2m three-dimensional space can be divided into 6×6×2 voxels, that is, this three-dimensional space corresponds to 6×6×2 grids.

[0016] Optionally, the larger the range of the target area, the larger the grid granularity that can be selected, and the smaller the range of the target area, the smaller the grid granularity that can be selected. In this way, more occupancy grid information can be displayed in a larger target area, and more detailed occupancy grid information can be displayed in a smaller target area.

[0017] In a possible implementation, in a set of display information, the target area may refer to the area within 6m around the host vehicle, that is, the length, width, and height of the target area are 12m×12m×2m, and the grid granularity may be 20cm.

[0018] In a possible implementation, in a set of display information, the target area may refer to the area within 3m around the host vehicle, that is, the length, width, and height of the target area are 6m×6m×2m, and the grid granularity may be 5cm.

[0019] In a possible implementation, in a set of display information, the target area may refer to the area within 1m around the host vehicle, that is, the length, width, and height of the target area are 2m×2m×2m, and the grid granularity may be 2cm.

[0020] In the embodiments of the present application, the correspondence between the target area and the grid granularity can be a pre-established correspondence (which can be used as the configuration for default recommendation), or the correspondence between the target area and the grid granularity can be determined in real time, and no specific limitation is made thereto.

[0021] In the embodiments of the present application, the vehicle-mounted computer can obtain display information through the following multiple methods:

[0022] 1. The display information is pre-set information.

[0023] Multiple groups of display information can be pre-set and stored. When the display information is needed, the vehicle-mounted computer can directly read the memory to obtain the multiple groups of display information. For example, two groups of display information are set. In one group of display information, the target area refers to the area within 6 m around the vehicle itself, that is, the length, width, and height of the target area are 12 m × 12 m × 2 m, and the grid granularity is 20 cm. In the other group of display information, the target area refers to the area within 3 m around the vehicle itself, that is, the length, width, and height of the target area are 6 m × 6 m × 2 m, and the grid granularity is 5 cm.

[0024] 2. The display information is dynamically generated based on the operations of the driver on the touch screen.

[0025] The above zoom operations may include the operations of the driver pressing the touch screen with two fingers and sliding the two fingers apart or sliding the two fingers together inward, for changing the range of the target area and the target resolution of the obstacle; or, the above zoom operations may include the operations of the driver pressing the touch screen with two fingers and performing rotation, translation, or dragging, for changing the display orientation and range of the target area. The foregoing operations can refer to when the driver uses a picture application. To enlarge the details of a picture, the relevant position can be pressed with the thumb and index finger, and then the two fingers are slid apart; or when reducing the picture to view the whole picture, the relevant position can be pressed with the thumb and index finger, and then the two fingers are slid together inward; or, to view the picture content in a different direction, the relevant position can be pressed with two fingers, and then the picture is dragged for translation or the picture direction is rotated.

[0026] In the embodiments of the present application, a human-machine interaction interface is provided for the driver. On the initial target area, the driver can autonomously zoom in / out the target area through this human-machine interaction interface, so as to obtain the final target area and the grid granularity corresponding to the target area. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the above multiple groups of display information.

[0027] In the embodiments of the present application, the multi-modal perception information and multiple groups of display information can be input into a pre-trained perception prediction network to output multiple occupancy grid information.

[0028] The above-mentioned perception prediction network can refer to the relevant content of the neural network in this application, and its principle and specific implementation can be referred to its content. The perception prediction network can be pre-trained and downloaded to the in-vehicle computer. Then, when using the perception prediction network, the parameters in the perception prediction network can be updated based on real-time input and output to make the prediction result of the perception prediction network more accurate, so as to obtain occupancy grid information that better conforms to the actual driving situation and assist the driver to drive the vehicle more safely.

[0029] For a single set of display information (including the target area and grid granularity), the occupancy situation and semantic information of the grids at all coordinate positions within the target area are the occupancy grid information corresponding to this display information. For example, for a set of display information (the target area can refer to the area within 6m around the vehicle, that is, the length, width, and height of the target area are 12m×12m×2m, and the grid granularity can be 20cm), it is possible to perceptually predict the occupancy situation and semantic information of the grids at all coordinate positions within the area within 6m around the vehicle, and display the grid occupancy situation with a grid granularity of 20cm; or, for another set of display information (the target area can refer to the area within 3m around the vehicle, that is, the length, width, and height of the target area are 6m×6m×2m, and the grid granularity can be 5cm), it is possible to perceptually predict the occupancy situation and semantic information of the grids at all coordinate positions within the area within 3m around the vehicle, and display the grid occupancy situation with a grid granularity of 5cm.

[0030] In a possible implementation, the in-vehicle computer can display a first occupancy grid screen on the center control screen, and this first occupancy grid screen corresponds to a first target area and a first grid granularity; display a second occupancy grid screen, and this second occupancy grid screen corresponds to a second target area and a second grid granularity; where the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.

[0031] Optionally, the above-mentioned first occupancy grid screen includes grids corresponding to n1 obstacles, and the above-mentioned second occupancy grid screen includes grids corresponding to n2 obstacles, and both n1 and n2 are positive integers.

[0032] The first occupancy grid screen displays the picture of the first target area, and it displays the grids of n1 obstacles perceptually predicted within the first target area with the first grid granularity; the second occupancy grid screen displays the picture of the second target area, and it displays the grids of n2 obstacles perceptually predicted within the second target area with the second grid granularity.

[0033] Optionally, the grids of the first obstacle in the first occupied grid image and the grids of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle refer to the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles include the second obstacle.

[0034] Both the first target area and the second target area are areas around the vehicle. The difference between them lies in the range. Specifically, the range of the first target area is larger than that of the second target area. Therefore, the first target area and the second target area may include the same obstacles. In other words, the obstacles in the second target area may also exist in the first target area, and some of the obstacles in the first target area may not appear in the second target area.

[0035] For the convenience of the driver's identification, in the embodiments of the present application, the grids of different obstacles can be displayed in different colors, different wireframes, etc., or the grids of different obstacles can be marked with different information (such as text, numbers, etc.). In this way, the driver can intuitively see the number, position, type, etc. of the obstacles included in the target area, and then take avoidance measures in a timely manner.

[0036] Based on this, in the first occupied grid image and the second occupied grid image, the grids corresponding to the same obstacle can be identified with the same information. For example, in the first occupied grid image and the second occupied grid image, the grids of the first obstacle and the second obstacle are respectively displayed in the same color (such as yellow grid), the same wireframe (such as solid wireframe), or in the first occupied grid image and the second occupied grid image, the grids of the first obstacle and the second obstacle (the obstacles that exist in both the first target area and the second target area) are marked with the same information (such as the same text, the same number, etc.). The aforementioned first obstacle and the second obstacle refer to the same obstacle. In this way, the driver can intuitively see which are the same obstacles in the target areas with different ranges, and compare the distribution and distance of the obstacles in different fields of view, and then take avoidance measures in a timely manner.

[0037] Optionally, the grids of the third obstacle in the second occupied grid image are identified with specific information. The third obstacle is the one closest to the vehicle among the n2 obstacles.

[0038] For the convenience of driver identification, in the embodiments of the present application, the grid of the obstacle closest to the vehicle can be marked with specific information. For example, the grid of the obstacle closest to the vehicle can be highlighted in a special color (red), or the grid of the obstacle closest to the vehicle can be displayed in a flashing manner, or the grid of the obstacle closest to the vehicle can be marked with text nearby indicating that it is the closest obstacle, or the distance between the grid of the obstacle closest to the vehicle and the vehicle can be marked with a number nearby. Further, when the distance between the vehicle and the closest obstacle is less than a set threshold (such as 0.5 m), the vehicle's audio and lights will give a synchronous warning.

[0039] It should be noted that in the embodiments of the present application, other ways can also be used to display the grid of the obstacle, and no specific limitation is made thereto.

[0040] In a possible implementation manner, the screen of the above-mentioned human-machine interaction interface can be displayed on the driver's terminal device (such as a mobile phone, a tablet, etc.) so as to facilitate the driver to synchronously view the distribution of obstacles around the vehicle and the driving status of the vehicle.

[0041] In the embodiments of the present application, the driver can install an assisted driving application program on the terminal device, and through this application program, the interconnection between the vehicle machine and the terminal device can be realized. The vehicle machine can transmit the screen displayed on the vehicle machine to the terminal device in real time, and when the driver opens the assisted driving application program, the screen displayed on the vehicle machine can be seen.

[0042] In addition, the human-machine interaction interface provided by the assisted driving application program can also allow the user to operate on the screen, such as zooming in / out the screen, changing the screen direction, and so on.

[0043] In a possible implementation manner, the human-machine interaction interface of the embodiments of the present application further includes a control for selecting a driving mode, and the driving mode includes a driving mode and a parking mode. Therefore, two controls, namely "Driving" and "Parking", are provided on the human-machine interaction interface.

[0044] Among them, in the driving mode, the human-machine interaction interface further includes a control for selecting an assisted mode, and the assisted mode includes an automatic mode and a manual mode. At this time, two controls, namely "Automatic" and "Manual", are provided on the human-machine interaction interface.

[0045] In the parking mode, the human-machine interface further includes a control for selecting an auxiliary mode. The auxiliary modes include an automatic mode, a manual mode, an Automated Valet Parking (AVP) mode, and an Auto Parking Asist (APA) mode. At this time, there are four controls, namely "Automatic", "Manual", "AVP", and "APA", set on the human-machine interface. The automatic mode and the manual mode can be selected in the case of human driving (i.e., the driver drives the vehicle), and the APA mode and the AVP mode can be selected in the case of intelligent driving (i.e., the driver activates the automatic driving and the vehicle completes the parking instruction by itself. At this time, the driver can be in the vehicle or outside the vehicle). The APA mode can activate the automatic parking function. During the process of parking into the parking space, the display logic of the panoramic occupancy display area can refer to the automatic mode in the case of human driving.

[0046] Optionally, the vehicle can display the occupancy grid situation and semantic information of the obstacles around the vehicle on the central control screen with multiple grid granularities, and adaptively adjust the fineness of the human-machine interface in real time as the environment changes during the driving and parking of the vehicle, so that the driver can take preventive measures in advance to avoid scratches.

[0047] Optionally, the driver can manually determine the target area, and the vehicle can display the occupancy grid situation and semantic information of the obstacles around the vehicle on the central control screen with multiple grid granularities, so as to perform refined display on the area of interest to the driver, and adaptively adjust the fineness of the human-machine interface in real time as the environment changes during the driving and parking of the vehicle, so that the driver can take preventive measures in advance to avoid scratches.

[0048] Optionally, the panoramic perception of the entire garage / parking lot can be performed according to the pose change of the vehicle, so as to provide effective full-scene static environment information for the driver to search for a parking space or automatic parking.

[0049] In a second aspect, the present application provides an auxiliary driving device, including: an acquisition module for acquiring multi-modal perception information, where the multi-modal perception information is acquired by multiple sensors arranged on the vehicle; an acquisition module for acquiring multiple sets of display information, where the display information includes the target area to be displayed and the grid granularity; a prediction module for performing perception prediction according to the multi-modal perception information and the multiple sets of display information to obtain multiple occupancy grid information, where the occupancy grid information is used to describe the distribution of the obstacles around the vehicle; and a display module for displaying multiple occupancy grid pictures on the human-machine interface according to the multiple occupancy grid information.

[0050] In a possible implementation, the display module is specifically configured to display a first occupied grid screen corresponding to a first target area and a first grid granularity; display a second occupied grid screen corresponding to a second target area and a second grid granularity; the range of the first target area is larger than that of the second target area, and the first grid granularity is larger than the second grid granularity.

[0051] In a possible implementation, the first occupied grid screen includes grids corresponding to n1 obstacles, and the second occupied grid screen includes grids corresponding to n2 obstacles, where both n1 and n2 are positive integers.

[0052] In a possible implementation, the grids of the first obstacle in the first occupied grid screen and the grids of the second obstacle in the second occupied grid screen are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.

[0053] In a possible implementation, the grids of the third obstacle in the second occupied grid screen are identified with specific information. The third obstacle is the one closest to the vehicle among the n2 obstacles.

[0054] In a possible implementation, the display module is further configured to display a panoramic occupied grid screen on the human-machine interaction interface according to the panoramic occupied grid information, where the panoramic occupied grid information is obtained based on the historical occupied grid information and driving trajectory of the vehicle's driving environment.

[0055] In a possible implementation, the display information is pre-set information.

[0056] In a possible implementation, the display information is dynamically generated based on the operations of the driver on the touch screen.

[0057] In a possible implementation, the operations include at least one of the following operations: the operation of the driver pressing two fingers on the touch screen and sliding the two fingers apart or together; or the operation of the driver pressing two fingers on the touch screen and performing rotation, translation, or dragging.

[0058] In a possible implementation, the human-machine interaction interface further includes a control for selecting a driving mode, and the driving mode includes one or more of the following modes: driving mode or parking mode.

[0059] In a possible implementation, in the driving mode, the human-machine interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode or manual mode.

[0060] In a possible implementation, in the parking mode, the human-machine interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode, or automatic parking assist APA mode.

[0061] In a possible implementation, the prediction module is specifically configured to input the multi-modal perception information and the multiple groups of display information into a pre-trained perception prediction network to output the multiple occupancy grid information.

[0062] In a possible implementation, the functions implemented by the perception prediction network include: multi-modal feature fusion based on multi-modal perception information; feature fusion based on time series; perception prediction based on the display information.

[0063] In a possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fish-eye network, a point cloud network, or a multi-modal fusion network based on the multi-modal perception information; wherein, the pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain the pinhole 3D image features, the fish-eye network is used to extract the visual and spatial features of the fish-eye image collected by the fish-eye camera to obtain the fish-eye 3D image features, the point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter-wave radar data to obtain the point cloud voxel features, and the multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fish-eye 3D image features, and the point cloud voxel features.

[0064] In a possible implementation, the perception prediction network further includes a time series fusion network, and the time series fusion network is used to align the multi-frame 3D spatial features in sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.

[0065] In a possible implementation, the perception prediction network further includes a feature enhancement network, and the feature enhancement network is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.

[0066] In a possible implementation, the perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multiple scales of resolution according to the target area selected on the human-computer interaction interface and the grid granularity, obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction at any resolution.

[0067] In a possible implementation, the multiple sensors include a pinhole camera and a fish-eye camera.

[0068] In a possible implementation, the multiple sensors further include a lidar and / or a millimeter-wave radar.

[0069] In a third aspect, the present application provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of the above first aspects.

[0070] In a fourth aspect, the present application provides a computer-readable storage medium, including a computer program, when the computer program is executed on a computer, the computer executes the method according to any one of the above first aspects.

[0071] In a fifth aspect, the present application provides a computer program product, the computer program product includes computer program code, when the computer program code runs on a computer, the computer executes the method according to any one of the above first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 is an exemplary functional block diagram of a vehicle 100 according to an embodiment of the present application;

[0073] Figure 2 is an exemplary functional block diagram of an in-vehicle assisted driving system according to an embodiment of the present application;

[0074] Figure 3 is a software structure block diagram of a car machine according to an embodiment of the present application;

[0075] Figure 4 is a flowchart of a process 400 of an assisted driving method provided by an embodiment of the present application;

[0076] Figure 5 is a schematic flow diagram of a perception prediction process according to an embodiment of the present application;

[0077] Figure 6 is a schematic diagram of a human-computer interaction interface of a central control screen according to an embodiment of the present application;

[0078] Figure 7 Schematic diagram of switching the display perspective of the human - machine interaction interface according to an embodiment of the present application;

[0079] Figure 8 Schematic diagram of the refined mode according to an embodiment of the present application;

[0080] Figure 9 Schematic diagram of the human - machine interaction interface of the central control screen according to an embodiment of the present application;

[0081] Figure 10 Schematic diagram of the refined mode according to an embodiment of the present application;

[0082] Figure 11 and Figure 12 Schematic diagram of the zoom operation in the manual mode according to an embodiment of the present application;

[0083] Figure 13 Schematic diagram of the human - machine interaction interface of the central control screen according to an embodiment of the present application;

[0084] Figure 14 Schematic diagram of the structure of the assisted driving device 1400 according to the present application. Detailed implementation manners

[0085] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0086] The terms "first", "second", etc. in the description, claims and drawings of the present application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0087] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0088] Before describing the technical solutions of the embodiments of the present application, the vehicle of the embodiments of the present application will be described first with reference to the accompanying drawings.

[0089] Figure 1 It is a schematic functional block diagram of a vehicle 100 according to an embodiment of the present application. As Figure 1 shown, the components coupled to or included in the vehicle 100 may include a propulsion system 110, a sensor system 120, a control system 130, a peripheral device 140, a power supply 150, a computing device 160, and a driver interface 170. The components of the vehicle 100 may be configured to work in a manner that is interconnected with each other and / or interconnected with other components coupled to the respective systems. For example, the power supply 150 may supply power to all components of the vehicle 100. The computing device 160 may be configured to receive data from the propulsion system 110, the sensor system 120, the control system 130, and the peripheral device 140 and control them. The computing device 160 may also be configured to generate a display of an image on the driver interface 170 and receive input from the driver interface 170.

[0090] It should be noted that in other examples, the vehicle 100 may include more, fewer, or different systems, and each system may include more, fewer, or different components. In addition, the illustrated systems and components may be combined or divided in any manner, and the present application does not make specific limitations thereon.

[0091] The computing device 160 may include a processor 161, a transceiver 162, and a memory 163. The computing device 160 may be a controller of the vehicle 100 or a part of the controller. The memory 163 may store instructions 1631 that run on the processor 161 to perform various functional applications and data processing of the vehicle 100. It may also store data created during the use of the vehicle 100 (e.g., map data 1632, etc.), an operating system (e.g., an embedded operating system such as Android, iOS, Windows, or Linux), and application programs required for at least one function. The processor 161 included in the computing device 160 may include one or more general-purpose processors and / or one or more dedicated processors (e.g., an image processor, a digital signal processor, etc.). In the case where the processor 161 includes more than one processor, such processors may work independently or in combination. The computing device 160 may implement functions to control the vehicle 100 based on inputs received through the driver interface 170. The transceiver 162 is used for communication between the computing device 160 and various systems. The memory 163 may further include one or more volatile storage components and / or one or more non-volatile storage components, such as optical, magnetic, and / or organic storage devices, and the memory 163 may be integrated with the processor 161 in whole or in part. The memory 163 may contain instructions 1631 (e.g., program logic) executable by the processor 161 to run various vehicle functions, including any of the functions or methods described herein.

[0092] The propulsion system 110 may provide powered movement for the vehicle 100. As Figure 1 shown, the propulsion system 110 may include an engine 114, an energy source 113, a transmission 112, and wheels / tires 111. Additionally, the propulsion system 110 may additionally or alternatively include other components in addition to Figure 1 the components shown. The present application does not make specific limitations in this regard.

[0093] The sensor system 120 may include several sensors for sensing information about the environment in which the vehicle 100 is located. As Figure 1As shown, the sensors of the sensor system 120 include a Global Positioning System (GPS) 126, an Inertial Measurement Unit (IMU) 125, a lidar sensor 124, a camera sensor 123, a millimeter wave radar sensor 122, and a brake 121 for modifying the position and / or orientation of the sensors. The GPS 126 can be any sensor for estimating the geographical location of the vehicle 100. To this end, the GPS 126 can include a transceiver that estimates the position of the vehicle 100 relative to the Earth based on satellite positioning data. In an example, the computing device 160 can be used to estimate the road on which the vehicle 100 travels by using the GPS 126 in combination with the map data 1632. The IMU 125 can be used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration and any combination thereof. In some examples, the combination of sensors in the IMU 125 can include, for example, an accelerometer and a gyroscope. Additionally, other combinations of sensors in the IMU 125 are possible. The lidar sensor 124 can be regarded as an object detection system that uses optical sensing to detect objects in the environment where the vehicle 100 is located. Generally, the lidar sensor 124 can measure the distance to a target or other properties of the target through optical remote sensing technology by using light to irradiate the target. As an example, the lidar sensor 124 can include a light source and / or a laser scanner configured to emit laser pulses, and a detector for receiving the reflection of the laser pulses. For example, the lidar sensor 124 can include a laser rangefinder reflected by a rotating mirror and scan the laser around a digital scene in one dimension or two dimensions to collect distance measurements at specified angular intervals. In an example, the lidar sensor 124 can include components such as a light (e.g., laser) source, a scanner and an optical system, a light detector and receiver electronics, as well as a position and navigation system. The lidar sensor 124 determines the distance of an object by scanning the laser reflected from the object and can form a 3D environmental map with an accuracy of up to centimeter level. The camera sensor 123 can include any camera (e.g., a pinhole camera, a fish-eye camera, a static camera, a video camera, etc.) for acquiring an image of the environment where the vehicle 100 is located. To this end, the camera sensor 123 can be configured to detect visible light, or can be configured to detect light from other parts of the spectrum (such as infrared light or ultraviolet light). Other types of camera sensors 123 are possible. The camera sensor 123 can be a two-dimensional detector, or can have a three-dimensional spatial range detection function. In some examples, the camera sensor 123 can be, for example, a distance detector configured to generate a two-dimensional image indicating the distance from the camera sensor 123 to several points in the environment. To this end, the camera sensor 123 can use one or more distance detection techniques.For example, the camera sensor 123 may be configured to use structured light technology, in which the vehicle 100 irradiates an object in the environment with a predetermined light pattern, such as a grid or checkerboard pattern, and uses the camera sensor 123 to detect the reflection of the predetermined light pattern from the object. Based on the distortion in the reflected light pattern, the vehicle 100 may be configured to detect the distance to a point on the object. The predetermined light pattern may include infrared light or light of other wavelengths. The millimeter-wave radar sensor 122 generally refers to an object detection sensor with a wavelength of 1 to 10 mm, and the frequency generally ranges from 10 GHz to 200 GHz. The measurement value of the millimeter-wave radar sensor 122 has depth information and can provide the distance to the target. Secondly, due to the obvious Doppler effect of the millimeter-wave radar sensor 122, it is very sensitive to speed and can directly obtain the speed of the target. The speed of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream application frequency bands of in-vehicle millimeter-wave radars are 24 GHz and 77 GHz respectively. The former has a wavelength of about 1.25 cm and is mainly used for short-range perception, such as the environment around the vehicle body, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4 mm and is used for medium- and long-range measurement, such as automatic following, adaptive cruise control (ACC), autonomous emergency braking (AEB), etc.

[0094] The sensor system 120 may also include additional sensors, including, for example, sensors that monitor the internal systems of the vehicle 100 (e.g., O2 monitor, fuel gauge, engine oil temperature, etc.). The sensor system 120 may also include other sensors. This application does not make specific limitations in this regard.

[0095] The control system 130 may be configured to control the operation of the vehicle 100 and its components. To this end, the control system 130 may include a steering unit 136, an accelerator 135, a braking unit 134, a sensor fusion algorithm 133, a computer vision system 132, and a navigation / route control system 131. The control system 130 may additionally or alternatively include other components in addition to Figure 1 the components shown. This application does not make specific limitations in this regard.

[0096] The peripheral device 140 may be configured to allow the vehicle 100 to interact with external sensors, other vehicles, and / or drivers. To this end, the peripheral device 140 may include, for example, a lighting system 145, a wireless communication system 144, a touch screen 143, a microphone 142, and / or a speaker 141. The peripheral device 140 may additionally or alternatively include other components in addition to Figure 1 the components shown. This application does not make specific limitations in this regard.

[0097] Power source 150 may be configured to supply power to some or all components of vehicle 100. To this end, power source 150 may include, for example, a rechargeable lithium-ion or lead-acid battery. In some examples, one or more battery packs may be configured to supply power. Other power source materials and configurations are also possible. In some examples, power source 150 and energy source 113 may be implemented together, as in some all-electric vehicles.

[0098] Components of vehicle 100 may be configured to operate in a manner interconnected with other components inside and / or outside their respective systems. To this end, components and systems of vehicle 100 may be communicatively linked together via a system bus, network, and / or other connection mechanisms.

[0099] Figure 2 For the exemplary functional block diagram of the in-vehicle assisted driving system of the embodiments of this application, as Figure 2 shown, components coupled to or included in the in-vehicle assisted driving system may include a computing unit, sensors, a central control screen, a lighting system, and an audio system. Among them, the computing unit corresponds to Figure 1 control system 130 in the shown embodiment, the sensors correspond to Figure 1 sensor system 120 in the shown embodiment, mainly involving camera sensors 123 (including pinhole cameras, fisheye cameras), millimeter-wave radar sensors 122, lidar sensors 124. The central control screen corresponds to Figure 1 touch screen 143 in the shown embodiment, providing a human-machine interaction interface for the driver. The lighting system corresponds to Figure 1 lighting system 145 in the shown embodiment, and the audio system corresponds to Figure 1 speaker 141 in the shown embodiment.

[0100] The in-vehicle assisted driving system in the embodiments of this application may also be referred to as a car machine, a vehicle central control system (abbreviated as central control), etc., and no specific limitation is made thereto.

[0101] Figure 3 is the software structure block diagram of the car machine of the embodiments of this application.

[0102] The layered architecture of the car machine divides the software into several layers, and each layer has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0103] The application layer may include a series of application packages.

[0104] such as Figure 3As shown, the application package may include applications such as calls, maps, navigation, WLAN, Bluetooth, music, videos, and assisted driving.

[0105] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0106] Such as Figure 3 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.

[0107] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0108] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, maps, audio, dialed and answered calls, etc.

[0109] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The human-computer interaction interface can be composed of one or more views. For example, the human-computer interaction interface including a notification icon may include a view for displaying text and a view for displaying pictures.

[0110] The phone manager is used to provide the communication function of the in-vehicle unit. For example, the management of call status (including answering, hanging up, etc.).

[0111] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.

[0112] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without driver interaction. For example, the notification manager is used to inform that the download is complete, message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears on the screen in the form of a dialogue window. For example, prompting text information in the status bar, emitting a prompt sound, and the indicator light flashing, etc.

[0113] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for the scheduling and management of the Android system.

[0114] The core library consists of two parts: one is the functional functions that the Java language needs to call, and the other is the core library of Android.

[0115] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0116] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.

[0117] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0118] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0119] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0120] The 2D graphics engine is the graphics engine for 2D drawing.

[0121] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0122] It can be understood that Figure 3 The components shown in the system framework layer, the system library, and the runtime layer do not constitute a specific limitation on the in-vehicle computer. In other embodiments of the present application, the in-vehicle computer may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements.

[0123] Since the embodiments of the present application involve the application of neural networks, for the sake of easy understanding, some nouns or terms used in the embodiments of the present application will be explained below, and these nouns or terms are also part of the invention content.

[0124] (1) Neural network

[0125] A neural network (NN) is a machine learning model. A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs and an intercept of 1 as inputs. The output of this operation unit can be:

[0126]

[0127] where s = 1, 2, …, n, and n is a natural number greater than 1. W s is the weight of x s , and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be a region composed of several neural units.

[0128] (2) Deep neural network

[0129] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the (i + 1)-th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: where, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs the following simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and the offset vector is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscript corresponds to the third-layer index 2 of the output and the second-layer index 4 of the input. In summary, the coefficient from the k-th neuron in the (L-1)-th layer to the j-th neuron in the L-th layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. In theory, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can perform more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).

[0130] (3) Convolutional Neural Network

[0131] A convolutional neural network (CNN) is a deep neural network with a convolutional structure and a deep learning architecture. A deep learning architecture refers to multiple levels of learning at different levels of abstraction through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network in which each neuron can respond to the input image. A convolutional neural network contains a feature extractor composed of convolutional layers and pooling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane.

[0132] A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. The convolutional layer can include many convolutional operators, also known as kernels, which act as filters for extracting specific information from the input image matrix in image processing. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolutional operation on the image, the weight matrix typically processes the input image pixel by pixel (or two pixels at a time... depending on the value of the stride) along the horizontal direction to complete the work of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image, and during the convolutional operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix produces a convolution output with a single depth dimension. However, in most cases, multiple weight matrices of the same size (rows × columns), i.e., multiple matrices of the same type, are used instead of a single weight matrix. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, where the dimension can be understood as determined by the above-mentioned "multiple". Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these multiple weight matrices of the same size also have the same size. Then, the multiple feature maps of the same size that are extracted are combined to form the output of the convolutional operation. The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can be used to extract information from the input image, enabling the convolutional neural network to make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network increases, the features extracted by the subsequent convolutional layers become more and more complex, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.

[0133] Since it is often necessary to reduce the number of training parameters, a pooling layer is often periodically introduced after the convolutional layer. It can be a pooling layer following a convolutional layer, or one or more pooling layers following multiple convolutional layers. In the process of image processing, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of the pixel values in the image within a specific range as the result of average pooling. The max pooling operator can take the pixel with the maximum value within the specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after passing through the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel point in the image output by the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.

[0134] After being processed by the convolutional layer / pooling layer, the convolutional neural network is still not sufficient to output the required output information. As mentioned before, the convolutional layer / pooling layer only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network needs to use neural network layers to generate one or a set of outputs with the number of classes required. Therefore, the neural network layer can include multiple hidden layers, and the parameters contained in these multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, etc.

[0135] Optionally, after the multiple hidden layers in the neural network layer, there is also an output layer of the entire convolutional neural network. This output layer has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, backpropagation will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network, that is, the error between the result output by the convolutional neural network through the output layer and the ideal result.

[0136] (4) Recurrent Neural Network

[0137] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, it is from the input layer to the hidden layer and then to the output layer, and there is a full connection between layers, while there is no connection between each node within each layer. Although this ordinary neural network has solved many problems, it is still powerless for many issues. For example, if you want to predict what the next word in a sentence is, you generally need to use the previous words because the words before and after in a sentence are not independent. The reason why RNN is called a recurrent neural network is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes within the hidden layer itself are no longer unconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNN can process sequential data of any length. The training of RNN is the same as that of traditional CNN or DNN. The error backpropagation algorithm is also used, but there is one difference: that is, if the RNN is unfolded, the parameters in it, such as W, are shared; while the traditional neural network mentioned above is not like this. And in the use of the gradient descent algorithm, the output of each step depends not only on the network of the current step but also on the states of the networks of several previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).

[0138] Since there are already convolutional neural networks, why do we still need recurrent neural networks? The reason is very simple. In convolutional neural networks, there is a prerequisite assumption that elements are independent of each other, and the input and output are also independent, such as cats and dogs. However, in the real world, many elements are interconnected, such as the change of stocks over time, or for example, a person said: "I like traveling, and the favorite place is Yunnan. I must go there if I have a chance in the future." Here, for filling in the blank, humans should all know that it is "Yunnan". Because humans will make inferences based on the context, but how to make machines do this step, RNN came into being. RNN aims to enable machines to have the ability to remember like humans. Therefore, the output of RNN needs to depend on the current input information and historical memory information.

[0139] (5) Loss function

[0140] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we truly want to predict, we can compare the predicted value of the current network with the target value we truly desire, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, configuring parameters for each layer in the deep neural network in advance). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the target value we truly want or a value very close to the target value we truly want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0141] (6) Backpropagation algorithm

[0142] The convolutional neural network can use the error backpropagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial super-resolution model parameters are updated by backpropagating the error loss information, so as to converge the error loss. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.

[0143] (7) Generative adversarial network

[0144] Generative adversarial networks (GAN) is a deep learning model. This model includes at least two modules: one module is the Generative Model, and the other module is the Discriminative Model. Through the mutual game learning of these two modules, better outputs can be generated. Both the Generative Model and the Discriminative Model can be neural networks, specifically, they can be deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Taking the GAN for generating pictures as an example, assume there are two networks, G (Generator) and D (Discriminator). Among them, G is a network for generating pictures. It receives a random noise z and generates pictures through this noise, denoted as G(z); D is a discriminant network used to determine whether a picture is "real". Its input parameter is x, where x represents a picture, and the output D(x) represents the probability that x is a real picture. If it is 1, it means the picture is 100% real. If it is 0, it means the picture cannot be real. During the training process of this generative adversarial network, the goal of the generative network G is to generate as real pictures as possible to deceive the discriminant network D, while the goal of the discriminant network D is to try to distinguish the pictures generated by G from the real pictures. In this way, G and D constitute a dynamic "game" process, that is, the "adversarial" in the "generative adversarial network". Finally, in the ideal state, the result of the game is that G can generate pictures G(z) that are "indistinguishable from real ones", and D is difficult to determine whether the pictures generated by G are real, that is, D(G(z)) = 0.5. In this way, an excellent generative model G is obtained, which can be used to generate pictures.

[0145] In complex road conditions and driving processes, there are some scenarios that require relatively high driving skills or have relatively high driving difficulties, such as parking in a narrow parking space, passing through a narrow road, making a U-turn on a narrow road, meeting another vehicle on a narrow road, Automatic Emergency Braking (AEB), Rear Automatic Emergency Braking (RAEB), etc. For the aforementioned scenarios, some auxiliary systems on vehicles (such as 360 panoramic view) may not be able to effectively identify and display obstacles at close range (including objects, pedestrians, other vehicles, etc.), and thus cannot provide information about obstacles at close range to the driver, making it difficult to provide good and safe driving assistance capabilities for the driver. To solve this problem, this application provides an assisted driving method and device.

[0146] Figure 4This is a flowchart of process 400 for the assisted driving method provided by the embodiments of this application. Process 400 can be executed by the vehicle 100 (especially the in-vehicle computer) described above. Process 400 can be used for the low-speed driving and parking assistance function in vehicle assisted driving. Process 400 is described as a series of steps or operations. It should be understood that process 400 can be executed in various orders and / or occur simultaneously, not limited to Figure 4 the execution order shown. Process 400 may include:

[0147] Step 401, obtain multimodal perception information, which is collected by a variety of sensors provided on the vehicle.

[0148] In the embodiments of this application, the driver can trigger the vehicle to start the low-speed driving and parking assistance function in the following various ways:

[0149] 1. The driver inputs a start command through the vehicle button or the vehicle central control screen.

[0150] In a possible implementation, an assisted driving mode button is provided in the vehicle, and the driver can press the button to trigger the start of the low-speed driving and parking assistance function. It should be understood that there are various specific implementation manners for starting the target function (such as the low-speed driving and parking assistance function) through the button. For example, pressing the button corresponding to the target function, or pressing the button multiple times briefly to poll to the target function, etc. The embodiments of this application do not make specific limitations on this.

[0151] In a possible implementation, the central control screen displays the main interface, which includes an icon corresponding to the low-speed driving and parking assistance function. The driver clicks the icon to trigger the start of the low-speed driving and parking assistance function. It should be understood that there are various specific implementation manners for starting the target function (such as the low-speed driving and parking assistance function) through the central control screen (with the ability of touch screen). For example, directly clicking the icon of the target function, or selecting the option of the target function from the menu of the vehicle assisted driving mode, etc. The embodiments of this application do not make specific limitations on this.

[0152] 2. The driver inputs a start command through voice.

[0153] In a possible implementation, the driver says "Please start the low-speed driving and parking assistance function" to the speaker of the in-vehicle computer. When the in-vehicle computer receives and recognizes the semantics of the voice, the low-speed driving and parking assistance function is started.

[0154] It should be noted that in addition to the above two methods, the embodiments of this application can also use other methods to trigger the vehicle to start the low-speed driving and parking assistance function, and no specific limitations are made on this.

[0155] After the vehicle starts the low-speed driving and parking assistance function, the in-vehicle computer can first obtain multimodal perception information.

[0156] Reference Figure 1 Referring to the embodiments shown, a vehicle is provided with a variety of sensors, including a camera sensor (hereinafter referred to as a camera), a lidar sensor (hereinafter referred to as lidar), and a millimeter-wave radar sensor (hereinafter referred to as millimeter-wave radar). Among them,

[0157] The camera may include any camera for obtaining an image of the environment in which the vehicle is located (for example, a pinhole camera, a fish-eye camera, a static camera, a video camera, etc.). To this end, the camera may be configured to detect visible light, or may be configured to detect light from other parts of the spectrum (such as infrared light or ultraviolet light). Other types of cameras are also possible. The camera may be a two-dimensional detector, or may have a three-dimensional spatial range detection function. In some examples, the camera may be, for example, a distance detector configured to generate a two-dimensional image indicating the distance from the camera to several points in the environment. To this end, the camera may use one or more distance detection techniques. For example, the camera may be configured to use structured light technology, in which the vehicle irradiates an object in the environment with a predetermined light pattern, such as a grid, and uses the camera to detect the reflection of the predetermined light pattern from the object. Based on the distortion in the reflected light pattern, the vehicle may be configured to detect the distance to the points on the object. The predetermined light pattern may include infrared light or light of other wavelengths.

[0158] Lidar can be regarded as an object detection system that uses optical sensing to detect objects in the environment where the vehicle is located. Generally, lidar can measure the distance to a target or other attributes of the target through optical remote sensing technology by irradiating the target with light. As an example, lidar may include a laser source and / or a laser scanner configured to emit laser pulses, and a detector for receiving the reflection of the laser pulses. For example, lidar may include a laser rangefinder reflected by a rotating mirror and scan the laser in one or two dimensions around a digital scene to collect distance measurements at specified angular intervals. In an example, lidar may include components such as a light (for example, laser) source, a scanner and an optical system, a light detector and receiver electronics, as well as a position and navigation system. Lidar determines the distance of an object by scanning the laser reflected from the object and can form a 3D environmental map with an accuracy of up to centimeter level. Millimeter-wave radar generally refers to an object detection sensor with a wavelength of 1-10 mm and a frequency range of approximately 10 GHz to 200 GHz.

[0159] The measurement values of millimeter-wave radars have depth information and can provide the distance to the target. Secondly, due to the obvious Doppler effect of millimeter-wave radars, they are very sensitive to speed and can directly obtain the speed of the target. The speed of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream application frequency bands of in-vehicle millimeter-wave radars are 24 GHz and 77 GHz respectively. The former has a wavelength of about 1.25 cm and is mainly used for short-distance perception, such as the surrounding environment of the vehicle body, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4 mm and is used for medium- and long-distance measurement, such as automatic following, adaptive cruise control (ACC), autonomous emergency braking (AEB), etc.

[0160] In the embodiments of the present application, the multi-modal perception information can be obtained based on the information collected by one or more of the above-mentioned multiple sensors.

[0161] Step 402: Obtain multiple groups of display information, where the display information includes the target area to be displayed and the grid granularity.

[0162] A group of display information may include the target area to be displayed and the grid granularity. Among them, the target area may refer to the range of the area of the real world displayed on the central control screen. For example, in the real world, the area within 3 m around the vehicle (i.e., the host vehicle), centered on the vehicle (i.e., the host vehicle), a 6 m×6 m square plane area 3 m forward, 3 m backward, 3 m left, and 3 m right, and then a 2 m height is added to form a 6 m×6 m×2 m three-dimensional space, and this three-dimensional space is the target area. The length, width, and height of the target area can be expressed as W (length)×H (width)×D (height). When representing the target area, the host vehicle can be used as the origin, and the three-dimensional coordinates of the 8 vertices of the target area can be used to represent the range of the target area. Alternatively, the three-dimensional coordinates of a certain vertex of the target area and the length, width, and height of the target area can also be used to represent the range of the target area. The embodiments of the present application do not make specific limitations on the representation form of the range of the target area.

[0163] The grid granularity refers to the principle of occupancy grid. Occupancy grid means that the three-dimensional space is divided in the form of voxels, and each voxel is characterized by binary values 0 and 1 to indicate whether the voxel is occupied. It can be seen that a voxel can correspond to a grid, and the grid granularity can be used to represent the granularity of dividing the three-dimensional space into voxels. For example, the size of a voxel is 1 m×1 m×1 m, and the above 6 m×6 m×2 m three-dimensional space can be divided into 6×6×2 voxels, that is, this three-dimensional space corresponds to 6×6×2 grids.

[0164] Optionally, the larger the range of the target area, the larger the grid granularity that can be selected, and the smaller the range of the target area, the smaller the grid granularity that can be selected. In this way, more occupancy grid information can be displayed in a larger target area, while more detailed occupancy grid information can be displayed in a smaller target area.

[0165] In a possible implementation, in a set of display information, the target area may refer to the area within 6 m around the host vehicle, that is, the length, width, and height of the target area are 12 m × 12 m × 2 m, and the grid granularity may be 20 cm.

[0166] In a possible implementation, in a set of display information, the target area may refer to the area within 3 m around the host vehicle, that is, the length, width, and height of the target area are 6 m × 6 m × 2 m, and the grid granularity may be 5 cm.

[0167] In a possible implementation, in a set of display information, the target area may refer to the area within 1 m around the host vehicle, that is, the length, width, and height of the target area are 2 m × 2 m × 2 m, and the grid granularity may be 2 cm.

[0168] In the embodiments of the present application, the corresponding relationship between the target area and the grid granularity may be a pre-established relationship (which can be used as a default recommended configuration), or the corresponding relationship between the target area and the grid granularity may be determined in real time. No specific limitation is made thereto.

[0169] In the embodiments of the present application, the vehicle-mounted computer can obtain display information through the following multiple methods:

[0170] 1. The display information is pre-set information.

[0171] Multiple sets of display information can be pre-set and stored. When the display information is needed, the vehicle-mounted computer can directly read the memory to obtain the multiple sets of display information. For example, two sets of display information are set. In one set of display information, the target area refers to the area within 6 m around the host vehicle, that is, the length, width, and height of the target area are 12 m × 12 m × 2 m, and the grid granularity is 20 cm. In the other set of display information, the target area refers to the area within 3 m around the host vehicle, that is, the length, width, and height of the target area are 6 m × 6 m × 2 m, and the grid granularity is 5 cm.

[0172] 2. The display information is dynamically generated based on the operations of the driver on the touch screen.

[0173] The above zoom operation may include an operation in which the driver presses the touch screen with two fingers and slides the two fingers apart or slides the two fingers together inward, for changing the range of the target area and the target resolution of the obstacle; or, the above zoom operation may include an operation in which the driver presses the touch screen with two fingers and performs rotation, translation or dragging, for changing the display orientation and range of the target area. The foregoing operations may refer to the situation when the driver uses a picture application. To magnify the details of a picture, the driver may press the relevant position with the thumb and index finger and then slide the two fingers apart; or to shrink the picture to view the whole picture, the driver may press the relevant position with the thumb and index finger and then slide the two fingers together inward; or, to view the picture content in a different direction, the driver may press the relevant position with two fingers and then drag the picture to translate or rotate the picture direction.

[0174] In an embodiment of the present application, a human-machine interaction interface is provided for the driver. On the initial target area, the driver can autonomously zoom in / zoom out the target area through the human-machine interaction interface, so as to obtain the final target area and the grid granularity corresponding to the target area. The initial target area and its corresponding grid granularity, as well as the final target area and its corresponding grid granularity, constitute the above-mentioned multiple groups of display information.

[0175] Step 403: Perform perception prediction based on the multi-modal perception information and the multiple groups of display information to obtain multiple occupancy grid information.

[0176] In an embodiment of the present application, the multi-modal perception information and the multiple groups of display information may be input into a pre-trained perception prediction network to output multiple occupancy grid information.

[0177] The above perception prediction network may refer to the content of the neural network above, and its principle and specific implementation may both refer to its content. The perception prediction network may be pre-trained and downloaded to the vehicle-mounted computer. Then, when using the perception prediction network, the parameters in the perception prediction network may be updated based on the real-time input and output, so that the prediction result of the perception prediction network is more accurate, and thus occupancy grid information more in line with the actual driving situation can be obtained to assist the driver in driving the vehicle more safely.

[0178] Exemplarily, Figure 5 is a schematic flowchart of the perception prediction process of an embodiment of the present application. As Figure 5 shown, the perception prediction process (corresponding to the perception prediction network) includes the following steps:

[0179] (1) The multi-modal perception information obtained by the in-vehicle computer includes M pinhole images (from pinhole cameras) and N fisheye images (from fisheye cameras). Feature extraction is performed on the M pinhole images and N fisheye images respectively, and according to the poses calibrated by the cameras that collected each image, the image features are transformed from the 2D perspective view (PV) perspective to the 3D space.

[0180] That is to say, the perception prediction network includes a pinhole network and a fisheye network based on multi-modal perception information; among them, the pinhole network is used to extract the visual and spatial features of the pinhole images collected by the pinhole cameras to obtain pinhole 3D image features, and the fisheye network is used to extract the visual and spatial features of the fisheye images collected by the fisheye cameras to obtain fisheye 3D image features.

[0181] a. The resolutions of the pinhole images and the fisheye images can be different, so the pinhole images and the fisheye images are separated and feature extraction is performed separately. Among the M pinhole images, feature extraction can also be performed on the pinhole images collected respectively according to the positions of the pinhole cameras (for example, in front of the vehicle, behind the vehicle, on the side of the vehicle). Among the N fisheye images, feature extraction can also be performed on the fisheye images collected respectively according to the positions of the fisheye cameras (for example, in front of the vehicle, behind the vehicle, on the side of the vehicle).

[0182] The above feature extraction of the images can adopt a method based on a convolutional neural network (CNN), such as a residual network (Resnet) or a feature pyramid network (FPN), etc., or, it can also be a method based on a Transformer structure, such as a Vision Transformer.

[0183] b. Based on the feature extraction, a feature map pv_feat in the unified PV perspective is obtained, with a dimension of C×H×W, where C is the number of feature channels, and H and W are the image sizes.

[0184] c. Consistent with step a, a neural network model is used to perform monocular depth estimation on the depth structure information of the images, so as to obtain a depth distribution map of the images, with a dimension of D×H×W, where D is the number of categories of the depth distribution.

[0185] d. The above depth distribution map is used to perform weighted mapping (such as broadcast multiplication) on the above pv_feat to obtain a depth visual feature map, with a dimension of C×D×H×W. And according to the internal and external camera parameter information (such as the position, focal length, etc. of the camera), the features in the depth visual feature map are projected to different positions in the 3D space to obtain a 3D feature map, with a dimension of C1×X×Y×Z, where X, Y, and Z are the length, width, and height in the three-dimensional space.

[0186] Referring to the above steps a - b, the M pinhole images and N fisheye images can be processed respectively to obtain their respective 3D feature maps. It should be noted that the dimensions of the 3D feature maps obtained from different frame images can be exactly the same, not exactly the same, or completely different, and no specific limitation is made in this regard.

[0187] (2) The multi-modal perception information obtained by the vehicle-mounted computer includes lidar point cloud (from lidar) and millimeter-wave point cloud (from millimeter-wave radar), and feature extraction is performed on the lidar point cloud and the millimeter-wave point cloud respectively.

[0188] That is to say, the perception and prediction network includes a point cloud network based on multi-modal perception information; the point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter-wave radar data to obtain point cloud voxel features.

[0189] a. For the lidar point cloud, voxelization processing is performed using the correspondence between the point cloud and the grid (i.e., the correspondence between the 3D points in the point cloud and X, Y, Z in the three-dimensional space). Each voxel can include point cloud features such as the number of point clouds, the maximum height of the 3D points, the minimum height of the 3D points, and the point cloud reflection intensity. After convolution operation, a lidar feature map is obtained.

[0190] b. For the millimeter-wave point cloud, voxelization processing is performed using the correspondence between the point cloud and the grid (i.e., the correspondence between the 3D points in the point cloud and X, Y, Z in the three-dimensional space), and after convolution operation, a millimeter-wave radar (radar) feature map is obtained.

[0191] c. The lidar feature map and the radar feature map are feature-fused (for example, spliced, added, etc.) to obtain a radar feature map with dimensions of C2×X×Y×Z.

[0192] (3) Multi-modal feature fusion is performed on the 3D feature map, lidar feature map, and radar feature map in the same three-dimensional space.

[0193] That is to say, the perception and prediction network includes a multi-modal fusion network based on multi-modal perception information; the multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, fisheye 3D image features, and point cloud voxel features.

[0194] a. The 3D feature maps corresponding to the M pinhole images and N fisheye images are weighted-fused respectively to obtain an image 3D feature map.

[0195] b. The radar feature map and the image 3D feature map are post-fused to obtain a fused 3D space feature map.

[0196] The above steps correspond to the function of temporal feature fusion of the perception and prediction network. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitations are imposed thereon.

[0197] (4) Further perform temporal alignment and spatial feature extraction on the features of the fused 3D spatial feature map.

[0198] That is to say, the perception and prediction network further includes a temporal fusion network, which is used to align multi-frame 3D spatial features in sequence to the same coordinate system based on the vehicle's motion parameters, and then perform multi-frame weighted fusion.

[0199] a. For the fused 3D spatial feature map of the current frame and the 3D spatial feature map of the historical T-1 frame, perform coordinate transformation on the 3D features of the historical frame according to the vehicle motion parameters (ego motion) so that they are in the same coordinate system as the 3D features at the current moment, and obtain the 3D spatial feature map of the T frames aligned to the same coordinate system.

[0200] b. Fuse the 3D spatial feature maps of the T frames in sequence into a 3D spatial feature map of the current frame.

[0201] c. Use a multi-stage feature extractor for the 3D spatial feature map of the current frame to extract spatial features again and obtain an enhanced 3D spatial feature map.

[0202] The above steps correspond to the function of multi-modal feature fusion of the perception and prediction network based on multi-modal perception information. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitations are imposed thereon.

[0203] (5) Perform multiple upsamplings on the enhanced 3D spatial feature map and fuse image features again for auxiliary feature enhancement;

[0204] That is to say, the perception and prediction network further includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse image features again to achieve auxiliary feature enhancement.

[0205] a. Pass the enhanced 3D spatial feature map through multiple upsamplings to obtain multiple resolutions corresponding to the above-mentioned multiple grid granularities;

[0206] b. Pass the image 3D feature map through multiple upsamplings to obtain multiple resolutions corresponding to the above-mentioned multiple grid granularities;

[0207] c. At the same resolution, perform weighted fusion on the features of the upsampled enhanced 3D spatial feature map and the upsampled image 3D feature map to obtain the 3D spatial feature map at this resolution;

[0208] (6) Crop sampling and weighted fusion are performed on the 3D spatial feature map at multiple scale resolutions according to the target area and grid granularity, and position features are obtained through interpolation to achieve any resolution.

[0209] That is, the perception and prediction network further includes a dynamic resolution prediction network based on display information feedback. The dynamic resolution prediction network is used to perform crop sampling and weighted fusion on the 3D spatial feature map at multiple scale resolutions according to the target area selected on the human-machine interaction interface and the grid granularity, obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction at any resolution.

[0210] a. Calculate the relative coordinates of the target area in the entire 3D feature space, and perform cropping on the 3D spatial feature map at multiple scale resolutions according to the relative coordinates.

[0211] b. Generate the coordinates of all positions that need to predict occupancy within the target area according to the grid granularity (i.e., its corresponding resolution).

[0212] c. Perform feature sampling and weighted fusion on the cropped 3D spatial feature map at multiple scale resolutions according to the coordinates of the positions that need to predict occupancy. The feature sampling is obtained through trilinear interpolation of adjacent features on the feature map.

[0213] (7) Perform Multilayer Perceptron (MLP) prediction on all the obtained position features, and then output the occupancy situation and semantic information of the grid at the corresponding coordinate positions. Among them, the occupancy situation is a binary classification prediction, that is, only two categories of occupied / unoccupied, and the semantic information is a multi-classification prediction, including road surface / moving target / other obstacles, etc.

[0214] For a single set of display information (including the target area and grid granularity), the occupancy situation and semantic information of the grids at all coordinate positions within the target area are the occupancy grid information corresponding to the display information. For example, for a set of display information (the target area can refer to the area within 6m around the vehicle, that is, the length, width, and height of the target area are 12m×12m×2m, and the grid granularity can be 20cm), the occupancy situation and semantic information of the grids at all coordinate positions within the area within 6m around the vehicle can be perceived and predicted, and the grid occupancy situation can be displayed at a grid granularity of 20cm; or, for another set of display information (the target area can refer to the area within 3m around the vehicle, that is, the length, width, and height of the target area are 6m×6m×2m, and the grid granularity can be 5cm), the occupancy situation and semantic information of the grids at all coordinate positions within the area within 3m around the vehicle can be perceived and predicted, and the grid occupancy situation can be displayed at a grid granularity of 5cm.

[0215] The above steps correspond to the function of perception and prediction based on display information in the perception and prediction network. It should be noted that, in addition to the above steps, this function can also be implemented in other ways, and no specific limitation is made thereto.

[0216] It should be noted that Figure 5 The illustrated embodiments only describe one implementation method of perception and prediction. The embodiments of the present application can also use other methods for perception and prediction, and no specific limitation is made thereto.

[0217] Step 404: Display a plurality of occupied grid pictures on the human-machine interaction interface according to a plurality of occupied grid information.

[0218] In a possible implementation manner, the vehicle-mounted computer can display a first occupied grid picture on the central control screen. The first occupied grid picture corresponds to a first target area and a first grid granularity; display a second occupied grid picture, and the second occupied grid picture corresponds to a second target area and a second grid granularity; wherein, the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.

[0219] Optionally, the above first occupied grid picture includes grids corresponding to n1 obstacles, and the above second occupied grid picture includes grids corresponding to n2 obstacles, where both n1 and n2 are positive integers.

[0220] The first occupied grid picture displays the picture of the first target area, which displays the grids of n1 obstacles perceived and predicted in the first target area with the first grid granularity; the second occupied grid picture displays the picture of the second target area, which displays the grids of n2 obstacles perceived and predicted in the second target area with the second grid granularity.

[0221] Optionally, the grids of the first obstacle in the above first occupied grid picture and the grids of the second obstacle in the above second occupied grid picture are marked with the same information. The first obstacle and the second obstacle refer to the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles include the second obstacle.

[0222] Both the first target area and the second target area are areas around the vehicle. The difference between them lies in the different ranges. Among them, the range of the first target area is larger than the range of the second target area. Therefore, the first target area and the second target area may include the same obstacles. In other words, the obstacles in the second target area may also exist in the first target area, and some obstacles in the first target area may not appear in the second target area.

[0223] For the convenience of driver identification, in the embodiments of the present application, the grids of different obstacles can be displayed in different colors, different wireframes, etc., or the grids of different obstacles can be marked with different information (such as text, numbers, etc.). In this way, the driver can intuitively see the number, position, type, etc. of the obstacles included in the target area, and then take avoidance measures in a timely manner.

[0224] Based on this, in the first occupancy grid screen and the second occupancy grid screen, the grids corresponding to the same obstacle can be identified with the same information. For example, in the first occupancy grid screen and the second occupancy grid screen, the grids of the first obstacle and the second obstacle are respectively displayed in the same color (such as a yellow grid), the same wireframe (such as a solid line frame), or, in the first occupancy grid screen and the second occupancy grid screen, the grids of the first obstacle and the second obstacle (the obstacles existing in both the first target area and the second target area) are marked with the same information (such as the same text, the same number, etc.). The aforementioned first obstacle and second obstacle refer to the same obstacle. In this way, the driver can intuitively see which are the same obstacles in the target areas of different ranges, and compare the obstacle distributions and distances in different fields of view, and then take avoidance measures in a timely manner.

[0225] Optionally, the grids of the third obstacle in the above-mentioned second occupancy grid screen are identified with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.

[0226] For the convenience of driver identification, in the embodiments of the present application, the grids of the obstacle closest to the vehicle around the vehicle can be identified with specific information. For example, the grids of the obstacle closest to the vehicle are highlighted in a special color (red), or the grids of the obstacle closest to the vehicle are displayed in a flashing manner, or the text indicating that it is the closest obstacle is marked near the grids of the obstacle closest to the vehicle, or the distance between it and the vehicle is marked with a number near the grids of the obstacle closest to the vehicle. Further, when the distance between the vehicle and the closest obstacle is less than a set threshold (such as 0.5 m), the vehicle's audio and lights will give a synchronous warning.

[0227] It should be noted that in the embodiments of the present application, other ways can also be used to display the grids of the obstacles, and no specific limitations are made in this regard.

[0228] Exemplarily, Figure 6 is a schematic diagram of the human-machine interaction interface of the central control screen in the embodiments of the present application. As Figure 6 shown, the human-machine interaction interface is composed of a surround-view image display area and a panoramic occupancy display area. Among them,

[0229] The surround-view image display area mainly displays the image information collected by the pinhole cameras and fisheye cameras around the vehicle itself.

[0230] The panoramic occupancy display area includes two areas on the left and right. The left area displays a relatively wide range of target areas and coarse-grained occupancy grid perception information, which can provide the driver with a relatively complete distribution of surrounding obstacles. The right area displays a relatively small range of target areas and fine-grained occupancy grid perception information, which can provide the driver with a refined contour of nearby obstacles, thereby intuitively showing the distance between the vehicle itself and nearby obstacles.

[0231] The above-mentioned left area corresponds to the above-mentioned first occupancy grid screen, and the above-mentioned right area corresponds to the above-mentioned second occupancy grid screen.

[0232] The above two occupancy grid screens display the target areas in two ranges with different grid granularities. The relatively wide range of target areas can improve the driver's panoramic perception ability of obstacles in a large range, and the relatively small range of target areas can improve the driver's refined perception ability of obstacles in a short range. The combination of the two can improve the driver's perception ability of the surrounding environment and avoid rubbing.

[0233] It should be noted that the above embodiments are described by taking two occupancy grid screens as an example. The number of occupancy grid screens displayed on the central control screen in the embodiments of the present application is not specifically limited, and can be equal to or greater than two. Each occupancy grid screen corresponds to a display information, that is, a target area is displayed with one grid granularity.

[0234] In a possible implementation manner, the screen of the above-mentioned human-computer interaction interface can be displayed on the driver's terminal device (for example, mobile phone, tablet, etc.) so as to facilitate the driver to synchronously view the distribution of obstacles around the vehicle and the driving status of the vehicle.

[0235] In the embodiments of the present application, the driver can install an assisted driving application program on the terminal device, and through this application program, the interconnection between the vehicle machine and the terminal device can be realized. When the vehicle machine executes step 404, the screen displayed on the vehicle machine can be transmitted to the terminal device in real time. When the driver opens the assisted driving application program, the screen displayed on the vehicle machine can be seen, referring to Figure 6 the embodiments shown.

[0236] In addition, the human-computer interaction interface provided by the assisted driving application program can also allow the user to operate on the screen, for example, zoom in / out the screen, change the screen direction, etc.

[0237] In the embodiments of the present application, by displaying multiple occupancy grid images corresponding to multiple grid granularities on the center control screen of the vehicle, the multiple occupancy grid images display target areas of different ranges with different grid granularities, which can not only improve the driver's panoramic perception ability of obstacles in a large range, but also improve the driver's refined perception ability of obstacles in a short range. The combination of the two can improve the driver's perception ability of the surrounding environment and avoid scratching.

[0238] In a possible implementation manner, in addition to the Figure 6 image shown in the embodiment, the human-machine interaction interface of the embodiments of the present application further includes controls for selecting driving modes, and the driving modes include a driving mode and a parking mode. Therefore, two controls, namely "Driving" and "Parking", are provided on the human-machine interaction interface.

[0239] Among them, in the driving mode, the human-machine interaction interface further includes controls for selecting auxiliary modes, and the auxiliary modes include an automatic mode and a manual mode. At this time, two controls, namely "Automatic" and "Manual", are provided on the human-machine interaction interface.

[0240] In the parking mode, the human-machine interaction interface further includes controls for selecting auxiliary modes. The auxiliary modes include an automatic mode, a manual mode, an Automated Valet Parking (AVP) mode, and an Auto Parking Asist (APA) mode. At this time, four controls, namely "Automatic", "Manual", "AVP", and "APA", are provided on the human-machine interaction interface. The automatic mode and the manual mode can be selected in the case of human driving (i.e., the driver drives the vehicle), and the APA mode and the AVP mode can be selected in the case of intelligent driving (i.e., the driver turns on the automatic driving, and the vehicle completes the parking instruction by itself. At this time, the driver can be in the vehicle or outside the vehicle). The APA mode can activate the automatic parking function. During the process of parking into the parking space, the display logic of the panoramic occupancy display area can refer to the Figure 9 automatic mode in the case of human driving shown in the embodiment.

[0241] In one embodiment, in the scenario of passing through a narrow road at a low speed, the implementation process of the assisted driving method is as follows:

[0242] 1. As shown in Figure 6 , the driver starts the low-speed driving and parking assistance function through the vehicle button or the icon on the center control screen of the vehicle, and then selects the "Driving" mode in the human-machine interaction interface displayed on the center control screen.

[0243] 2. The vehicle processes the data collected by the sensors (including pinhole cameras, fisheye cameras, lidar, and millimeter-wave radars) around the vehicle itself.

[0244] 3. Based on the sensor data collected by the vehicle, the environment around the host vehicle is perceived and predicted in real time. In the embodiments of the present application, this step is implemented through a perception and prediction network. The input of the perception and prediction network includes: 1) multi-modal perception information, including pinhole images, fisheye images, laser point clouds, millimeter-wave point clouds, etc.; 2) multiple display information (including target areas and grid granularities); the output of the perception and prediction network includes multiple occupancy grid information (including the occupancy situation and semantic information of the grids at all coordinate positions within the target area).

[0245] The implementation process of the perception and prediction network can refer to Figure 5 the embodiments shown, which will not be elaborated here.

[0246] 4. The vehicle performs real-time display on the central control screen according to multiple occupancy grid information.

[0247] As Figure 6 shown, the human-machine interaction interface is composed of a surround-view image display area and a panoramic occupancy display area. Among them, the surround-view image display area mainly displays the image information collected by the pinhole camera and fisheye camera around the host vehicle. The panoramic occupancy display area includes two areas on the left and right. The left area displays a wider range of target areas and coarse-grained occupancy grid perception information, which can provide the driver with a more comprehensive distribution of surrounding obstacles. The right area displays a smaller range of target areas and fine-grained occupancy grid perception information, which can provide the driver with a refined contour of nearby obstacles, and thus intuitively show the distance between the host vehicle and nearby obstacles.

[0248] Optionally, the low-speed driving and parking assistance function of the embodiments of the present application can support the driver to adjust the viewing angles of different grid displays. For example, Figure 7 ( Figure 7 is a schematic diagram of switching the display viewing angle of the human-machine interaction interface in the embodiments of the present application) shown, the driver can perform a drag and rotation operation on the central control screen to adjust to the top-down bird's-eye view (Bird's Eye View, BEV) angle. This can allow the driver to experience the driving experience of viewing the distribution of surrounding obstacles from multiple perspectives.

[0249] 5. The low-speed driving and parking assistance function of the embodiments of the present application can provide two ways to perform refined display on a specific area, namely "automatic" and "manual" modes. For example, Figure 8 ( Figure 8 is a schematic diagram of the refined mode in the embodiments of the present application) shown, the two modes can be switched in the upper right corner of the human-machine interaction interface. The driver can click the gesture button to switch to the "automatic" mode.

[0250] Optionally, the low-speed driving and parking assistance function can be default set to the "Automatic" mode. The "Automatic" mode can correspond to the situation where "the displayed information is pre-set information" in step 402 above.

[0251] In this embodiment, in the "Driving" and "Automatic" modes, in the human-machine interaction interface of the central control screen, the occupancy grid information of the area within 6m around the vehicle (i.e., the target area) is displayed in the left area of the panoramic occupancy display area in a coarse grid granularity of 20cm; the occupancy grid information of the area within 3m around the vehicle (i.e., the target area) is displayed in the right area of the panoramic occupancy display area in a fine grid granularity of 5cm.

[0252] Through the occupancy grid information of the obstacles displayed on the central control screen, the driver can know which direction of the vehicle has obstacles, the distance between the vehicle and each obstacle, and which position of the vehicle is already close to the obstacles, etc., so as to take preventive measures in advance to avoid rubbing.

[0253] 6. As Figure 8 shown, the distance between the obstacle closest to the vehicle and the vehicle can also be displayed on the central control screen, so that the driver can more intuitively see how far the closest obstacle is, and take corresponding driving means to avoid the obstacle and avoid rubbing.

[0254] In this embodiment, the vehicle can display the occupancy grid situation and semantic information of the obstacles around the vehicle on the central control screen with multiple grid granularities, and adaptively adjust the fineness of the human-machine interaction interface in real time during the driving and parking process of the vehicle, so that the driver can take preventive measures in advance to avoid rubbing.

[0255] In another embodiment, in the scenario of parking in a narrow parking space, the implementation process of the assisted driving method is as follows:

[0256] 1. As Figure 9 ( Figure 9 is a schematic diagram of the human-machine interaction interface of the central control screen of the embodiment of the present application), the driver activates the low-speed driving and parking assistance function through the vehicle button or the icon on the central control screen of the vehicle, and then selects the "Parking" mode in the human-machine interaction interface displayed on the central control screen.

[0257] Steps 2-4 of this embodiment can refer to Figure 6 Steps 2-4 of the shown embodiment, which will not be elaborated here.

[0258] 5. In the case of human driving, the low-speed driving and parking assistance function of the embodiment of the present application can provide two ways to finely display a specific area, namely the "Automatic" and "Manual" modes, as Figure 10 ( Figure 10As shown in the schematic diagram of the refined mode of the embodiment of the present application, the two modes can be switched in the upper right corner of the human-machine interaction interface. The driver can click the gesture button to switch to the "manual" mode. The "manual" mode can correspond to the situation where "the displayed information is dynamically generated based on the driver's operation on the touch screen" in step 402 above.

[0259] In this embodiment, in the "parking" and "manual" modes, as Figure 10 shown, in the initial state, in the human-machine interaction interface of the central control screen, the occupancy grid information of the area within 6m around the vehicle (i.e., the target area) is displayed in the left and right areas of the panoramic occupancy display area in the form of a coarse grid granularity of 20cm. And the right area can be provided for the driver to perform zoom operations thereon.

[0260] As Figure 11 and Figure 12 ( Figure 11 and Figure 12 As shown in the schematic diagram of the zoom operation in the manual mode of the embodiment of the present application), the driver can manually move and zoom the right area of the panoramic occupancy display area on the human-machine interaction interface. The vehicle-mounted computer extracts the zoomed area and multiple according to the driver's operation, and uses the zoomed area as the target area. This multiple can be used as the basis for obtaining the grid granularity. The driver locates the target area within 4m around the vehicle (i.e., the target area) through the zoom operation. At this time, in the human-machine interaction interface of the central control screen, the occupancy grid information of the area within 4m around the vehicle (i.e., the target area) is displayed in the right area of the panoramic occupancy display area in the form of a coarse grid granularity of 8cm.

[0261] In this embodiment, the driver can manually determine the target area, and the vehicle can display the occupancy grid situation and semantic information of the obstacles around the vehicle on the central control screen with multiple grid granularities, so as to perform refined display on the area of interest to the driver, and adaptively adjust the fineness of the human-machine interaction interface in real time as the environment changes during the driving and parking processes of the vehicle, enabling the driver to take preventive measures in advance to avoid scratches.

[0262] In another embodiment, the vehicle-mounted computer can display a panoramic occupancy grid picture on the human-machine interaction interface according to the panoramic occupancy grid information, which is obtained based on the historical occupancy grid information and driving trajectory of the vehicle's driving environment.

[0263] In the AVP garage / parking lot parking scenario, the implementation process of the assisted driving method is as follows:

[0264] 1. As Figure 13 ( Figure 13As shown in the schematic diagram of the human-machine interaction interface of the central control screen in the embodiment of the present application, after the vehicle enters the garage / parking lot, the driver activates the low-speed parking and driving assistance function through the vehicle buttons or the icons on the central control screen of the vehicle, and then selects the "Parking" mode and the AVP mode in the human-machine interaction interface displayed on the central control screen.

[0265] Step 2-3 of this embodiment can refer to Figure 6 Steps 2-3 of the embodiment shown, which will not be elaborated here.

[0266] 4. As the vehicle travels in the garage / parking lot, multi-frame result fusion is performed on the predicted occupancy grid situation and semantic information, so as to realize the reconstruction of the panoramic static structure. The steps are as follows:

[0267] (1) Remove dynamic targets according to the predicted semantic information, and only retain static structure obstacles;

[0268] (2) Convert the remaining occupancy grids to a unified world coordinate system according to the vehicle's pose;

[0269] (3) Accumulate the occupancy grid information in the world coordinate system of the historical N frames, perform probability weighted fusion, and obtain the fused 3D occupancy grid information (corresponding to the panoramic occupancy grid information);

[0270] (4) Convert the 3D occupancy grid information to the vehicle's own coordinate system at the current moment;

[0271] (5) Similarly convert the historical trajectory points of the vehicle to the vehicle's own coordinate system at the current moment;

[0272] 5. Perform real-time display and output according to the 3D occupancy grid information in the vehicle's own coordinate system. Refer to Figure 7 and Figure 8 In the embodiment shown, the driver can perform drag and zoom operations on the human-machine interaction interface on the central control screen.

[0273] 6. In the AVP mode, the driver can view the basement reconstruction result, vehicle driving and parking trajectory through the assisted driving application program on the terminal device (such as mobile phone, tablet, etc.), and its picture can be the same as the picture of the vehicle computer.

[0274] This embodiment can perform panoramic perception of the entire garage / parking lot according to the change of the vehicle's pose, so as to provide effective full-scene static environment information for the driver to search for parking or automatic parking.

[0275] Figure 14 It is a schematic diagram of the structure of the assisted driving device 1400 of the present application, as Figure 14As shown in the figure, the assisted driving device 1400 of this embodiment can be applied to the above-mentioned terminal device. The assisted driving device 1400 may include: an acquisition module 1401, an acquisition module 1402, a prediction module 1403, and a display module 1404. Among them,

[0276] The acquisition module 1401 is used to obtain multi-modal perception information, and the multi-modal perception information is acquired by a variety of sensors arranged on the vehicle; the acquisition module 1402 is used to obtain multiple groups of display information, and the display information includes a target area to be displayed and a grid granularity; the prediction module 1403 is used to perform perception prediction based on the multi-modal perception information and the multiple groups of display information to obtain multiple occupancy grid information, and the occupancy grid information is used to describe the distribution of obstacles around the vehicle; the display module 1404 is used to display multiple occupancy grid pictures on the human-computer interaction interface according to the multiple occupancy grid information.

[0277] In a possible implementation manner, the display module 1404 is specifically used to display a first occupancy grid picture, which corresponds to a first target area and a first grid granularity; display a second occupancy grid picture, which corresponds to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.

[0278] In a possible implementation manner, the first occupancy grid picture includes grids corresponding to n1 obstacles, and the second occupancy grid picture includes grids corresponding to n2 obstacles, where n1 and n2 are both positive integers.

[0279] In a possible implementation manner, the grids of the first obstacle in the first occupancy grid picture and the grids of the second obstacle in the second occupancy grid picture are marked with the same information, the first obstacle and the second obstacle are the same obstacle, the n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.

[0280] In a possible implementation manner, the grids of the third obstacle in the second occupancy grid picture are marked with specific information, and the third obstacle is the one closest to the vehicle among the n2 obstacles.

[0281] In a possible implementation manner, the display module 1404 is further used to display a panoramic occupancy grid picture on the human-computer interaction interface according to the panoramic occupancy grid information, and the panoramic occupancy grid information is obtained based on the historical occupancy grid information and the driving trajectory of the vehicle's driving environment.

[0282] In a possible implementation, the display information is pre-set information.

[0283] In a possible implementation, the display information is dynamically generated based on the operations of the driver on the touch screen.

[0284] In a possible implementation, the operations include at least one of the following operations: the operation of the driver pressing two fingers on the touch screen and sliding the two fingers apart or sliding the two fingers together; or, the operation of the driver pressing two fingers on the touch screen and performing rotation, translation or dragging.

[0285] In a possible implementation, the human-machine interaction interface further includes a control for selecting a driving mode, and the driving mode includes one or more of the following modes: driving mode or parking mode.

[0286] In a possible implementation, in the driving mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode or manual mode.

[0287] In a possible implementation, in the parking mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode or automatic parking assist APA mode.

[0288] In a possible implementation, the prediction module 1403 is specifically configured to input the multi-modal perception information and the multi-group display information into a pre-trained perception prediction network to output the multiple occupancy grid information.

[0289] In a possible implementation, the functions implemented by the perception prediction network include: multi-modal feature fusion based on multi-modal perception information; feature fusion based on time series; perception prediction based on the display information.

[0290] In a possible implementation, the perception prediction network includes one or more of the following networks: a pinhole network, a fish-eye network, a point cloud network or a multi-modal fusion network based on the multi-modal perception information; wherein, the pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain the pinhole 3D image features, the fish-eye network is used to extract the visual and spatial features of the fish-eye image collected by the fish-eye camera to obtain the fish-eye 3D image features, the point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter wave radar data to obtain the point cloud voxel features, and the multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fish-eye 3D image features and the point cloud voxel features.

[0291] In a possible implementation manner, the perception and prediction network further includes a temporal fusion network, which is configured to align multi-frame 3D spatial features in sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.

[0292] In a possible implementation manner, the perception and prediction network further includes a feature enhancement network, which is configured to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.

[0293] In a possible implementation manner, the perception and prediction network further includes a dynamic resolution prediction network based on the display information feedback. The dynamic resolution prediction network is configured to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area selected on the human-machine interaction interface and the grid granularity, obtain the position features through interpolation, and then predict the occupancy and semantic information based on the position features to achieve prediction at any resolution.

[0294] In a possible implementation manner, the multiple sensors include a pinhole camera and a fish-eye camera.

[0295] In a possible implementation manner, the multiple sensors further include a lidar and / or a millimeter-wave radar.

[0296] The device of this embodiment can be used to execute Figure 4 the technical solutions of the method embodiment shown. The implementation principle and technical effects are similar, and will not be elaborated here.

[0297] In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware-encoded processor, or can be executed and completed by the combination of the hardware and software modules in the encoded processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0298] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0299] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0300] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0301] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0302] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0303] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0304] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0305] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application and should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An assisted driving method, characterized in that, Including: Obtaining multi-modal perception information, which is collected by a variety of sensors disposed on a vehicle; Obtaining multiple groups of display information, where the display information includes a target area to be displayed and a grid granularity; Performing perception prediction based on the multi-modal perception information and the multiple groups of display information to obtain multiple occupancy grid information, where the occupancy grid information is used to describe the distribution of obstacles around the vehicle; Displaying multiple occupancy grid pictures on a human-machine interaction interface according to the multiple occupancy grid information.

2. The method according to claim 1, wherein The displaying multiple occupancy grid pictures according to the multiple occupancy grid information includes: Displaying a first occupancy grid picture corresponding to a first target area and a first grid granularity; Displaying a second occupancy grid picture corresponding to a second target area and a second grid granularity; The range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.

3. The method according to claim 2, characterized in that The first occupancy grid picture includes grids corresponding to n1 obstacles, and the second occupancy grid picture includes grids corresponding to n2 obstacles, where both n1 and n2 are positive integers.

4. The method according to claim 3, wherein The grids of a first obstacle in the first occupancy grid picture and the grids of a second obstacle in the second occupancy grid picture are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.

5. The method according to any one of claims 2 - 4, characterized in that The grids of a third obstacle in the second occupancy grid picture are identified with specific information. The third obstacle is the closest one to the vehicle among the n2 obstacles.

6. The method according to any one of claims 1-5, characterized in that, Further including: Displaying a panoramic occupancy grid picture on the human-machine interaction interface according to panoramic occupancy grid information, where the panoramic occupancy grid information is obtained based on historical occupancy grid information and a driving trajectory of the vehicle's driving environment.

7. The method according to any one of claims 1-6, characterized in that, The display information is pre-set information.

8. The method according to any one of claims 1 to 6, characterized in that The display information is dynamically generated based on operations of a driver on a touch screen.

9. The method according to claim 8, characterized in that, The operations include at least one of the following operations: An operation where the driver presses the touch screen with two fingers and slides the two fingers outward or inward together; or An operation where the driver presses the touch screen with two fingers and performs rotation, translation, or dragging.

10. The method according to any one of claims 1-9, characterized in that, The human-machine interaction interface further includes a control for selecting a driving mode, where the driving mode includes one or more of the following modes: a driving mode or a parking mode.

11. The method according to claim 10, wherein In the driving mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, where the auxiliary mode includes one or more of the following modes: an automatic mode or a manual mode.

12. The method according to claim 10, characterized in that, In the parking mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, where the auxiliary mode includes one or more of the following modes: an automatic mode, a manual mode, an autonomous valet parking AVP mode, or an automatic parking assistance APA mode.

13. The method according to any one of claims 1-12, characterized in that, The performing perception prediction based on the multi-modal perception information and the multiple groups of display information to obtain multiple occupancy grid information includes: Input the multi-modal perception information and the multi-group display information into a pre-trained perception prediction network to output the multiple occupancy grid information.

14. The method according to claim 13, wherein The perception prediction network includes one or more of the following networks: a pinhole network, a fisheye network, a point cloud network, or a multi-modal fusion network based on the multi-modal perception information; wherein, The pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain the pinhole 3D image features, the fisheye network is used to extract the visual and spatial features of the fisheye image collected by the fisheye camera to obtain the fisheye 3D image features, the point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter wave radar data to obtain the point cloud voxel features, and the multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fisheye 3D image features, and the point cloud voxel features.

15. The method according to claim 13 or 14, characterized in that, The perception prediction network further includes a temporal fusion network, which is used to align multi-frame 3D spatial features in sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.

16. The method according to any one of claims 13-15, characterized in that, The perception prediction network further includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.

17. The method according to any one of claims 13-16, characterized in that, The perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area selected on the human-computer interaction interface and the grid granularity, and obtain the position features through interpolation, and then predict the occupancy and semantic information according to the position features to achieve prediction at any resolution.

18. The method according to any one of claims 1-17, characterized in that, The multiple sensors include a pinhole camera and a fisheye camera.

19. The method according to claim 18, wherein The multiple sensors further include a lidar and / or a millimeter wave radar.

20. An assisted driving device, characterized in that, Including: An acquisition module, configured to acquire multi-modal perception information, where the multi-modal perception information is acquired by multiple sensors disposed on the vehicle; An acquisition module, configured to acquire multi-group display information, where the display information includes the target area to be displayed and the grid granularity; A prediction module, configured to perform perception prediction according to the multi-modal perception information and the multi-group display information to obtain multiple occupancy grid information, where the occupancy grid information is used to describe the distribution of obstacles around the vehicle; A display module, configured to display multiple occupancy grid pictures on the human-computer interaction interface according to the multiple occupancy grid information.

21. The device according to claim 20, characterized in that, The display module is specifically configured to display a first occupancy grid picture corresponding to a first target area and a first grid granularity; display a second occupancy grid picture corresponding to a second target area and a second grid granularity; the range of the first target area is larger than the range of the second target area, and the first grid granularity is larger than the second grid granularity.

22. The device according to claim 21, characterized in that, The first occupancy grid picture includes grids corresponding to n1 obstacles, and the second occupancy grid picture includes grids corresponding to n2 obstacles, where both n1 and n2 are positive integers.

23. The device according to claim 22, characterized in that, The grids of the first obstacle in the first occupied grid image and the grids of the second obstacle in the second occupied grid image are identified with the same information. The first obstacle and the second obstacle are the same obstacle. The n1 obstacles include the first obstacle, and the n2 obstacles all include the second obstacle.

24. The device according to any one of claims 21 to 23, characterized in that, The grids of the third obstacle in the second occupied grid image are identified with specific information. The third obstacle is the one closest to the vehicle among the n2 obstacles.

25. The device according to any one of claims 20-24, characterized in that, The display module is further configured to display an omnidirectional occupied grid image on the human-machine interaction interface according to the omnidirectional occupied grid information, where the omnidirectional occupied grid information is obtained based on the historical occupied grid information and the driving trajectory of the vehicle's driving environment.

26. The device according to any one of claims 20-25, characterized in that, The display information is pre-set information.

27. The device according to any one of claims 20-25, characterized in that, The display information is dynamically generated based on the driver's operations on the touch screen.

28. The device according to claim 27, wherein The operations include at least one of the following operations: The operation that the driver presses the touch screen with two fingers and slides the two fingers outward or inward together; or, The operation that the driver presses the touch screen with two fingers and rotates, translates or drags.

29. The device according to any one of claims 20 - 28, characterized in that, The human-machine interaction interface further includes a control for selecting a driving mode, and the driving mode includes one or more of the following modes: driving mode or parking mode.

30. The device according to claim 29, wherein In the driving mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode or manual mode.

31. The device according to claim 29, characterized in that, In the parking mode, the human-machine interaction interface further includes a control for selecting an auxiliary mode, and the auxiliary mode includes one or more of the following modes: automatic mode, manual mode, autonomous valet parking AVP mode or automatic parking assistance APA mode.

32. The device according to any one of claims 20-31, characterized in that, The prediction module is specifically configured to input the multi-modal perception information and the multiple groups of display information into a pre-trained perception prediction network to output the multiple occupied grid information.

33. The device according to claim 32, wherein, The perception prediction network includes one or more of the following networks: a pinhole network, a fish-eye network, a point cloud network or a multi-modal fusion network based on the multi-modal perception information; where The pinhole network is used to extract the visual and spatial features of the pinhole image collected by the pinhole camera to obtain the pinhole 3D image features. The fish-eye network is used to extract the visual and spatial features of the fish-eye image collected by the fish-eye camera to obtain the fish-eye 3D image features. The point cloud network is used to extract the 3D voxel features corresponding to the lidar and / or millimeter-wave radar data to obtain the point cloud voxel features. The multi-modal fusion network is used to perform weighted fusion on the pinhole 3D image features, the fish-eye 3D image features and the point cloud voxel features.

34. The device according to claim 32 or 33, characterized in that, The perception prediction network further includes a temporal fusion network, and the temporal fusion network is used to align the multi-frame 3D spatial features in sequence to the same coordinate system based on the motion parameters of the vehicle, and then perform multi-frame weighted fusion.

35. The device according to any one of claims 32 - 34, characterized in that The perception prediction network further includes a feature enhancement network, which is used to perform multi-scale feature fusion and multiple upsamplings on the 3D spatial feature map, and fuse the image features again to achieve auxiliary feature enhancement.

36. The device according to any one of claims 32 - 35, characterized in that, The perception prediction network further includes a dynamic resolution prediction network based on the feedback of the display information. The dynamic resolution prediction network is used to perform cropping sampling and weighted fusion on the 3D spatial feature map at multi-scale resolutions according to the target area selected on the human-computer interaction interface and the grid granularity, obtain position features through interpolation, and then predict occupancy and semantic information based on the position features to achieve prediction at any resolution.

37. The device according to any one of claims 20-36, characterized in that, The multiple sensors include a pinhole camera and a fish-eye camera.

38. The device according to claim 37, characterized in that, The multiple sensors further include a lidar and / or a millimeter-wave radar.

39. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-19.

40. A computer-readable storage medium, characterized in that, Including a computer program, when the computer program is executed on a computer, the computer executes the method according to any one of claims 1-19.

41. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on a computer, the computer executes the method according to any one of claims 1-19.