Perception information acquisition method and system, vehicle and product

By recognizing driving modes and using adaptive image perception and occupancy network models, the problem of inaccurate perception information acquisition in existing technologies has been solved, achieving accurate perception information acquisition in driving and parking modes and improving driving safety.

CN121799403APending Publication Date: 2026-04-07BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing perception information acquisition technologies are difficult to adapt to complex and ever-changing driving scenarios, resulting in missing or incorrect perception information acquired in different driving modes. In particular, in parking mode, blind spots are easily caused, increasing the risk of accidents.

Method used

By identifying the current driving mode, different image perception methods and occupancy network models are used to acquire perception information in driving and parking modes. In driving mode, a surround-view camera is used to acquire low-resolution information at a greater distance, while in parking mode, a fisheye camera is used to acquire high-resolution information at a closer distance. These images are then processed using an occupancy network model to obtain perception information adapted to different driving modes.

Benefits of technology

It enables the acquisition of more accurate and reliable perception information in different driving scenarios, avoids information loss or errors, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121799403A_ABST
    Figure CN121799403A_ABST
Patent Text Reader

Abstract

The invention provides a perception information acquisition method and system, a vehicle and a product, and relates to the technical field of vehicles, and the method comprises the steps: determining a current driving mode of the vehicle, and obtaining image data corresponding to the current driving mode; and on the basis of the image data corresponding to the current driving mode, perception information of the vehicle is obtained, and the perception information is used for representing the environment around the vehicle. According to the method, the image data corresponding to the current driving mode of the vehicle is obtained, the perception information of the vehicle is obtained from the image data, and more accurate and reliable perception information is extracted for different driving modes such as a driving mode and a parking mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a method, system, vehicle, and product for acquiring sensing information. Background Technology

[0002] Perception information acquisition refers to acquiring data about the vehicle's surrounding environment through sensors, extracting perception information such as the location, shape, type, speed, and trajectory of surrounding obstacles, and then using the acquired perception information to display the vehicle's driving environment to the driver.

[0003] However, during driving, the surrounding environment is complex and ever-changing, and existing perception information acquisition technologies struggle to adapt to these complex and changing driving scenarios, easily leading to incomplete or erroneous perception information. Therefore, a perception information acquisition method, system, vehicle, and product are needed to ensure accurate and reliable perception information acquisition in different driving scenarios. Summary of the Invention

[0004] In view of the above problems, embodiments of this application provide a method, system, vehicle, and product for acquiring perception information to ensure accurate and reliable perception information is acquired in different driving scenarios.

[0005] A first aspect of this application provides a method for acquiring sensory information, the method comprising: Determine the vehicle's current driving mode and acquire the image data corresponding to the current driving mode; Based on the image data corresponding to the current driving mode, the vehicle's perception information is obtained, and the perception information is used to characterize the environment around the vehicle.

[0006] A second aspect of this application also provides a sensing information acquisition system, applied to perform the sensing information acquisition method as described in the first aspect of this application, the system comprising: The image data acquisition module is used to determine the current driving mode of the vehicle and acquire the image data corresponding to the current driving mode; The perception information acquisition module is used to acquire the perception information of the vehicle based on the image data corresponding to the current driving mode, and the perception information is used to characterize the environment around the vehicle.

[0007] A third aspect of this application also provides a vehicle, the vehicle including a perception information acquisition system, the perception information acquisition system being used to perform the steps of the perception information acquisition method described in the first aspect of this application.

[0008] A fourth aspect of this application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the perception information acquisition method described in the first aspect of this application.

[0009] A fifth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the perception information acquisition method described in the first aspect of this application.

[0010] A sixth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the perception information acquisition method as described in the first aspect of this application.

[0011] This application provides a method for acquiring perception information, the method comprising: determining the current driving mode of a vehicle and acquiring image data corresponding to the current driving mode; acquiring perception information of the vehicle based on the image data corresponding to the current driving mode, the perception information being used to characterize the environment surrounding the vehicle.

[0012] The specific benefits are as follows: This application obtains image data corresponding to the current driving mode of the vehicle and extracts the vehicle's perception information from it. This enables the extraction of more accurate and reliable perception information for different driving modes, such as driving mode and parking mode, in order to adapt to complex and ever-changing driving environments and avoid the loss or error of the obtained perception information. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating the steps of a method for acquiring sensing information provided in an embodiment of this application; Figure 2 This is a schematic diagram of the location distribution of a vehicle-mounted camera provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an occupancy network model provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating intelligent interaction based on perceptual information, provided in an embodiment of this application. Figure 5This is a schematic diagram of a driving mode recognition process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a sensing information acquisition system provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0015] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0016] Perception information acquisition refers to acquiring data about the vehicle's surrounding environment through sensors and extracting perception information from it, such as the position, shape, type, speed, and trajectory of surrounding obstacles. This acquired perception information can then be used to present the vehicle's driving environment to the driver. However, during driving, the surrounding environment is complex and ever-changing, and existing perception information acquisition technologies struggle to adapt to these complex and changing driving scenarios, easily leading to incomplete or erroneous perception information.

[0017] In view of the above problems, embodiments of this application propose a method, system, vehicle, and product for acquiring perception information to ensure accurate and reliable acquisition of perception information in different driving scenarios. The following, in conjunction with the accompanying drawings, provides a detailed description of the perception information acquisition method, system, vehicle, and product provided by embodiments of this application through some examples and application scenarios.

[0018] The first aspect of this application proposes a method for acquiring sensory information. The method for acquiring sensory information in the first aspect will be described in sections 1.1-1.6 below.

[0019] 1.1 A brief overview of methods for acquiring sensory information: Reference Figure 1 , Figure 1 A flowchart illustrating the steps of a method for acquiring sensory information is shown, such as... Figure 1 As shown, the method includes: Step S101: Determine the current driving mode of the vehicle and obtain the image data corresponding to the current driving mode.

[0020] Step S102: Based on the image data corresponding to the current driving mode, obtain the perception information of the vehicle, which is used to characterize the environment around the vehicle.

[0021] In this embodiment, the current driving mode can be divided into driving mode or parking mode. Specifically, a vehicle's driving mode is generally divided into driving mode and parking mode. Driving mode refers to the mode in which the vehicle travels normally on the road at a higher speed, while parking mode refers to the mode in which the vehicle parks, turns, or drives slowly on some complex and difficult road conditions at a lower speed. Currently, relevant vehicle perception information acquisition technologies are mainly for acquiring perception information of the vehicle in driving mode and cannot be directly applied to parking mode. This can lead to errors or omissions in the perception information acquired in parking mode. For example, the acquired perception information may show an excessively large blind spot, and if the driver directly parks based on the perception information, it may easily cause accidents such as collisions.

[0022] In this embodiment, when executing step S102, the image data corresponding to the current driving mode can be processed using the image perception method corresponding to the current driving mode to obtain the vehicle's perception information; wherein, the driving mode corresponds to a first image perception method with a first perception resolution, and the parking mode corresponds to a second image perception method with a second perception resolution. The second perception resolution is greater than the first perception resolution.

[0023] In this embodiment, after identifying the vehicle's current driving mode (driving mode or parking mode), image data can be acquired according to the corresponding driving mode. For example, the image acquisition range of the first set of cameras corresponding to the driving mode is larger than the image acquisition range of the second set of cameras corresponding to the parking mode. A larger image acquisition range allows information from farther locations to be obtained, but the corresponding accuracy or resolution is lower. In driving mode, the vehicle speed is higher, and to ensure vehicle safety, it is necessary to acquire perception information to understand objects at greater distances (e.g., vehicles or obstacles 30 meters away). Therefore, a longer perception range is required (i.e., a longer corresponding image acquisition range), and correspondingly, the requirements for perception resolution and accuracy are relatively lower (generally requiring a perception accuracy of 40-50cm). In parking mode, the vehicle moves slower and other objects in the scene are closer to the vehicle. To avoid accidents such as collisions, it is necessary to accurately perceive the information of various obstacles around the vehicle by acquiring perception information, so as to minimize the blind spot of the vehicle. Therefore, the perception range needs to be closer (i.e., the corresponding image acquisition range is closer), and the requirements for perception resolution and accuracy are relatively high (generally requiring a perception accuracy of 5-10cm).

[0024] Therefore, the two different driving modes have different requirements for perception information. In driving mode, image data from a greater distance is acquired and processed at a lower perception resolution (first perception resolution) to obtain perception information. In parking mode, image data from a closer distance is acquired and processed at a higher perception resolution (second perception resolution) to obtain perception information. This system addresses the different requirements of driving and parking scenarios regarding the range, resolution, and accuracy of perception information by processing image data using different perception methods (first image perception method or second first image perception method) to obtain more accurate and reliable perception information (e.g., the position, shape, type, speed, and trajectory of obstacles around the vehicle), outputting perception results adapted to the corresponding driving mode. Specifically, the image perception method involves selecting the corresponding perception hyperparameters (including perception range, perception accuracy, and perception resolution) according to the current driving mode, and then inputting the image data into the occupancy network model, allowing the occupancy network model to extract the perception information corresponding to the perception hyperparameters from the image data.

[0025] Different driving modes have different requirements for perception information and require different amounts of image data. Driving modes need to obtain perception information over a longer range from image data, while parking modes need to obtain perception information with fewer blind spots and higher resolution. This application obtains the vehicle's perception information by acquiring image data corresponding to the vehicle's current driving mode, thus achieving more accurate and reliable perception information extraction for different driving modes, such as driving and parking modes. This adapts to complex and changing driving environments and avoids missing or erroneous perception information.

[0026] 1.2 Acquire image data according to the method corresponding to the current driving mode: Different driving modes have different image data requirements. This embodiment adaptively switches the driving / parking occupancy perception mode based on the determined current driving mode (driving mode or parking mode), acquiring image data according to the method corresponding to the current driving mode to meet its perception information needs. Specifically, in driving mode, multiple-view surround-view cameras can be used to capture images of the vehicle's surroundings, obtaining image data corresponding to the driving mode. Surround-view cameras have a longer perception range, which is beneficial for acquiring information from more distant locations. In parking mode, multiple-view fisheye cameras can be used in conjunction with surround-view cameras to capture images of the vehicle's surroundings, obtaining image data corresponding to the parking mode. The arrangement of surround-view cameras can easily result in blind spots exceeding 2 meters around the vehicle. Fisheye cameras have a wider shooting angle, which can reduce the blind spots created by surround-view cameras, allowing for more accurate information about the vehicle's close-range location.

[0027] In one possible implementation, acquiring the image data corresponding to the current driving mode includes: When the current driving mode is driving mode, acquire the image data corresponding to the driving mode; When the current driving mode is parking mode, acquire the image data corresponding to the parking mode.

[0028] In one possible implementation, acquiring image data corresponding to the parking mode when the current driving mode is parking mode includes: When the current driving mode of the vehicle is parking mode, the first set of cameras of the vehicle is used to acquire image data corresponding to the parking mode. The first set of cameras includes: a first type of camera with front view, rear view and side view of the vehicle, and a second type of camera with front view and rear view of the vehicle.

[0029] The first group of cameras includes a type 1 camera with four perspectives: front view, left view, rear view, and right view, and a type 2 camera with front view and rear view. The type 1 camera is a fisheye camera, and the type 2 camera is a panoramic camera.

[0030] The step of acquiring image data corresponding to the current driving mode, which is driving mode, includes: When the current driving mode of the vehicle is driving mode, the second set of cameras of the vehicle is used to acquire image data corresponding to the driving mode. The second set of cameras includes: a second type of camera with a front view, a rear view and a side view of the vehicle; wherein, the field of view of the first type of camera is greater than the field of view of the second type of camera, and the maximum working distance of the first type of camera is less than the maximum working distance of the second type of camera.

[0031] The second group of cameras includes six perspectives: front-view, left-front-view, left-rear-view, rear-view, right-front-view, and right-rear-view. These second-type cameras can be panoramic cameras. The field of view of the first type of camera is larger than that of the second type of camera, and the maximum working distance of the first type of camera is smaller than that of the second type of camera. The field of view refers to the angle of view of the camera; the larger the field of view, the smaller the blind spot. The maximum working distance is the distance between the camera and the furthest point within the camera's information acquisition range.

[0032] In this embodiment, when image data is captured using a first type of camera or a second type of camera, the captured camera parameters, including camera intrinsic parameters and camera extrinsic parameters, can be obtained. The camera parameter information (first camera parameter information or second camera parameter information) includes at least: original camera intrinsic parameters, distortion coefficients, and camera pose information relative to the vehicle. This camera parameter information includes camera intrinsic parameters and camera extrinsic parameters. The camera intrinsic parameters include the original camera intrinsic parameters and distortion coefficients, and the camera extrinsic parameters include the camera pose information relative to the vehicle (including rotation and translation matrices).

[0033] Reference Figure 2 , Figure 2 A schematic diagram showing the location distribution of vehicle-mounted cameras is provided, such as... Figure 2 As shown, four side-view (left front, left rear, right front, right rear) surround-view cameras are installed on the vehicle fenders. The fields of view of the surround-view cameras from different angles are interspersed. Given a camera configuration, it can be ensured that the camera coverage blind spots and dead angles are small.

[0034] Here, A represents the surround-view camera. Surround-view cameras have a longer field of view, enabling them to capture objects at greater distances, but their narrower angle of view can easily lead to blind spots exceeding 2 meters around the vehicle. In driving mode, images are captured using a second set of cameras, specifically six surround-view cameras (A1 for front view, A2 for left front, A3 for left rear, A4 for rear view, A5 for right front, and A6 for right rear). This yields image data corresponding to the driving mode, which is pure visual data obtained through camera capture, along with camera parameter information from the second set of cameras (second camera parameter information). This multi-view surround-view imagery provides rich semantic information about the vehicle's surroundings.

[0035] Most technologies for acquiring network occupancy perception information use one or more of the following: surround view cameras, millimeter-wave radar, or lidar multimodal sensors to obtain perception information about the vehicle's surroundings in the driving environment. However, these sensors have large blind spots near the vehicle and it is difficult to predict the occupancy of voxels in the vehicle's close space. Therefore, they are not applicable in parking scenarios or when passing through narrow spaces.

[0036] To address the aforementioned problems, embodiments of this application propose using a second set of cameras for image capture in parking mode, such as... Figure 2As shown, B represents a fisheye camera, specifically a camera utilizing four fisheye views (front-view B1, rear-view B2, left-view B3, and right-view B4) plus two surround-view cameras (front-view A1 and rear-view A4) to capture images, resulting in a total of six perspectives corresponding to the parking mode. This represents pure visual data obtained through camera capture, along with the camera parameter information of the first set of cameras (first camera parameter information). Therefore, this embodiment uses four fisheye cameras and front-view and rear-view high-definition surround-view cameras to capture images in parking scenarios, and six surround-view cameras to capture images in driving scenarios. This ensures that the perception range and blind spots meet the usage requirements for different driving modes, making the solution applicable to parking scenarios. Furthermore, compared to related technologies that use surround-view cameras, millimeter-wave radar, or lidar multimodal sensors, the camera sensor used in this embodiment is a commonly equipped sensor device in vehicles, thus ensuring that the proposed technical solution has wider applicability and lower cost.

[0037] 1.3 Using an occupancy network to obtain perceptual information from image data: In one possible implementation, obtaining the vehicle's perception information based on image data corresponding to the current driving mode includes: The image data corresponding to the current driving mode is input into the occupancy network model to obtain the vehicle's perception information.

[0038] Occupancy network model, also known as occupancy prediction network, refers to the occupancy network model. In related technologies, occupancy networks are used to acquire perception information in driving modes, without considering that the perception information required in parking modes differs from that required in driving modes. The occupancy network model proposed in this application is obtained by training the model using pre-collected sample image data corresponding to parking modes (sample image data captured by a first set of cameras, i.e., sample image data captured by a fisheye camera and a surround-view camera) and the corresponding vehicle perception information as training samples. For example, the occupancy network model is trained using the first sample image data corresponding to parking modes and the second sample image data corresponding to driving modes, enabling it to acquire perception information from image data corresponding to different driving modes.

[0039] In one possible implementation, inputting the image data corresponding to the current driving mode into the occupancy network model includes: When the current driving mode is driving mode, the image data corresponding to the driving mode is input into the occupancy network model; When the current driving mode is parking mode, the image data corresponding to the parking mode is input into the occupancy network model.

[0040] Specifically, the occupancy network model can be viewed as a combination of two sub-models: the image data corresponding to the parking mode is input into the parking occupancy network sub-model of the occupancy network model to obtain the vehicle's parking perception information, which is used to characterize the vehicle's parking environment; the image data corresponding to the driving mode is input into the driving occupancy network sub-model of the occupancy network model to obtain the vehicle's driving perception information, which is used to characterize the vehicle's driving environment.

[0041] The following describes in detail the process of the occupancy network model processing the input image data. In this embodiment, the occupancy network model (based on selected perception hyperparameters) is used to obtain the vehicle's perception information from the image data corresponding to the current driving mode (six-view vehicle environment images and camera parameter information).

[0042] In one possible implementation, the image data corresponding to the current driving mode is input into an occupancy network model to obtain the vehicle's perception information, including: Step S201: Based on the distortion coefficient of the camera used to collect the image data corresponding to the current driving mode, perform distortion removal processing on the image data corresponding to the current driving mode to obtain distortion-free image data.

[0043] Specifically, because different cameras (fisheye cameras or surround-view cameras) are used to capture images in different driving modes, the data format and resolution of the image data obtained in the two driving modes differ. (Refer to...) Figure 3 , Figure 3 A schematic diagram of the structure of an occupancy network model is shown, such as... Figure 3 As shown, firstly, the image data corresponding to the current driving mode is input into the distortion correction module of the network model to perform distortion correction on the image data. Specifically, as shown in Formula 1: (Formula 1); in, This represents the image data from the x-th viewpoint (the raw image captured by the camera). This represents the distortion correction coefficients used to correct distortion in the image data from the x-th viewpoint. This represents the distortion-free image data after distortion correction from the nth viewpoint. This represents the camera's extrinsic parameters (camera pose information relative to the vehicle). This represents the camera's intrinsic parameters (original camera intrinsic parameters, distortion coefficients).

[0044] Specifically, for image data captured by a surround-view camera, a surround-view camera distortion correction algorithm is used for distortion correction; for image data captured by a fisheye camera, a fisheye camera distortion correction algorithm is used for distortion correction, resulting in distortion-free camera images (distortion-free image data) and distortion-corrected camera parameter information. Furthermore, the resolution of the distortion-free image data is adjusted to meet the input data requirements of the network model. After adjustment, the camera parameter information (intrinsic parameters) is recalculated to obtain adjusted camera parameter information.

[0045] Step S202: Extract feature maps from the distortion-free image data.

[0046] Specifically, based on the backbone network, two-dimensional image features are extracted from the distortion-free camera images (distortion-free image data) at each viewpoint to obtain a feature map. As shown in Formula 2: (Formula 2); in, Let represent the distortion-free camera image (distortion-free image data) from the nth viewpoint, 𝐵𝑎𝑐𝑘𝑏𝑜𝑛𝑒() represent the backbone network, and 𝐼𝑚𝑔𝐹𝑒𝑎𝑡𝑢𝑟𝑒(𝑗) represent the feature map extracted by the backbone network from the distortion-free image data of the nth viewpoint. In this embodiment, the backbone network architecture can be a ResNet101 network structure.

[0047] Step S203: Based on the intrinsic and extrinsic parameters of the camera used to collect image data corresponding to the current driving mode, the feature map is converted into a bird's-eye view BEV feature map.

[0048] Specifically, using the camera pose information from the intrinsic and extrinsic parameters of the camera used to collect image data corresponding to the current driving mode, and based on the mapping relationship between 2D images and 3D space, multi-view 2D image features (i.e., feature maps) are aggregated into 3D space. Then, a pooling module converts the 3D spatial features into BEV features in 3D space, reducing the number of network parameters and improving the model's inference speed. As shown in Formula 3: (Formula 3); In the formula, Let represent the feature map extracted by the backbone network from the distortion-free image data of the ith viewpoint, let represent the camera extrinsic and intrinsic parameters of the j-th viewpoint, let LSS() represent the BEV feature extraction algorithm, and let BEVFeature represent the extracted BEV features.

[0049] Step S204: Target detection is performed based on the BEV feature map to obtain target detection information in the three-dimensional space where the vehicle is located; and occupancy prediction is performed based on the BEV feature map to obtain voxel occupancy status and occupancy semantic information in the three-dimensional space where the vehicle is located. The vehicle's perception information includes: target detection information in the three-dimensional space where the vehicle is located, voxel occupancy status in the three-dimensional space where the vehicle is located, and occupancy semantic information.

[0050] Specifically, such as Figure 3 As shown, based on the occupancy-aware feature learning network, voxel occupancy information (i.e., voxel occupancy status and occupancy semantic information) in 3D space is inferred from the BEV features in 3D space. As shown in Formula 4: (Formula 4); Here, BEVFeature represents the extracted BEV features. This represents the occupancy-aware feature learning network used to extract occupancy-aware information, and OCC represents the voxel occupancy information obtained through inference.

[0051] Furthermore, based on object detection feature learning networks, object detection information (i.e., object detection information), including bounding boxes and orientations, can be inferred from BEV features in 3D space. This can be done by inferring the object detection information in 3D space from the BEV features. (See Formula 5.) (Formula 5); Here, BEVFeature represents the extracted BEV features. represents the object detection feature learning network, and 𝐷𝑒𝑡 represents the object detection information in the three-dimensional space obtained through inference.

[0052] Finally, the occupancy network model, after passing through a post-processing module, outputs multi-paradigm fused 3D spatial visual perception information (i.e., vehicle perception information), including voxel occupancy in 3D space, occupancy mesh semantic information, and object detection information. This information is then transmitted to downstream modules for the vehicle to perform intelligent interaction tasks. This embodiment utilizes an occupancy network model to acquire perception information from multi-angle image data (image data corresponding to the current driving mode), determining the position, shape, and speed of objects in the vehicle's 3D space. This enables the acquisition of perception information from image data corresponding to different driving modes, improving the accuracy and reliability of the acquired perception information.

[0053] 1.4 Obtain the perception hyperparameter set corresponding to the driving mode, and obtain perception information based on the perception hyperparameter set: Different driving modes have different requirements for perception information. In this embodiment, the occupancy perception model uses the perception hyperparameters corresponding to the current driving mode (e.g., the perception resolution corresponding to parking mode is greater than that corresponding to driving mode) to process image data and obtain perception information. This achieves the different requirements for perception information in driving and parking modes (e.g., higher resolution information on the close position of the vehicle is required in parking mode). The image data is processed using different perception methods (e.g., different sub-models in the occupancy network model, or different perception hyperparameters) to obtain more accurate and reliable perception information.

[0054] In one possible implementation, the step of inputting the image data corresponding to the current driving mode into the occupancy network model to obtain the vehicle's perception information includes: Determine the perception hyperparameter set corresponding to the current driving mode; the perception hyperparameter set includes at least one of the following: perception range, number of voxel grids, and side length of voxel grids; The occupancy network model, based on the perception hyperparameter set, extracts the vehicle's perception information from the image data corresponding to the current driving mode.

[0055] In this embodiment, different driving modes have different requirements for perception information. In driving mode, a longer perception range, lower perception accuracy, and perception resolution (i.e., a smaller voxel grid side length) are used to extract perception information from image data. In parking mode, a closer perception range, higher perception accuracy, and perception resolution (i.e., a larger voxel grid side length) are used to extract perception information from image data. This allows the occupancy network model to obtain more accurate and reliable perception information for the different requirements of driving and parking scenarios in terms of the range, resolution, and accuracy of perception information (i.e., the corresponding perception hyperparameter set), and output perception results adapted to the corresponding driving mode.

[0056] In one possible implementation, determining the perception hyperparameter set corresponding to the current driving mode includes: When the current driving mode is parking mode, determine one of the parking perception hyperparameter groups corresponding to the current parking scenario from multiple parking perception hyperparameter groups; When the current driving mode is driving mode, determine one of the driving perception hyperparameter groups corresponding to the current driving scenario from multiple driving perception hyperparameter groups; The sensing range of the parking perception hyperparameter group is smaller than that of the driving perception hyperparameter group; the side length of the voxel grid of the parking perception hyperparameter group is smaller than that of the voxel grid of the driving perception hyperparameter group.

[0057] In this embodiment, the driving mode requires a longer perception range and relatively lower requirements for perception resolution and accuracy, while the parking mode requires a shorter perception range and relatively higher requirements for perception resolution and accuracy. This embodiment determines the appropriate set of perception hyperparameters for the occupancy network model based on different scenarios.

[0058] In one possible implementation, when the current driving mode is a driving mode, determining one of the driving perception hyperparameter sets corresponding to the current driving scenario from multiple driving perception hyperparameter sets includes: Based on the vehicle's driving scenario information, a driving perception hyperparameter group corresponding to the current driving scenario is determined from the first perception hyperparameter group, the second perception hyperparameter group, and the third perception hyperparameter group; wherein, the perception range of the first perception hyperparameter is farther than that of the second perception hyperparameter, and the perception range of the second perception hyperparameter is farther than that of the third perception hyperparameter.

[0059] Specifically, driving scenario information can include vehicle speed, road conditions, and perception information obtained from the previous perception acquisition. In driving mode, a suitable combination of perception hyperparameters can be selected from three preset combinations based on the driving scenario information. These are the first, second, and third perception hyperparameter groups. The first perception hyperparameter group is for long-range perception, including a perception range of 75m front and rear, 25m left and right, and 5m up and 3m down, with a voxel grid of 300*100*16 and each voxel grid having a side length of 0.5m. The second perception hyperparameter group is for mid-range perception, including a perception range of 50m front and rear, 50m left and right, and 5m up and 3m down, with a voxel grid of 200*200*16 and each voxel grid having a side length of 0.5m. The third perception hyperparameter group belongs to the mid-range perception hyperparameter group, which includes: a perception range of 40m in front and behind, 40m in the left and right, 4.4m above and 2m below, with the rear axle center of the vehicle as the origin, and a voxel grid of 200*200*16, with each voxel grid having a side length of 0.4m.

[0060] When the current driving mode is parking mode, determine one of the parking perception hyperparameter sets corresponding to the current parking scenario from multiple parking perception hyperparameter sets, including: Based on the vehicle's parking scenario information, a parking perception hyperparameter group corresponding to the current parking scenario is determined from the fourth and fifth perception hyperparameter groups. The third perception hyperparameter group has a longer perception range than the fourth perception hyperparameter group, and the fourth perception hyperparameter group has a longer perception range than the fifth perception hyperparameter group. Furthermore, the parking perception hyperparameter group has a higher perception accuracy than the driving perception hyperparameter group.

[0061] Specifically, parking scenario information can include vehicle speed, gear, road conditions, and perception information obtained from the previous perception acquisition. In parking mode, based on the parking scenario information, a suitable perception hyperparameter can be selected from two preset combinations: the fourth perception hyperparameter group and the fifth perception hyperparameter group. The fourth perception hyperparameter group is a mid-range hyperparameter group, including a perception range of 12.5m in front and behind, 12.5m to the left and right, and 3m above and 1m below, with a voxel grid of 250*250*40 and each voxel grid having a side length of 0.1m. The fifth perception hyperparameter group is a short-range hyperparameter group, including a perception range of 0.8m in front and behind, 0.8m to the left and right, and 2m above and 0.5m below, with a voxel grid of 320*320*50 and each voxel grid having a side length of 0.05m.

[0062] This application addresses the different requirements of driving and parking scenarios regarding the range, resolution, and accuracy of perception information. It can acquire corresponding perception hyperparameter sets based on the characteristics of the driving mode, outputting matching perception range and resolution for different scenarios. Therefore, it obtains more accurate and reliable perception information and outputs perception results adapted to the current driving mode, catering to the different requirements of driving and parking scenarios for the range, resolution, and accuracy of perception information (i.e., the corresponding perception hyperparameter sets).

[0063] 1.5 Utilizing vehicle perception information for intelligent interaction: In this embodiment, after acquiring the vehicle's perception information (multi-paradigm fusion of three-dimensional spatial visual perception information), the perception information is processed and transmitted to the in-vehicle interaction model used to perform interactive tasks. The acquired perception information is widely used to realize intelligent interaction of vehicle-side network occupancy perception, so as to improve user experience.

[0064] In one possible implementation, the method further includes: Based on the vehicle's perception information, a perception interaction is performed.

[0065] Reference Figure 4 , Figure 4 A schematic diagram of intelligent interaction based on perceptual information is shown, such as... Figure 4As shown, based on the acquired perceptual information, three aspects of interaction (voice interaction, visual interaction, and intelligent interaction) can be performed. Sections 1.5.1, 1.5.2, and 1.5.3 below provide specific explanations for each aspect of interaction.

[0066] 1.5.1 Voice Interaction: In one possible implementation, the perception interaction based on the vehicle's perception information includes: In response to a user's voice query request regarding the vehicle's surrounding environment, a voice reply is generated based on the vehicle's perception information.

[0067] In this embodiment, the user can initiate a voice query request for perceived information about the vehicle's surrounding environment through dialogue (e.g., how many vehicles are in front? / Approximately how far is the rear of the vehicle from the wall when reversing?). The voice interaction module can obtain the corresponding information from the perceived information, output the query results, and convert the text into voice response through the text-to-speech module.

[0068] In one possible implementation, a voice response is generated based on the vehicle's perception information, including: Using a large vision-language model, image language features are extracted from the vehicle's perceived information; Based on the image language features, the geometric and semantic information of each target is encoded to obtain a language-enhanced map; The voice information in the voice query request is converted into text information and input into the visual-language big model; Based on the language enhancement map, the visual-language big model generates a voice response to the text information.

[0069] In this embodiment, the Visual-Language Large Model (LVLM) can compute image-language features (i.e., visual-language features) from perceived information. These features are used to generate object descriptions of relevant objects around the vehicle. Each target (referring to various target objects in the vehicle's surrounding environment obtained from the target detection information of the perceived information) is aligned with its image-language features to enhance the generated occupancy network information, thereby constructing a language-enhanced map. In the language-enhanced map, the geometry (location, area, centroid) and semantics (target and image description) of each target are encoded. The language-enhanced map can be directly used as the context for LVLM to answer user target-level and scene-level queries. That is, the vehicle-side user can input questions via voice, the system converts the voice information into text information through a speech-to-text module, and inputs it into the LVLM model for querying. The LVLM model extracts information from the perceived information to provide an answer, and the natural language explanation returned by the LVLM model is converted into speech (i.e., a voice response) for playback.

[0070] 1.6.2 Intelligent Interaction: In one possible implementation, the perception interaction based on the vehicle's perception information includes: Based on the vehicle's sensor information, control commands are sent to the vehicle's onboard electrical systems.

[0071] In this embodiment, the vehicle's electrical system acquires information about the driving scenario through sensing information. Specifically, by extracting key information for the application of electrical functions (such as application information), the system intelligently adjusts the electrical functions (by sending control commands to the vehicle's electrical system) to achieve intelligent interaction. Taking a vehicle-mounted pixel headlight as an example, the acquired sensing information enables intelligent lighting for different driving modes. For instance, the system can determine the position of nearby objects based on the sensing information, providing high-brightness illumination for these objects to alert the driver to their presence. Furthermore, it can adjust the direct beam range based on the grid positions occupied by pedestrians and oncoming vehicles to avoid glare for pedestrians and drivers of oncoming vehicles.

[0072] 1.6.3 Visual Interaction: In one possible implementation, the vehicle's perception information includes: voxel occupancy information and target detection information; the perception interaction based on the vehicle's perception information includes: Replace the voxel occupancy information corresponding to the position in the object detection box of the target detection information with the three-dimensional model of the target object represented by the voxel occupancy information; The vehicle's driving environment is displayed on the screen by combining the vehicle's perception information and the target object's three-dimensional model.

[0073] In this embodiment, since the perceptual information obtained through occupancy networks is generally displayed as object information in three-dimensional space in the form of voxel occupancy semantics, the information is too discrete and loses the overall information of the object. Therefore, this embodiment proposes to combine voxel occupancy information and target detection information to perform visualization of fused perceptual information. Figure 4 The visualization replacement of the 3D object model shown replaces the voxel occupancy information of the corresponding positions in the object detection bounding boxes (object detection information) with the 3D model of the target object (fine-grained mesh model). This ensures a consistent visual style while providing a more intuitive display of occupancy network information. Users can directly observe the visualized driving space fused from the occupancy network (voxel occupancy information) and object detection (object detection information). Centered on the vehicle, it can display global and local spatial information according to user needs, enabling a better overall perception of the vehicle's driving environment and assisting driving.

[0074] 1.6 Methods for identifying driving modes: Regarding step S101 above, this embodiment proposes an adaptive recognition method for identifying driving modes as either driving or parking modes. If the required perception information in the current driving scenario is close-range scene and object information, and the distance resolution accuracy requirement is high, then the driving mode is classified as parking mode. If the required perception information in the current driving scenario is medium- to long-range scene and object information, and the distance resolution accuracy requirement is lower, then the driving mode is classified as driving mode. Thus, by recognizing the driving mode, the driving / parking occupancy perception mode can be flexibly switched, and the perception result adapted to the corresponding driving mode can be output.

[0075] In one possible implementation, determining the vehicle's current driving mode includes: Step S1011: Detect the driver's manual selection result and determine the current driving mode of the vehicle; or Step S1012: Determine the current driving mode of the vehicle based on the vehicle driving information.

[0076] Specifically, refer to Figure 5 , Figure 5 A schematic diagram of a driving mode recognition process is shown, such as Figure 5 As shown, driving mode recognition is divided into automatic recognition or manual selection (where the driver manually selects based on their perception needs). In manual selection mode, the driver can select the current driving mode as driving mode or parking mode (e.g., driving mode) by triggering relevant buttons on the display screen. Figure 5 The manual parking mode and manual driving mode are shown. If the driver does not operate the system, the system defaults to automatic recognition mode to identify the driving mode. During vehicle operation, driving mode recognition can be performed at fixed intervals (i.e., step S1012), or when a manual selection by the driver is detected, driving mode recognition can be performed (i.e., step S1011). The automatic recognition mode can be based on a rule-based algorithm using a decision tree to identify the current driving mode (e.g., ...). Figure 5 (The rule-based driving mode determination shown).

[0077] In one possible implementation, the vehicle driving information includes: vehicle gear information and vehicle speed information; step S1012, determining the current driving mode of the vehicle based on the vehicle driving information, includes: If the vehicle is currently in reverse or neutral, the current driving mode will be set to parking mode. Specifically, for example... Figure 5As shown, the automatic recognition mode first determines the vehicle's gear. If the current gear is reverse or neutral, it means that the vehicle is not in normal driving mode and the current driving mode can be directly classified as parking mode. If the current gear is drive, further judgment is made based on the vehicle's speed.

[0078] If the vehicle is currently in drive: Based on the vehicle speed information, if the vehicle speed is below a first speed, the current driving mode is determined to be parking mode. Specifically, the first speed can be 10 km / h. If the vehicle's speed at the current moment is below 10 km / h, the vehicle is in a low-speed state, and the judgment of information about nearby objects is more important, so the current driving mode is classified as parking mode.

[0079] Based on the vehicle speed information, if the vehicle speed is determined to be above a second speed and the duration exceeds a first duration, the current driving mode is determined to be a driving mode. Specifically, the second speed can be 20 km / h, and the first duration can be 5 seconds. If the vehicle speed is above 20 km / h and lasts for more than 5 seconds, the current driving scenario is classified as a driving mode.

[0080] Otherwise, the current driving mode will be set to parking mode. Specifically, if the vehicle speed does not meet either of the above two conditions, the current driving mode will be directly set to parking mode.

[0081] Optionally, the method further includes: If it is determined that the current driving mode of the vehicle is different from the driving mode at the previous moment, the duration since the last driving mode switch is determined.

[0082] If the duration is less than the second duration, wait until the second duration is reached, then switch driving modes.

[0083] like Figure 5As shown in the "Anti-Jump Module," this embodiment can also set an anti-jump mechanism to avoid frequent changes in the vehicle's driving mode. Specifically, if the vehicle's driving mode was driving mode at the previous moment, and the current driving mode is identified as parking mode (i.e., the driving mode has changed, and it is determined that the vehicle's driving mode at the current moment is different from the previous moment's driving mode), then it is necessary to first determine the duration after the last driving mode switch, that is, the duration after the last switch from parking mode to driving mode. If the duration does not exceed the second duration (the second duration can be 5 seconds), it means that the last driving mode change was recent. In this case, the driving mode is not changed temporarily (from driving mode to parking mode at the previous moment). It can wait until the second duration (i.e., driving mode is maintained for 5 seconds) before switching the driving mode (changing the driving mode from driving mode to parking mode at the previous moment), thus avoiding frequent changes in the vehicle's driving mode and ensuring the accuracy and reliability of the recognition results.

[0084] A second aspect of this application also provides a sensing information acquisition system, applied to perform the sensing information acquisition method described in the first aspect, with reference to... Figure 6 , Figure 6 A schematic diagram of the structure of a sensory information acquisition system is shown, such as... Figure 6 As shown, the system includes: The image data acquisition module is used to determine the current driving mode of the vehicle and acquire the image data corresponding to the current driving mode; The perception information acquisition module is used to acquire the perception information of the vehicle based on the image data corresponding to the current driving mode, and the perception information is used to characterize the environment around the vehicle.

[0085] In one possible implementation, the sensing information acquisition module includes: An occupancy network model is used to obtain the vehicle's perception information based on the image data corresponding to the current driving mode.

[0086] In one possible implementation, the sensing information acquisition module includes: The first input submodule is used to input the image data corresponding to the driving mode into the occupancy network model when the current driving mode is driving mode. The second input submodule is used to input the image data corresponding to the parking mode into the occupancy network model when the current driving mode is parking mode.

[0087] In one possible implementation, the system further includes: A perception hyperparameter group determination module is used to determine the perception hyperparameter group corresponding to the current driving mode; the perception hyperparameter group includes at least one of the following: perception range, number of voxel grids, and side length of voxel grids; The occupancy network model extracts the vehicle's perception information from the image data corresponding to the current driving mode based on the perception hyperparameter set.

[0088] In one possible implementation, the hyperparameter set determination module includes: The parking perception hyperparameter group determination submodule is used to determine one of the parking perception hyperparameter groups corresponding to the current parking scenario from multiple parking perception hyperparameter groups when the current driving mode is parking mode. The vehicle perception hyperparameter group determination submodule is used to determine one of the vehicle perception hyperparameter groups corresponding to the current driving scenario from multiple vehicle perception hyperparameter groups when the current driving mode is driving mode. The sensing range of the parking perception hyperparameter group is smaller than that of the driving perception hyperparameter group; the side length of the voxel grid of the parking perception hyperparameter group is smaller than that of the voxel grid of the driving perception hyperparameter group.

[0089] In one possible implementation, the image data acquisition module includes: The driving image data acquisition submodule is used to acquire image data corresponding to the driving mode when the current driving mode is driving mode; The parking image data acquisition submodule is used to acquire image data corresponding to the parking mode when the current driving mode is parking mode.

[0090] In one possible implementation, the system further includes: The first type of camera provides the vehicle's front, rear, and side views; the first type of camera is used to acquire image data corresponding to the parking mode. The vehicle has a second type of camera for front, rear and side views; the second type of camera is used to acquire image data corresponding to parking mode and / or driving mode; wherein, the field of view of the first type of camera is greater than the field of view of the second type of camera, and the maximum working distance of the first type of camera is less than the maximum working distance of the second type of camera.

[0091] In one possible implementation, the first type of camera is a fisheye camera, and the second type of camera is a panoramic camera.

[0092] In one possible implementation, the first type of camera for the side view is located at the fender position of the vehicle.

[0093] In one possible implementation, the occupancy network model includes: The distortion correction module is used to perform distortion correction processing on the image data corresponding to the current driving mode based on the distortion coefficient of the camera used to collect the image data corresponding to the current driving mode, so as to obtain distortion-free image data. A backbone network is used to extract feature maps from the distortion-free image data; The BEV feature conversion module is used to convert the feature map into a BEV feature map based on the intrinsic and extrinsic parameters of the camera used to collect image data corresponding to the current driving mode. An object detection feature learning network is used to perform object detection based on the BEV feature map to obtain object detection information in the three-dimensional space where the vehicle is located. An occupancy-aware feature learning network is used to predict occupancy based on the BEV feature map, thereby obtaining the voxel occupancy status and occupancy semantic information of the three-dimensional space in which the vehicle is located.

[0094] In one possible implementation, the system further includes: The perception and interaction module is used to perform perception and interaction based on the perception information of the vehicle.

[0095] In one possible implementation, the perception interaction module includes: The voice interaction submodule is used to respond to user-initiated voice query requests regarding the vehicle's surrounding environment and generate a voice response based on the vehicle's perception information.

[0096] In one possible implementation, the voice interaction module includes: A large vision-language model is used to extract image language features from the vehicle's perception information; based on the image language features, the geometric and semantic information of each target is encoded to obtain a language-enhanced map; The voice information in the voice query request is converted into text information; based on the language enhancement map, a voice response is generated for the text information.

[0097] In one possible implementation, the perception interaction module includes: The intelligent interaction submodule is used to send control commands to the vehicle's onboard electrical systems based on the vehicle's perception information.

[0098] In one possible implementation, the vehicle's perception information includes voxel occupancy information and target detection information; the perception interaction module includes a visual interaction submodule, used for... Replace the voxel occupancy information corresponding to the position in the object detection box of the target detection perception information with the three-dimensional model of the target object represented by the voxel occupancy information; The vehicle's driving environment is displayed on the screen by combining the vehicle's perception information and the target object's three-dimensional model.

[0099] In one possible implementation, the image data acquisition module includes: A manual recognition submodule is used to detect the driver's manual selection and determine the vehicle's current driving mode; or The automatic identification submodule is used to determine the current driving mode of the vehicle based on the vehicle's driving information.

[0100] In one possible implementation, the vehicle driving information includes: vehicle gear information and vehicle speed information; determining the current driving mode of the vehicle based on the vehicle driving information includes: If the vehicle is currently in reverse or neutral, the current driving mode is determined to be parking mode. If the vehicle is currently in drive: Based on the vehicle speed information, if the vehicle speed is determined to be below a first speed, the current driving mode is determined to be parking mode; Based on the vehicle speed information, if the vehicle speed is determined to be above the second speed and the duration exceeds the first duration, the current driving mode is determined as the driving mode. Otherwise, the current driving mode will be set as parking mode.

[0101] A third aspect of this application also provides a vehicle, the vehicle including a perception information acquisition system, the perception information acquisition system being used to perform the steps of the perception information acquisition method described in the first aspect of this application.

[0102] This application also provides an electronic device, see embodiments thereof. Figure 7 , Figure 7 This is a schematic diagram of the structure of the electronic device proposed in the embodiments of this application. Figure 7 As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus for communication. The memory 110 stores a computer program that can run on the processor 120 to implement the steps of the sensing information acquisition method described in the first aspect of the present application.

[0103] This application also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the perception information acquisition method described in the first aspect of this application.

[0104] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the perception information acquisition method as described in the first aspect of this application.

[0105] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0106] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0107] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0109] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0110] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0111] The above provides a detailed description of a method, system, vehicle, and product for acquiring perception information provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for acquiring sensory information, characterized in that, The method includes: Determine the vehicle's current driving mode and acquire the image data corresponding to the current driving mode; Based on the image data corresponding to the current driving mode, the vehicle's perception information is obtained, and the perception information is used to characterize the environment around the vehicle.

2. The method according to claim 1, characterized in that, The step of obtaining the vehicle's perception information based on the image data corresponding to the current driving mode includes: The image data corresponding to the current driving mode is input into the occupancy network model to obtain the vehicle's perception information.

3. The method for acquiring sensory information according to claim 2, characterized in that, The step of inputting the image data corresponding to the current driving mode into the occupancy network model includes: When the current driving mode is driving mode, the image data corresponding to the driving mode is input into the occupancy network model; When the current driving mode is parking mode, the image data corresponding to the parking mode is input into the occupancy network model.

4. The method for acquiring sensory information according to claim 2, characterized in that, The step of inputting the image data corresponding to the current driving mode into the occupancy network model to obtain the vehicle's perception information includes: Determine the perception hyperparameter set corresponding to the current driving mode; the perception hyperparameter set includes at least one of the following: perception range, number of voxel grids, and side length of voxel grids; The occupancy network model, based on the perception hyperparameter set, extracts the vehicle's perception information from the image data corresponding to the current driving mode.

5. The method for acquiring sensory information according to claim 4, characterized in that, Determining the perception hyperparameter set corresponding to the current driving mode includes: When the current driving mode is parking mode, determine one of the parking perception hyperparameter groups corresponding to the current parking scenario from multiple parking perception hyperparameter groups; When the current driving mode is driving mode, determine one of the driving perception hyperparameter groups corresponding to the current driving scenario from multiple driving perception hyperparameter groups; The sensing range of the parking perception hyperparameter group is smaller than that of the driving perception hyperparameter group; the side length of the voxel grid of the parking perception hyperparameter group is smaller than that of the voxel grid of the driving perception hyperparameter group.

6. The method for acquiring sensory information according to claim 2, characterized in that, The image data corresponding to the current driving mode is input into the occupancy network model to obtain the vehicle's perception information, including: Based on the distortion coefficients of the camera used to collect image data corresponding to the current driving mode, distortion removal processing is performed on the image data corresponding to the current driving mode to obtain distortion-free image data. Extract feature maps from the distortion-free image data; Based on the intrinsic and extrinsic parameters of the camera used to collect image data corresponding to the current driving mode, the feature map is converted into a bird's-eye view BEV feature map; Target detection is performed based on the BEV feature map to obtain target detection information in the three-dimensional space where the vehicle is located. Occupation prediction is performed based on the BEV feature map to obtain voxel occupancy status and occupancy semantic information in the three-dimensional space where the vehicle is located.

7. The method for acquiring sensory information according to claim 1, characterized in that, The step of obtaining the image data corresponding to the current driving mode includes: When the current driving mode is driving mode, acquire the image data corresponding to the driving mode; When the current driving mode is parking mode, acquire the image data corresponding to the parking mode.

8. The method for acquiring sensory information according to claim 7, characterized in that, The step of acquiring image data corresponding to the parking mode when the current driving mode is parking mode includes: When the current driving mode of the vehicle is parking mode, the first set of cameras of the vehicle is used to acquire image data corresponding to the parking mode. The first set of cameras includes: a first type of camera with front view, rear view and side view of the vehicle, and a second type of camera with front view and rear view of the vehicle. The step of acquiring image data corresponding to the current driving mode, which is driving mode, includes: When the current driving mode of the vehicle is driving mode, the second set of cameras of the vehicle is used to acquire image data corresponding to the driving mode. The second set of cameras includes: a second type of camera with a front view, a rear view and a side view of the vehicle; wherein, the field of view of the first type of camera is greater than the field of view of the second type of camera, and the maximum working distance of the first type of camera is less than the maximum working distance of the second type of camera.

9. The method for acquiring sensory information according to any one of claims 1-8, characterized in that, The method further includes: Based on the vehicle's perception information, a perception interaction is performed.

10. The method for acquiring sensory information according to claim 9, characterized in that, The step of performing perception interaction based on the vehicle's perception information includes: In response to a user's voice query request regarding the vehicle's surrounding environment, a voice reply is generated based on the vehicle's perception information.

11. The method for acquiring sensory information according to claim 10, characterized in that, Based on the vehicle's perception information, a voice response is generated, including: Using a large vision-language model, image language features are extracted from the vehicle's perceived information; Based on the image language features, the geometric and semantic information of each target is encoded to obtain a language-enhanced map; The voice information in the voice query request is converted into text information and input into the visual-language big model; Based on the language enhancement map, the visual-language big model generates a voice response to the text information.

12. The method for acquiring sensory information according to claim 9, characterized in that, The step of performing perception interaction based on the vehicle's perception information includes: Based on the vehicle's sensor information, control commands are sent to the vehicle's onboard electrical systems.

13. The method for acquiring sensory information according to claim 9, characterized in that, The vehicle's perception information includes: voxel occupancy information and target detection information; the perception interaction based on the vehicle's perception information includes: Replace the voxel occupancy information corresponding to the position in the object detection box of the target detection perception information with the three-dimensional model of the target object represented by the voxel occupancy information; The vehicle's driving environment is displayed on the screen by combining the vehicle's perception information and the target object's three-dimensional model.

14. The method for acquiring sensory information according to any one of claims 1-8, characterized in that, Determine the vehicle's current driving mode, including: The driver's manual selection result is detected, and the current driving mode of the vehicle is determined; or Based on the vehicle's driving information, determine the vehicle's current driving mode.

15. The method for acquiring sensory information according to claim 14, characterized in that, The vehicle driving information includes: vehicle gear information and vehicle speed information; based on the vehicle driving information, the current driving mode of the vehicle is determined, including: If the vehicle is currently in reverse or neutral, the current driving mode is determined to be parking mode. If the vehicle is currently in drive: Based on the vehicle speed information, if the vehicle speed is determined to be below a first speed, the current driving mode is determined to be parking mode; Based on the vehicle speed information, if the vehicle speed is determined to be above the second speed and the duration exceeds the first duration, the current driving mode is determined as the driving mode. Otherwise, the current driving mode will be set as parking mode.

16. A sensing information acquisition system, characterized in that, The system is applied to performing the perceptual information acquisition method as described in any one of claims 1-15, the system comprising: The image data acquisition module is used to determine the current driving mode of the vehicle and acquire the image data corresponding to the current driving mode; The perception information acquisition module is used to acquire the perception information of the vehicle based on the image data corresponding to the current driving mode, and the perception information is used to characterize the environment around the vehicle.

17. The sensing information acquisition system according to claim 16, characterized in that, The sensing information acquisition module includes: An occupancy network model is used to obtain the vehicle's perception information based on the image data corresponding to the current driving mode.

18. The sensing information acquisition system according to claim 16, characterized in that, The system also includes: The first type of camera provides the vehicle's front, rear, and side views; the first type of camera is used to acquire image data corresponding to the parking mode. The vehicle has a second type of camera for front, rear and side views; the second type of camera is used to acquire image data corresponding to parking mode and / or driving mode; wherein, the field of view of the first type of camera is greater than the field of view of the second type of camera, and the maximum working distance of the first type of camera is less than the maximum working distance of the second type of camera.

19. The sensing information acquisition system according to claim 18, characterized in that, The first type of camera is a fisheye camera, and the second type of camera is a panoramic camera.

20. The sensing information acquisition system according to claim 18, characterized in that, The first type of camera for the side view is located on the fender of the vehicle.

21. The sensing information acquisition system according to claim 17, characterized in that, The occupancy network model includes: The distortion correction module is used to perform distortion correction processing on the image data corresponding to the current driving mode based on the distortion coefficient of the camera used to collect the image data corresponding to the current driving mode, so as to obtain distortion-free image data. A backbone network is used to extract feature maps from the distortion-free image data; The BEV feature conversion module is used to convert the feature map into a BEV feature map based on the intrinsic and extrinsic parameters of the camera used to collect image data corresponding to the current driving mode. An object detection feature learning network is used to perform object detection based on the BEV feature map to obtain object detection information in the three-dimensional space where the vehicle is located. An occupancy-aware feature learning network is used to predict occupancy based on the BEV feature map, thereby obtaining the voxel occupancy status and occupancy semantic information of the three-dimensional space in which the vehicle is located.

22. A vehicle, characterized in that, The vehicle includes a perception information acquisition system, which is used to perform the steps of the perception information acquisition method according to any one of claims 1-15.

23. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the perceptual information acquisition method as described in any one of claims 1-15.

24. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the perceptual information acquisition method as described in any one of claims 1-15.