Vehicle control method and device, electronic equipment and storage medium

By collecting and processing image data around the vehicle to generate elevation feature maps, the problem of accuracy fluctuation of traditional sensors under complex road conditions is solved, enabling vehicles to accurately perceive and make safe decisions in complex environments.

CN121469550BActive Publication Date: 2026-07-24XIAOMI EV TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAOMI EV TECH CO LTD
Filing Date
2025-12-04
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional vehicle sensors, such as ultrasonic radar, are susceptible to interference from complex road conditions, leading to fluctuations in accuracy and making it difficult to meet the needs of real-time obstacle avoidance decision-making.

Method used

By collecting image data around the vehicle, extracting features and fusing them into a first voxel to generate an elevation feature map, which is used to characterize the height of ground protrusions or the depth of depressions, and controlling the vehicle to execute commands based on the elevation feature map.

Benefits of technology

It improves the vehicle's ability to perceive complex road conditions and enhances driving safety, ensuring that the vehicle makes intelligent and reasonable decisions in various environments and reducing the occurrence of collision accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121469550B_ABST
    Figure CN121469550B_ABST
Patent Text Reader

Abstract

The application provides a vehicle control method and device, electronic equipment and storage medium, wherein the method comprises: collecting image data of the surrounding of the vehicle, extracting features from the image data to obtain image feature data; fusing the image feature data into a first voxel body to obtain an elevation feature map; wherein the parameters in the first voxel body are parameters of the ground perceived by the vehicle; and controlling the vehicle to execute corresponding instructions based on the elevation feature map. By collecting and processing image data, an elevation feature map that can accurately represent the ground conditions is generated, enabling the vehicle to have more accurate perception of the surrounding environment, making up for the shortcomings of traditional environmental perception methods, and improving the accuracy and comprehensiveness of perception. The vehicle control system provides intuitive and key environmental information, and based on the elevation feature map, the vehicle can make more intelligent and reasonable decisions, effectively cope with various complex road conditions, and improve the safety and stability of vehicle driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the application of image data technology in the field of vehicles, and more particularly to a vehicle control method, device, electronic device, and storage medium. Background Technology

[0002] In the field of modern surveying and geographic information, accurate determination of ground elevation data plays an indispensable and crucial role in many important applications such as urban planning, terrain modeling, resource exploration, and environmental monitoring.

[0003] There are significant shortcomings in road height detection methods in related technologies: traditional vehicle-mounted sensors (such as ultrasonic radar) are easily affected by complex road conditions, resulting in fluctuations in accuracy, which makes it difficult to meet the needs of real-time obstacle avoidance decision-making. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, this application proposes a method, apparatus, electronic device, and storage medium.

[0006] One embodiment of this application proposes a vehicle control method, including: Collect image data of the area surrounding the vehicle, and extract features from the image data to obtain image feature data; The image feature data is fused into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of the ground protrusion or the depth of the depression; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. The vehicle is controlled to execute corresponding commands based on the elevation feature map.

[0007] By collecting and deeply analyzing image data around the vehicle, the generated elevation feature map accurately reflects the undulations of the ground, enabling the vehicle to perceive its driving environment more precisely and avoiding driving risks caused by insufficient environmental perception. The elevation feature map provides the vehicle control system with intuitive and crucial environmental information. Based on this information, the vehicle can make more intelligent and rational decisions, effectively cope with various complex road conditions, and significantly improve the safety and stability of vehicle operation.

[0008] Optionally, fusing the image feature data into a first voxel to obtain an elevation feature map includes: The image feature data is sampled and incorporated into the first voxel to obtain the second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

[0009] The sampling and fusion process can more accurately combine image feature data with ground parameters, generating an elevation feature map containing rich and accurate elevation information. This allows vehicles to have a more detailed understanding of ground conditions, accurately perceive even minor terrain changes, greatly improve their ability to perceive complex terrain, and provide a more precise basis for vehicle path planning.

[0010] Optionally, the method further includes: According to a preset number of iterations, the image feature data is sampled multiple times and sequentially fused into the first voxel of the current round to iteratively update the elevation feature map; wherein, the first voxel of the current round is constructed based on the elevation feature map of the previous round.

[0011] Through multiple iterative fusions, the elevation feature map can be continuously optimized, reducing errors and more accurately reflecting the true ground conditions. This provides vehicles with more precise environmental information, enabling them to more accurately assess road conditions and make reasonable decisions during driving, such as more accurately planning avoidance routes and adjusting speed, thereby effectively avoiding collisions and improving driving safety.

[0012] Optionally, the step of repeatedly fusing the image feature data into a first voxel to update the elevation feature map includes: The values ​​of each voxel in the reference plane in the updated first voxel volume are determined based on the elevation feature map; wherein the height range and height sampling interval in the updated first voxel volume are smaller than the height range and height sampling interval of the first voxel volume before the update.

[0013] By narrowing the height range and height sampling interval, the first voxel can capture subtle features and changes in the ground more precisely, enabling the generated elevation feature map to more accurately reflect the detailed information of the ground. This is crucial for vehicles driving on complex terrain, such as roads covered with small stones or with minor cracks. Vehicles can adjust their suspension system or speed in advance based on a more accurate elevation feature map, ensuring a smooth and safe ride.

[0014] Optionally, the image data includes single-frame image data or multi-frame image data, and the step of extracting features from the image data to obtain image feature data includes: The image data of one or more frames is input into the feature extraction module to perform multi-scale feature extraction to obtain multiple first feature data. Scale each of the first feature data to the same size to obtain the second feature data; The image feature data is obtained by fusing the various second feature data.

[0015] Scale feature extraction can extract feature information at different levels from image data, including both macroscopic scene structure features and microscopic detail features. This comprehensive feature extraction method makes the extracted image feature data richer and more complete, helping vehicles to more accurately understand image content and driving environment conditions, thus improving the accuracy and depth of environmental perception.

[0016] Optionally, sampling the image feature data and incorporating it into the first voxel to obtain the second voxel includes: The three-dimensional mesh sampling function is determined based on the feature matrix corresponding to the image feature data; The image feature data is filled into each voxel of the first voxel according to the three-dimensional network sampling function to obtain the second voxel.

[0017] Determining the 3D mesh sampling function based on the feature matrix enables precise mapping of image feature data to the first voxel. This precision ensures that image feature data is accurately integrated into the voxel, avoiding information loss or misallocation and improving data processing accuracy. Precise data mapping guarantees the generation of high-quality elevation feature maps, enabling them to accurately reflect the actual ground conditions and provide reliable environmental information for vehicles. By filling image feature data into each voxel of the first voxel, effective integration of image feature data information with ground parameter information within the first voxel is achieved. The second voxel not only contains the basic ground parameters represented by the first voxel but also incorporates environmental details carried by the image feature data, such as the shape and location of obstacles and ground texture. This integration makes the second voxel a richer, more information-rich data structure, providing comprehensive data support for subsequent elevation feature map generation and helping vehicles gain a more comprehensive understanding of the driving environment.

[0018] Optionally, the step of filling the image feature data into each voxel of the first voxel body according to the three-dimensional network sampling function includes any one of the following: For single-image feature data, the feature value in the image feature data is determined in the first voxel according to the three-dimensional network sampling function, and the feature value is filled into the corresponding voxel; For multiple image feature data, the feature value in each image feature data is determined in the first voxel according to the three-dimensional network sampling function, and the feature value corresponding to the same voxel is fused and filled into the corresponding voxel.

[0019] Direct infilling of single image feature data accurately preserves the original information of each individual feature, enabling the second voxel to precisely reflect the position and attributes of each detailed feature in the image on the ground. This is crucial for accurately depicting subtle changes and special features on the ground, such as cracks and potholes. This detailed information is essential for vehicles to assess driving safety, helping them make evasive decisions in advance and ensuring driving safety.

[0020] The fusion and filling method for multiple image feature data can effectively integrate various feature information at the same location, avoiding information omissions or conflicts. Through a reasonable fusion method, the second voxel can comprehensively reflect the integrated features of complex regions in the image, such as regions that simultaneously contain obstacle and ground texture information. This comprehensive information integration provides strong support for generating more accurate and detailed elevation feature maps, enabling vehicles to have a more comprehensive understanding of the driving environment and make more reasonable control decisions, such as accurately planning driving paths in complex road conditions.

[0021] Optionally, fusing the first voxel and the second voxel to determine the elevation feature map includes: The second voxel is input into the pre-trained first model to obtain the first elevation feature vector of the output. The elements in the first elevation feature vector are fused with the values ​​in the corresponding voxels in the first voxel to obtain the elevation values ​​of each voxel in the first voxel. Elevation values ​​corresponding to voxels with the same horizontal and vertical coordinates in the first voxel are merged to obtain elevation data, and the various elevation data are combined to obtain the elevation feature map.

[0022] By utilizing a pre-trained height feature extraction model, precise height features can be effectively extracted from complex voxel data. Fusing these features with the parameters of the first voxel further improves the accuracy of the elevation values. This allows the generated elevation feature map to accurately reflect the actual undulations of the ground, providing vehicles with reliable ground information and helping them accurately determine the safety of their driving paths. By fusing information from the first and second voxels, the advantages of both are fully utilized. The first voxel provides the basic parameter framework for the ground, while the second voxel incorporates rich image feature data. The elevation feature map generated by combining the two contains more comprehensive and accurate ground information. Compared to using either voxel alone, it more realistically reflects the actual ground conditions and enhances the vehicle's environmental perception capabilities.

[0023] Optionally, controlling the vehicle to execute corresponding instructions based on the elevation feature map includes at least one of the following: Control the vehicle to execute vehicle avoidance control commands; Control the vehicle to execute a collision risk warning command; Control the vehicle to execute the door locking command to prevent opening.

[0024] Vehicle avoidance control commands enable vehicles to avoid obstacles and poor road conditions in a timely manner, effectively reducing the probability of collisions and ensuring the safety of the vehicle and its occupants. Whether encountering pedestrians or vehicles suddenly appearing on urban roads or facing complex terrain in the countryside, vehicles can ensure driving safety through accurate avoidance maneuvers.

[0025] Collision risk warning commands provide drivers with an opportunity to anticipate potential hazards. Timely warnings allow drivers to be more alert and take preventative measures such as braking or swerving in advance, thus avoiding collisions. This warning mechanism is not only applicable to autonomous vehicles but also crucial for human-driven vehicles, assisting drivers in better assessing road conditions and improving driving safety.

[0026] The "Doors Lock and Do Not Open" instruction provides additional safety for occupants when the vehicle is in a dangerous area. It prevents accidental opening of the doors, avoiding injury in dangerous situations, especially in extreme circumstances where the vehicle may roll over or tilt; this instruction effectively protects the lives of those inside.

[0027] Optionally, the method further includes: The vehicle's behavior is displayed through an in-vehicle display device to inform the user of the instructions executed by the vehicle; The vehicle's actions are played through a speaker to inform the user of the commands executed by the vehicle; The information corresponding to the vehicle's behavior is sent to the target device to inform the user of the instructions executed by the vehicle.

[0028] Through multiple feedback methods, users can gain a comprehensive understanding of the vehicle's behavior and operating status from different perspectives, enhancing their perception and understanding of the vehicle's intelligent control system. Visual displays, voice prompts, and remote information push notifications enable users to obtain key vehicle information promptly, whether inside or outside the vehicle, improving the user-vehicle interaction experience.

[0029] Another embodiment of this application provides a vehicle control device, including: The feature extraction module is used to collect image data around the vehicle and extract features from the image data to obtain image feature data; An elevation determination module is used to fuse the image feature data into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of the ground protrusion or the depth of the depression; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. The control module is used to control the vehicle to execute corresponding instructions based on the elevation feature map.

[0030] Optionally, the elevation determination module includes: A sampling submodule is used to sample the image feature data and incorporate it into the first voxel to obtain a second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

[0031] Optionally, the device further includes: An iteration module is used to sample the image feature data multiple times according to a preset number of iterations and sequentially fuse it into the first voxel of the current round to iteratively update the elevation feature map; wherein the first voxel of the current round is constructed based on the elevation feature map of the previous round.

[0032] Another embodiment of this application proposes a vehicle including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method described in the foregoing aspect.

[0033] Another embodiment of this application proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0034] Another embodiment of this application proposes a chip including processing circuitry configured to perform the method described in one aspect above.

[0035] Another embodiment of this application proposes a computer program product that, when executed by a processor, implements the method described in the foregoing aspect.

[0036] The vehicle control method, device, electronic equipment, chip, and storage medium proposed in this application generate elevation feature maps that accurately characterize ground conditions by acquiring and processing image data. This enables vehicles to have a more precise perception of their surrounding environment, compensating for the shortcomings of traditional environmental perception methods and improving the accuracy and comprehensiveness of perception. The vehicle control system provides intuitive and crucial environmental information. Based on the elevation feature maps, vehicles can make more intelligent and rational decisions, effectively cope with various complex road conditions, and improve the safety and stability of vehicle operation.

[0037] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0038] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a schematic flowchart of a vehicle control method provided in an embodiment of this application; Figure 2 This is a schematic diagram of an autonomous driving scenario provided in this embodiment; Figure 3 This is a schematic diagram of an automatic parking scenario provided in this embodiment; Figure 4 This is a schematic diagram of another vehicle control system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a vehicle control device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a chip proposed in an embodiment of this application. Detailed Implementation

[0039] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0040] Ground elevation estimation (GPE) technology, a crucial research area in autonomous driving and robotics, aims to accurately reconstruct surface elevation information through perception systems. This enables refined modeling of complex terrains and provides intelligent systems with reliable environmental understanding capabilities. In autonomous driving scenarios, GPE is widely used to detect potholes, curbs, and other elevation changes, supporting real-time 3D road surface reconstruction, adaptive suspension control, and path planning, thereby improving vehicle safety and comfort. In robotics, this technology helps mobile robots perceive elevation differences in complex terrains (such as uneven ground or obstacle areas), achieving stable navigation and obstacle avoidance. Currently, mainstream methods for GPE include vision-based perspective monocular depth estimation, binocular depth estimation, LiDAR-based point cloud processing, and vision-based BEV (Bird's EyeView) feature methods. Monocular depth estimation relies on a single camera, offering low cost but limited accuracy and susceptibility to lighting and texture variations. Binocular depth estimation improves accuracy through stereo vision, but suffers from high parallax computation complexity and is limited by baseline distance. LiDAR solutions excel in distance measurement thanks to high-precision point cloud data, but are costly and susceptible to weather conditions. Visual BEV methods convert perspective images into top-down elevation maps, balancing accuracy and computational efficiency, but still face challenges in occluded and long-distance scenarios. Each approach has its advantages and disadvantages. Monocular and BEV methods excel in cost and real-time performance, making them suitable for large-scale deployment, while binocular and LiDAR methods offer superior accuracy and robustness. Future development of ground elevation estimation technology will focus on multi-sensor fusion (such as combining vision and LiDAR), deep learning algorithm optimization, and real-time performance improvements to address diverse scenario requirements and propel autonomous driving and robotic systems towards higher levels of intelligence.

[0041] The vehicle control method, apparatus, electronic device, chip, and storage medium of this application are described below with reference to the accompanying drawings.

[0042] Figure 1 This is a schematic diagram of a vehicle control process provided in an embodiment of this application.

[0043] In one implementation, the vehicle control method of this application embodiment can be configured in a vehicle control device, which can be applied to any electronic device so that the electronic device can perform vehicle control functions.

[0044] Among them, electronic devices can be any device with computing capabilities, such as mobile terminals, which can be hardware devices with various operating systems, touch screens and / or displays, such as mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0045] As another implementation, the vehicle control method of this application embodiment can also be executed by a chip with processing capabilities. The chip includes an image signal processing chip (ISP), a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a system on a chip (SOC), a reduced instruction set computer (RISC), etc., which will not be listed here.

[0046] It should be noted that all data collection operations related to users in this application are conducted with the user's authorization and in strict compliance with relevant laws and regulations such as privacy and security.

[0047] like Figure 1 As shown, the method may include the following steps: Step 101: Collect image data of the area surrounding the vehicle and extract features from the image data to obtain image feature data; Step 102: The image feature data is fused into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of the ground protrusion or the depth of the depression; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. Step 103: Control the vehicle to execute corresponding instructions based on the elevation feature map.

[0048] In this embodiment, against the backdrop of today's intelligent transportation development, vehicle safety and intelligence have become key concerns. Traditional vehicle control methods often rely on relatively simple sensor data, making it difficult to comprehensively and accurately perceive the complex and ever-changing driving environment. Image-based vehicle control technology, however, can utilize the rich visual information collected by cameras to provide a more comprehensive and accurate basis for vehicle control decisions. By extracting features from the image data surrounding the vehicle and fusing them to generate an elevation feature map, the undulations of the ground can be intuitively presented. This allows the vehicle to execute appropriate control commands based on this detailed information, effectively improving vehicle safety and intelligence.

[0049] Image data of the vehicle's surrounding environment, captured by cameras installed around the vehicle, encompasses a wealth of information including roads, surrounding objects, and terrain, serving as the raw data foundation for subsequent analysis and decision-making.

[0050] Image feature data are representative data extracted from image data through specific feature extraction algorithms. They can reflect the key attributes of objects in the image, such as shape, texture, and edge features, which helps to deepen the understanding and analysis of image content.

[0051] The first voxel is a three-dimensional data structure whose parameters are related to the ground perceived by the vehicle, such as the size, spatial distribution, and attributes corresponding to ground features. It is used to carry and process data related to ground information and provides a basic framework for generating elevation feature maps.

[0052] An elevation feature map is an information map that visually displays the height of ground protrusions or the depth of depressions in a graphical manner. It is generated by fusing image feature data into a first voxel, providing key environmental terrain information for vehicle control systems and helping vehicles determine the safety of their driving paths.

[0053] The vehicle uses its onboard cameras to collect comprehensive image data of its surroundings. This image data acts like the vehicle's "eyes," recording the real-time conditions of the environment. Subsequently, specialized feature extraction algorithms process the collected image data, extracting essential image features that characterize the content of the complex images. For example, edge detection algorithms can extract features such as road boundaries and obstacle outlines, providing crucial information for subsequent analysis and decision-making.

[0054] The extracted image feature data is integrated into a first voxel, and through a series of complex data processing and fusion operations, an elevation feature map is generated. In this process, the image feature data interacts with preset ground parameters within the first voxel. For example, feature data representing protruding objects in the image is combined with the corresponding voxel in the first voxel, thus presenting the corresponding protrusion height in the elevation feature map; similarly, concave areas are represented by negative height values ​​in the map.

[0055] The vehicle control system performs intelligent analysis and decision-making based on the ground conditions reflected in the generated elevation feature map, and then controls the vehicle to execute corresponding commands. For example, when the elevation feature map shows that there are large bumps or depressions on the ground ahead that may affect the vehicle's driving safety, the vehicle will automatically execute control commands such as deceleration and avoidance to ensure a smooth and safe driving process.

[0056] Analysis of beneficial effects: Precise environmental perception: By collecting and deeply analyzing image data around the vehicle, the generated elevation feature map can accurately reflect the undulation of the ground, making the vehicle's perception of the driving environment more accurate and avoiding driving risks caused by insufficient environmental perception.

[0057] Intelligent decision support: Elevation feature maps provide intuitive and crucial environmental information for vehicle control systems. Based on this information, vehicles can make more intelligent and rational decisions, effectively cope with various complex road conditions, and significantly improve the safety and stability of vehicle operation.

[0058] Versatility and Adaptability: Based on widely used image data processing technology, this method has strong versatility and can adapt to a variety of different driving scenarios. Whether it is the complex road conditions in urban areas or special terrains such as rural areas and wilderness, it can generate accurate elevation feature maps through the analysis of image data, providing strong support for vehicle control.

[0059] Optionally, fusing the image feature data into a first voxel to obtain an elevation feature map includes: The image feature data is sampled and incorporated into the first voxel to obtain the second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

[0060] In this embodiment, generating an accurate elevation feature map that reflects ground conditions is a core step in achieving precise vehicle control. A detailed and reasonable operational procedure ensures the effective fusion of image feature data and ground parameters, thereby more accurately presenting ground undulation information. By first sampling image feature data and integrating it into a first voxel to generate a second voxel, and then fusing the two voxels to determine the elevation feature map, this step-by-step processing method helps to construct a more detailed elevation feature map, providing the vehicle with detailed and accurate ground information, and thus supporting the vehicle in making precise control decisions.

[0061] A first voxel is generated based on ground parameters perceived by the vehicle. These parameters may include information such as the approximate flatness of the ground and the expected height range, thus constructing a three-dimensional data framework. Then, image feature data is sampled, meaning discrete sample points are selected from continuous image feature data according to certain rules. Using a specific algorithm, these sampled data are integrated into the voxels of the first voxel, generating a second voxel. In this process, information from the image feature data is discretized and allocated into the three-dimensional space of the first voxel, so that the second voxel contains both basic ground parameter information and environmental detail features carried by the image. For example, if the image feature data contains road edge features, these features will be reflected at corresponding locations in the second voxel through sampling and integration operations.

[0062] The fusion process between the first and second voxels aims to combine the advantages of both to generate an accurate elevation feature map. A specially designed fusion algorithm, such as a weighted average or neural network-based fusion method, organically combines the information representing basic ground parameters from the first voxel with the image feature data incorporated from the second voxel. During the fusion process, the information of each voxel is recalculated and adjusted. In the final elevation feature map, each pixel corresponds to the precise elevation data of its corresponding location in the image data. This elevation data accurately reflects the height of ground undulations or the depth of ground depressions. For example, for a specific region in the image, its elevation value in the elevation feature map is obtained by fusing the corresponding voxel information from the first and second voxels, thus accurately presenting the actual undulations of the ground in that region.

[0063] Analysis of beneficial effects: High-precision elevation information acquisition: The multi-step sampling and fusion process can more accurately combine image feature data with ground parameters, generating an elevation feature map containing rich and accurate elevation information. This allows vehicles to have a more detailed understanding of ground conditions and accurately perceive even minor terrain changes, greatly improving their ability to perceive complex terrain and providing a more precise basis for vehicle path planning.

[0064] Data optimization and refinement: By first integrating image feature data into a first voxel to generate a second voxel, and then fusing them to determine the elevation feature map, this step-by-step processing method helps to optimize and refine the data. At each step, the data can be specifically adjusted and improved to compensate for potential information loss or inaccuracies, resulting in a final elevation feature map that more accurately reflects the actual ground conditions and provides more reliable decision support for vehicle control.

[0065] Optionally, the method further includes: According to a preset number of iterations, the image feature data is sampled multiple times and sequentially fused into the first voxel of the current round to iteratively update the elevation feature map; wherein, the first voxel of the current round is constructed based on the elevation feature map of the previous round.

[0066] In this embodiment, The system initiates an iterative fusion process based on a preset number of iterations. In each iteration, the voxel volume for the current iteration is first constructed based on the elevation feature map generated in the previous iteration. The elevation feature map from the previous iteration reflects the general ground conditions at that time. Based on this, the parameters of the voxel volume are adjusted and optimized. For example, parameters such as voxel size, height range, and spatial distribution are adjusted accordingly based on factors such as the complexity and undulation of the terrain shown in the elevation feature map. If the previous elevation feature map shows many minor undulations in the ground, the voxel size of the current iteration may be reduced to capture details more accurately; the height range may also be shrunk or expanded according to the actual terrain undulations. Then, image feature data is sampled, and the sampled values ​​are integrated into the voxel volume of the current iteration. The sampling method and integration rules are similar to the process of generating the second voxel volume, but are optimized accordingly based on the parameter adjustments of the current iteration voxel volume. In this way, the information in the voxel volume is continuously updated to more accurately reflect changes in the ground conditions. Finally, the voxels incorporating image feature data are processed to generate an elevation feature map again, which is further improved in terms of accuracy and detail compared to the previous one.

[0067] As the number of iterations increases, each updated elevation feature map more accurately reflects the actual ground conditions compared to the previous iteration. This is because each iteration fully considers previous results and new image feature data, allowing the elevation feature map to gradually converge to a more accurate state. For example, in the initial iterations, the elevation feature map may only roughly depict the main terrain changes, such as large bumps and depressions; however, with increasing iterations, the elevation feature map increasingly displays detailed information such as the specific height of bumps, the depth of depressions, their boundaries, and minor terrain undulations. Simultaneously, due to the dynamic changes in the vehicle's driving environment, information such as newly appearing obstacles or terrain changes can also be reflected in the elevation feature map in a timely manner through iterative fusion, ensuring that the vehicle always obtains the latest and most accurate environmental information.

[0068] Analysis of beneficial effects: Improving the accuracy of elevation feature maps: Through multiple iterative fusions, elevation feature maps can be continuously optimized, reducing errors and more accurately reflecting the true ground conditions. This provides vehicles with more precise environmental information, enabling them to more accurately judge road conditions and make reasonable decisions during driving, such as more accurately planning avoidance routes and adjusting speed, thereby effectively avoiding collisions and improving driving safety.

[0069] Adapting to dynamic environmental changes: During vehicle operation, the dynamic changes in the environment are unpredictable. Multiple iterative fusion methods enable the elevation feature map to keep pace with these changes, updating ground information in real time. Whether it's a sudden obstacle or a change in road conditions due to construction or other reasons, the elevation feature map can quickly adjust, providing the vehicle with the latest environmental data and ensuring safe driving in dynamic environments.

[0070] Enhanced system stability: The elevation feature map, optimized through multiple iterations, provides more reliable foundational information for the vehicle control system, reducing vehicle control errors caused by inaccurate environmental information. This enhances the stability and reliability of the entire vehicle control system, enabling stable operation of the vehicle in various complex environments, reducing the probability of system failure, and increasing user trust in the vehicle's intelligent control system.

[0071] Optionally, the step of repeatedly fusing the image feature data into a first voxel to update the elevation feature map includes: The values ​​of each voxel in the reference plane in the updated first voxel volume are determined based on the elevation feature map; wherein the height range and height sampling interval in the updated first voxel volume are smaller than the height range and height sampling interval of the first voxel volume before the update.

[0072] In this embodiment, image feature data is fused into the first voxel multiple times to update the elevation feature map, which is crucial for further optimizing the accuracy and adaptability of the elevation feature map. During each iteration, by configuring the parameters of the first voxel in this iteration and adjusting the reference plane voxel value of the first voxel based on the previously generated elevation feature map, particularly by narrowing the height range and height sampling interval, ground details can be captured more precisely. This improves the elevation feature map's ability to describe complex terrain, provides the vehicle with more accurate ground information, and enables it to make more precise control decisions in complex and changing driving environments.

[0073] Explanation of the principle: The values ​​of each voxel in the reference plane within the updated first voxel volume are determined based on the elevation feature map. The reference plane is typically a baseline plane used to construct the first voxel volume. Its voxel values ​​are adjusted using the elevation feature map to ensure the first voxel volume better reflects the actual ground conditions. For example, based on elevation information at different locations in the elevation feature map, the voxel values ​​of the reference plane are adjusted to better match the actual ground height variations.

[0074] The updated first voxel has a smaller height range and height sampling interval than the original first voxel. Reducing the height range focuses more on the actual height variations on the ground, avoiding the introduction of unnecessary noise or inaccurate information from an excessively large height range. For example, if the previous elevation feature map showed that ground height variations were mainly concentrated between -10 cm and 20 cm, the updated first voxel's height range might be reduced to -5 cm to 15 cm. Reducing the height sampling interval means that the number of voxels increases within the same height range, allowing for more detailed capture of height variations. For example, instead of sampling one voxel every 5 cm, the updated voxel might be sampled every 2 cm. This adjustment allows the first voxel to more accurately describe subtle undulations and changes in the ground, laying the foundation for subsequent fusion of image feature data to generate a more accurate elevation feature map.

[0075] Analysis of beneficial effects: Improved detail capture capability: By narrowing the height range and height sampling interval, the first voxel can capture subtle features and changes in the ground more precisely, enabling the generated elevation feature map to more accurately reflect the detailed information of the ground. This is crucial for vehicles driving on complex terrain, such as roads covered with small stones or with minor cracks. Vehicles can adjust their suspension system or speed in advance based on a more accurate elevation feature map, ensuring smooth and safe driving.

[0076] Enhanced terrain adaptability: The parameters of the first body are dynamically adjusted based on the previous elevation feature map, enabling the entire system to better adapt to different terrain conditions. Whether it is a flat urban road or a mountainous road with large undulations, the system can generate an accurate elevation feature map that reflects the terrain by reasonably adjusting the height range and sampling interval, thereby improving the vehicle's driving adaptability in various complex environments.

[0077] Optimizing decision-making accuracy: More accurate elevation feature maps provide vehicles with more precise ground information, enabling the vehicle control system to make more accurate control decisions. For example, when encountering road sections with complex gradient changes, the vehicle can adjust power output and braking force in advance based on the accurate elevation feature map, ensuring a safe and smooth driving process and effectively avoiding decision-making errors caused by inaccurate information.

[0078] Optionally, the image data includes single-frame image data or multi-frame image data, and the step of extracting features from the image data to obtain image feature data includes: The image data of one or more frames is input into the feature extraction module to perform multi-scale feature extraction to obtain multiple first feature data. Scale each of the first feature data to the same size to obtain the second feature data; The image feature data is obtained by fusing the various second feature data.

[0079] In this embodiment, accurate extraction of image feature data is a crucial foundation for achieving precise vehicle control. During actual driving, the image data acquired by the vehicle may be a single frame or multiple consecutive frames. Employing multi-scale feature extraction, size scaling, and feature fusion methods allows for comprehensive mining of information from different perspectives, overcoming differences in image data scale and resolution, and improving the quality and usability of image feature data. This provides strong support for subsequently generating accurate elevation feature maps and enabling the vehicle to make reasonable control decisions, ensuring that the vehicle can accurately perceive its environment in various complex scenarios and achieve safe and intelligent driving.

[0080] Single or multiple frames of image data are input into a dedicated feature extraction module. This module employs a multi-scale feature extraction method, achieved by setting convolutional kernels of different sizes in different convolutional layers. Large-sized convolutional kernels capture macroscopic features in the image, such as the overall direction of roads and the layout of surrounding buildings; these features reflect the general structure and scene information of the image. Small-sized convolutional kernels focus on extracting microscopic details in the image, such as pebbles, cracks, and details of traffic signs on the road surface. Through this multi-scale processing approach, multiple primary feature data are obtained from the image data. Each primary feature data represents unique feature information of the image at different scales, thus comprehensively covering various information levels in the image.

[0081] Since the first feature data extracted at different scales may differ in size, it is necessary to scale each first feature data to the same size to obtain the second feature data in order to facilitate subsequent fusion operations. The size scaling process typically employs interpolation algorithms, such as bilinear interpolation or bicubic interpolation. These algorithms, while preserving feature information, adjust feature data of different sizes to a uniform specification by recalculating and redistributing pixel values. In this way, all second feature data maintain a consistent size, providing convenient conditions for subsequent feature fusion and ensuring that feature information from different scales can be effectively integrated in the same dimension.

[0082] The various scaled second feature data are then fused. The fusion method can be chosen based on actual needs, with common methods including stitching and weighted summation. For example, stitching connects different second feature data along a specific dimension (such as the channel dimension) to form a new data structure containing feature information at all scales; alternatively, weighted summation assigns weights to each second feature data based on its importance to the overall image information, and then sums them to obtain the final image feature data. The fused image feature data integrates feature information from different scales, containing a richer and more comprehensive description of the image content, more accurately reflecting the environmental information represented by the image, and providing a more detailed data foundation for subsequent generation of elevation feature maps.

[0083] Analysis of beneficial effects: Rich Feature Information Acquisition: Multi-scale feature extraction can extract feature information at different levels from image data, including both macroscopic scene structure features and microscopic detail features. This comprehensive feature extraction method makes the extracted image feature data richer and more complete, helping vehicles to more accurately understand image content and driving environment conditions, and improving the accuracy and depth of environmental perception.

[0084] Improving feature availability: Size scaling and feature fusion operations ensure the effective integration of feature information at different scales, making the final image feature data more usable. Feature data with a unified size can be processed and analyzed within the same framework, while the fusion operation complements the advantages of features at different scales, forming a more representative and informative feature set. This allows the image feature data to better meet the needs of subsequent generation of elevation feature maps and vehicle control decisions, improving the overall system performance and accuracy.

[0085] Enhanced adaptability: This universal feature extraction method works effectively for both single-frame and multi-frame image data, enhancing the system's adaptability to different types of image data. In actual driving, vehicles may encounter various complex scenes and lighting conditions. This method can adapt to these changes, accurately extracting image feature data and ensuring stable and reliable vehicle operation in different environments.

[0086] Optionally, sampling the image feature data and incorporating it into the first voxel to obtain the second voxel includes: The three-dimensional mesh sampling function is determined based on the feature matrix corresponding to the image feature data; The image feature data is filled into each voxel of the first voxel according to the three-dimensional network sampling function to obtain the second voxel.

[0087] In this embodiment, accurately sampling and integrating image feature data into the first voxel is a crucial step in generating the elevation feature map. By determining a three-dimensional mesh sampling function based on the feature matrix corresponding to the image feature data, and filling the image feature data into each voxel of the first voxel according to this function, an effective combination of image feature data and the first voxel can be achieved. This step not only relates to the accuracy of the elevation feature map generation but also plays a decisive role in whether the vehicle can make reasonable control decisions based on accurate ground information. A reasonable and accurate sampling and filling process helps to retain key information in the image feature data and accurately map it into the three-dimensional space of the first voxel, laying a solid foundation for the subsequent generation of an elevation feature map that accurately reflects the ground conditions.

[0088] First, the feature matrix corresponding to the image feature data is analyzed in depth. This feature matrix contains rich image feature data information, such as the location, intensity, and orientation of the features. Based on the understanding and analysis of this information, a specific algorithm is used to determine the three-dimensional mesh sampling function. This sampling function defines how to sample from continuous image feature data and how to accurately map the sampled values ​​to the three-dimensional space of the first voxel. It fully considers the spatial relationship and data correspondence between the image feature data and the first voxel, ensuring that the sampling process can accurately capture the key information in the image feature data and rationally allocate it to the various voxel positions of the first voxel. For example, if a feature in the image feature data represents a protrusion on the ground, the sampling function will determine which voxel in the first voxel should carry the relevant information of the protrusion based on the location of the feature in the image and the spatial structure of the first voxel.

[0089] Based on a defined 3D network sampling function, image feature data is filled into each voxel of the first voxel. Specifically, the sampling function precisely determines the voxel position corresponding to each feature value in the image feature data within the first voxel. Then, this feature value is filled into the corresponding voxel, thus obtaining the second voxel. In this process, the image feature data is discretized and rationally distributed within the 3D space of the first voxel, enabling the second voxel to effectively carry the image feature data information. For example, for a feature value representing a road edge in the image, the sampling function determines its corresponding voxel in the first voxel and fills it in, so the second voxel contains the position and feature information of the road edge in 3D space. In this way, the second voxel integrates the image feature data with the ground parameter information of the first voxel, providing a rich data foundation for the subsequent generation of elevation feature maps.

[0090] Analysis of beneficial effects: Precise data mapping: By determining the 3D mesh sampling function based on the feature matrix, precise mapping of image feature data to the first voxel volume can be achieved. This precision ensures that image feature data information is accurately integrated into the voxel volume, avoiding information loss or misallocation and improving the accuracy of data processing. Precise data mapping provides a guarantee for generating high-quality elevation feature maps, enabling the elevation feature maps to accurately reflect the actual ground conditions and provide reliable environmental information for vehicles.

[0091] Effective Information Integration: By filling image feature data into each voxel of the first voxel, effective integration of image feature data information with ground parameter information within the first voxel is achieved. The second voxel not only contains the basic ground parameters represented by the first voxel but also incorporates environmental details carried by the image feature data, such as the shape and location of obstacles and ground texture. This integration makes the second voxel a richer and more information-rich data structure, providing comprehensive data support for subsequent elevation feature map generation and helping vehicles gain a more comprehensive understanding of the driving environment.

[0092] Optionally, the step of filling the image feature data into each voxel of the first voxel body according to the three-dimensional network sampling function includes any one of the following: For single-image feature data, the feature value in the image feature data is determined in the first voxel according to the three-dimensional network sampling function, and the feature value is filled into the corresponding voxel; For multiple image feature data, the feature value in each image feature data is determined in the first voxel according to the three-dimensional network sampling function, and the feature value corresponding to the same voxel is fused and filled into the corresponding voxel.

[0093] In this embodiment, different filling methods are used for single and multiple image feature data during the process of filling image feature data into the first voxel. This helps to process the image feature data more rationally and improve the accuracy and completeness of the data in the second voxel. This is crucial for generating high-quality elevation feature maps, as the accuracy of the elevation feature map directly affects the vehicle's judgment and control decisions regarding the driving environment. A reasonable filling method ensures that the information in the image feature data is accurately reflected in the second voxel, thereby providing reliable data support for generating accurate elevation feature maps and ensuring that the vehicle makes appropriate driving decisions based on accurate ground information.

[0094] When processing single image feature data, the voxel corresponding to the feature value in the image feature data in the first voxel volume is determined according to the 3D network sampling function. For example, assuming there is a feature representing a road surface crack in the image feature data, the 3D network sampling function can accurately calculate the voxel position corresponding to the crack feature value in the 3D space of the first voxel volume. Then, the feature value is directly filled into the corresponding voxel. This filling method is suitable for processing relatively independent image feature data with clear correspondences. It can accurately integrate single feature information into the first voxel volume, retain the original information of the feature, and enable the second voxel volume to accurately reflect the position and attributes of the single feature on the ground, providing accurate detailed information for the subsequent generation of elevation feature maps.

[0095] For multiple image feature data, the 3D network sampling function is used to determine the voxel corresponding to the feature value in each image feature data in the first voxel volume. Since multiple image feature data may correspond to the same voxel location—for example, in an image, a region may simultaneously contain ground texture features and small obstacle features, these two features may correspond to the same voxel in the first voxel volume—it is necessary to fuse the feature values ​​corresponding to the same voxel before filling them into the corresponding voxel. The fusion method can be chosen according to the actual situation, such as weighted averaging, where each feature value is assigned a corresponding weight based on the importance of different features and then summed; or using logical judgment, where different feature values ​​represent different types of information, and the final value to be filled into the voxel is determined according to certain rules. This method can reasonably handle the situation where multiple image feature data are at the same voxel location, avoid information conflicts, and ensure that the information carried by each voxel in the second voxel volume accurately and completely reflects the multiple features in the image, providing a comprehensive data foundation for generating accurate elevation feature maps.

[0096] Analysis of beneficial effects: Accurate detail preservation: This direct filling method for single image feature data accurately preserves the original information of each independent feature, enabling the second voxel to precisely reflect the position and attributes of each detailed feature in the image on the ground. This is crucial for accurately depicting subtle changes and special features on the ground, such as cracks and potholes. This detailed information is essential for vehicles to assess driving safety, helping them make evasive decisions in advance and ensuring driving safety.

[0097] Comprehensive Information Integration: The fusion and filling method for multiple image feature data can effectively integrate various feature information at the same location, avoiding information omissions or conflicts. Through a reasonable fusion method, the second voxel can comprehensively reflect the combined features of complex regions in the image, such as areas that simultaneously contain obstacle and ground texture information. This comprehensive information integration provides strong support for generating more accurate and detailed elevation feature maps, enabling vehicles to have a more comprehensive understanding of the driving environment and make more reasonable control decisions, such as accurately planning driving paths in complex road conditions.

[0098] Optionally, fusing the first voxel and the second voxel to determine the elevation feature map includes: The second voxel is input into the pre-trained first model to obtain the first elevation feature vector of the output. The elements in the first elevation feature vector are fused with the values ​​in the corresponding voxels in the first voxel to obtain the elevation values ​​of each voxel in the first voxel. Elevation values ​​corresponding to voxels with the same horizontal and vertical coordinates in the first voxel are merged to obtain elevation data, and the various elevation data are combined to obtain the elevation feature map.

[0099] In this embodiment, fusing the first and second voxels to determine the elevation feature map is a crucial step in generating accurate information reflecting ground conditions. By inputting the second voxel into a pre-trained first model to obtain elevation feature vectors, and fusing them with the corresponding voxel values ​​from the first voxel, the final elevation feature map is obtained. This series of operations fully utilizes the information advantages of the two voxels to generate a feature map containing precise elevation data. Accurate elevation feature maps provide crucial ground information for vehicles, helping them make reasonable control decisions and ensuring safe and smooth driving. Optionally, the first model is used to extract height or elevation-related features from each voxel in the second voxel.

[0100] The second voxel is input into a pre-trained first model. This model, trained on a large amount of data, is able to extract key height-related features from the image feature data and ground parameter fusion information contained in the second voxel, and output them as a first elevation feature vector. For example, the model may identify feature patterns in the second voxel that represent ground protrusions or depressions and convert them into elements in a vector that reflect height-related information at different locations.

[0101] The elements in the first elevation feature vector are fused with the values ​​of the corresponding voxels in the first voxel volume. In the first voxel volume, each voxel has its initially set ground-related parameter values, while the elements in the first elevation feature vector carry height information based on image feature data. Using a specific fusion algorithm, such as weighted fusion, different weights are assigned to the importance of determining the true elevation based on the first voxel volume parameters and the elevation feature vector elements. The values ​​of both are then combined to obtain the elevation values ​​of each voxel in the first voxel volume. For example, for a given voxel, its initial value in the first voxel volume might represent the average height of the ground at that location, while the corresponding element in the elevation feature vector might reflect height variations caused by objects in the image. The fusion of these two values ​​yields a more accurate elevation value for that voxel's location.

[0102] The elevation values ​​corresponding to voxels with the same horizontal and vertical coordinates in the first voxel volume are merged to obtain the first elevation data. This means that on a two-dimensional plane, the elevation values ​​of voxels at different height layers at the same location are integrated to obtain the final elevation data for that location. Then, the various first elevation data are combined and generated according to certain rules (such as arranging them in a matrix). In this elevation map, each pixel location corresponds to the actual ground location, and its pixel value represents the ground elevation at that location, intuitively showing the height of ground bulges or the depth of depressions, providing vehicles with clear information about ground conditions.

[0103] Analysis of beneficial effects: Precise Elevation Information Extraction: Utilizing a pre-trained first model, precise elevation features can be effectively extracted from complex voxel data. Fusing these features with the first voxel parameters further improves the accuracy of elevation values. This allows the generated elevation feature map to accurately reflect the actual undulations of the ground, providing vehicles with reliable ground information and helping them accurately determine the safety of their driving paths.

[0104] Advantages of Information Fusion: By fusing information from the first and second voxels, the advantages of both are fully utilized. The first voxel provides a basic parameter framework for the ground, while the second voxel incorporates rich image feature data. The resulting elevation feature map contains more comprehensive and accurate ground information. Compared to using either voxel alone, it can more realistically reflect the actual ground conditions and enhance the vehicle's environmental perception capabilities.

[0105] Intuitive Decision Support: The generated elevation feature map displays ground elevation information in an intuitive graphical form, allowing the vehicle control system to make decisions directly based on this information. For example, by analyzing the elevation feature map, the vehicle can quickly identify high ground, low ground, or obstacles ahead, thereby adjusting its speed and direction in a timely manner, improving the efficiency and accuracy of decision-making, and ensuring vehicle driving safety.

[0106] Optionally, controlling the vehicle to execute corresponding instructions based on the elevation feature map includes at least one of the following: Control the vehicle to execute vehicle avoidance control commands; Control the vehicle to execute a collision risk warning command; Control the vehicle to execute the door locking command to prevent opening.

[0107] In this embodiment, controlling the vehicle to execute corresponding commands based on elevation feature maps is a key step in achieving safe and intelligent vehicle driving. Different ground conditions are reflected in the elevation feature maps, and the vehicle needs to take corresponding measures based on this information to ensure driving safety, avoid collision risks, and deal with special situations. Vehicle avoidance control commands, collision risk warning commands, and door locking and opening prohibition commands are all important decisions made by the vehicle based on elevation feature maps, which help improve the vehicle's ability to cope with complex environments.

[0108] Vehicle Escape Control Command: When the elevation feature map indicates an obstacle or ground condition that may affect the vehicle's normal driving, such as a large bump, depression, or other vehicles or pedestrians, the vehicle control system will trigger a vehicle escape control command. Based on the obstacle's position and height on the elevation feature map and the vehicle's current driving status (such as speed and direction), the system calculates an appropriate escape path and maneuver. This may include adjusting the driving direction, decelerating, or accelerating to ensure the vehicle safely avoids the obstacle and continues to drive smoothly. For example, if the elevation feature map shows a large bump on the right ahead, the vehicle may slightly adjust its direction to the left and decelerate appropriately to avoid colliding with the bump and causing damage or loss of control.

[0109] Collision Risk Warning Command: Information from the elevation feature map is also used to assess the risk of a collision between the vehicle and surrounding objects. If the distance between the vehicle and an obstacle is detected to be close to a danger threshold, or if ground conditions may cause the vehicle to lose control and lead to a collision, the vehicle will execute a collision risk warning command. Warning methods may include displaying warning information on the vehicle's in-vehicle display, such as a flashing red icon or text indicating "Collision risk ahead"; simultaneously, an audible alarm may be emitted through the speakers to attract the attention of the driver or passengers. Furthermore, some advanced vehicle systems may send warning information to the driver's mobile device, ensuring that the driver receives timely alerts in any situation and can take preventative measures to avoid a collision.

[0110] Door Lock and Prevent Opening Command: In certain special circumstances, such as when the elevation feature map indicates that the vehicle is in a dangerous area, like near a cliff edge, stuck in a deep pit, or on unstable terrain, the vehicle will execute a door lock and prevent opening command to ensure the safety of the occupants. This command automatically locks the doors to prevent occupants from accidentally opening them, which could lead to them falling or other dangers. Simultaneously, the vehicle control system continuously monitors the elevation feature map, and will only allow the doors to unlock and resume normal opening and closing functions once the danger has passed.

[0111] Analysis of beneficial effects: Ensuring driving safety: Vehicle avoidance control commands enable vehicles to avoid obstacles and poor road conditions in a timely manner, effectively reducing the probability of collisions and ensuring the safety of the vehicle and its occupants. Whether encountering pedestrians or vehicles suddenly appearing on city roads or facing complex terrain in the countryside, vehicles can ensure driving safety through accurate avoidance maneuvers.

[0112] Early Risk Warning: Collision risk warning commands provide drivers with an opportunity to anticipate potential hazards. Timely alertness allows drivers to be more vigilant and take preventative measures such as braking or swerving in advance, thus avoiding collisions. This warning mechanism is not only applicable to autonomous vehicles but also crucial for human-driven vehicles, assisting drivers in better assessing road conditions and improving driving safety.

[0113] Special Situation Protection: The "Doors Lock and Do Not Open" instruction provides additional safety for occupants when the vehicle is in a dangerous area. It prevents accidental opening of the doors, avoiding injury in dangerous situations, especially in extreme circumstances where the vehicle may roll over or tilt. This instruction effectively protects the lives of those inside the vehicle.

[0114] Optionally, the method further includes: The vehicle's behavior is displayed through an in-vehicle display device to inform the user of the instructions executed by the vehicle; The vehicle's actions are played through a speaker to inform the user of the commands executed by the vehicle; The information corresponding to the vehicle's behavior is sent to the target device to inform the user of the instructions executed by the vehicle.

[0115] In this embodiment, it is crucial to promptly provide feedback on the vehicle's behavior to the user after the vehicle executes relevant commands. Through various methods, such as displaying information on the vehicle's in-vehicle display device, playing it through speakers, and sending information to target devices, the user can be ensured to have a comprehensive and timely understanding of the vehicle's operating status and executed commands. This enhances the user's trust in the vehicle's intelligent control system and also facilitates further actions or responses when necessary.

[0116] Internal display devices: Vehicles are equipped with displays such as a central control screen or a head-up display (HUD). When the vehicle executes a command, relevant information is displayed to the user in an intuitive graphical, textual, or animated format. For example, if the vehicle executes an obstacle avoidance control command, the display may show the vehicle's current trajectory changes and the location information of obstacles in the surrounding environment, accompanied by the text prompt "Vehicle has avoided the obstacle; driving is currently safe." This visual display allows users to clearly understand the vehicle's behavior and the surrounding environment, enhancing their perception of the vehicle's driving process.

[0117] Speaker Playback: The vehicle's speaker system can inform the user of vehicle commands via voice. For example, when the vehicle issues a collision risk warning command, the speaker will play a voice prompt, "Collision risk ahead, please remain vigilant"; when the vehicle issues a door locking command, the speaker will announce, "Vehicle is in a danger zone, doors are locked, do not attempt to open." Voice prompts are timely and direct, attracting the user's attention, especially when the user cannot immediately check the display device, ensuring that the user receives important information.

[0118] Sending information to target devices: In addition to in-vehicle feedback, the vehicle can also send information related to executed commands to target devices, such as the user's mobile phone or smartwatch. This requires the vehicle to establish a connection with the target device, typically via Bluetooth, Wi-Fi, or mobile network. After the vehicle executes the command, the relevant information is sent to the target device via SMS, push notification, or a specific application message. For example, if a user is away from the vehicle but remains connected via their mobile phone, when the vehicle executes a command to lock the doors, the phone will receive a push notification informing the user that "Your vehicle is in a danger zone and the doors have automatically locked." This method allows users to stay informed about important changes in the vehicle's status even when they are not inside the vehicle.

[0119] Analysis of beneficial effects: Enhanced User Perception: Through multiple feedback methods, users can gain a comprehensive understanding of the vehicle's behavior and operating status from different perspectives, enhancing their perception and understanding of the vehicle's intelligent control system. Visual displays, voice prompts, and remote information push notifications enable users to obtain key vehicle information in a timely manner, whether inside or outside the vehicle, improving the user-vehicle interaction experience.

[0120] Enhancing Trust: Timely and accurate feedback allows users to perceive the reliability and transparency of the vehicle's intelligent control system, strengthening their trust in the vehicle's automatic control functions. When users clearly understand why the vehicle executes a certain command and its current status, they will use the vehicle's intelligent functions with greater confidence, promoting the adoption and application of intelligent driving technology.

[0121] Facilitating user operation: By receiving information about the vehicle's execution commands, users can take further actions or countermeasures when necessary. For example, upon receiving a collision risk warning, the driver can drive more cautiously or manually assist the vehicle in obstacle avoidance maneuvers; when receiving information from outside the vehicle that the vehicle is in a dangerous area, the user can promptly contact relevant rescue personnel to ensure the safety of the vehicle and its occupants.

[0122] Figure 2 This is a schematic diagram of an autonomous driving scenario provided in this embodiment, such as... Figure 2 As shown, the vehicle is traveling on a city road with numerous potholes and open manholes. The road surface is also uneven due to the impact of passing heavy vehicles. The vehicle is in autonomous driving mode, designed to provide a comfortable and safe driving experience for passengers.

[0123] During operation, cameras around the vehicle continuously collect image data of the surrounding environment, covering the road ahead, the roadside on both sides, and the scene behind the vehicle. This image data is transmitted to the vehicle's image processing unit, where feature extraction is performed to obtain image feature data. Simultaneously, the 3D point cloud data acquired by the LiDAR is processed and converted into a coordinate system corresponding to the image data for subsequent fusion.

[0124] Image feature data is fused with a first voxel. The parameters of the first voxel are set based on the vehicle's initial perception of the ground, such as voxel size and height range. Using a 3D mesh sampling function determined by the feature matrix, image feature data is filled into each voxel of the first voxel to obtain a second voxel. In this process, data representing features such as potholes and manhole covers in the image are accurately mapped to the corresponding positions in the first voxel.

[0125] The first and second voxels are fused. A pre-trained height feature extraction model is used to extract height features from the second voxel, outputting them as a first elevation feature vector. Elements of this vector are then fused with the corresponding voxel values ​​in the first voxel to obtain the elevation values ​​of each voxel in the first voxel. Elevation values ​​corresponding to voxels with the same horizontal and vertical coordinates are merged to generate an elevation feature map. This map accurately displays the height of road surface bulges and the depth of depressions, including the location and depth information of potholes and manhole covers.

[0126] The vehicle's control system analyzes the elevation feature map in real time. When it detects minor undulations in the road surface, the system adjusts the chassis suspension in real time based on the degree and frequency of these undulations to provide a more comfortable driving experience. For example, if the elevation feature map shows a series of small bumps on the road ahead, similar to speed bumps, the vehicle control system will automatically adjust the damping and stiffness of the chassis suspension, softening the suspension to cushion vibrations when the vehicle passes over them and reduce the bumps felt by passengers inside the vehicle.

[0127] For areas with significant undulations or uneven surfaces, the system plans the vehicle's speed and posture in advance. If the elevation feature map indicates a large pothole ahead, the vehicle will reduce its speed in advance and adjust the suspension height to ensure sufficient ground clearance when approaching the pothole, preventing scraping of the chassis. Simultaneously, the suspension system will adjust damping force to ensure the vehicle quickly regains a stable ride after traversing the pothole.

[0128] When the elevation feature map clearly shows a pothole or open manhole ahead, the vehicle's avoidance system is immediately activated. The system first calculates the optimal avoidance path based on the location and size of the pothole or open manhole on the elevation feature map, as well as the vehicle's current speed and direction.

[0129] If a pothole or open manhole is located directly in front of the vehicle and the surrounding space allows, the vehicle will smoothly swerve to one side while slowing down appropriately to ensure a safe and stable avoidance process. During the avoidance process, the vehicle will monitor the surrounding environment in real time, including the dynamics of other vehicles and pedestrians. It will obtain relevant information through sensors around the vehicle and combine it with elevation feature maps to continuously adjust the avoidance path to avoid collisions with other objects.

[0130] For example, when a large pothole is detected on the right side ahead, the vehicle will slightly adjust its direction to the left while reducing its speed to a safe range. During the avoidance maneuver, if the vehicle detects a motorcycle approaching from the left rear using sensors, it will again slightly adjust its avoidance path to ensure that it avoids the pothole without affecting the motorcycle's movement. Figure 3 This is a schematic diagram of an automatic parking scenario provided in this embodiment, such as... Figure 3 As shown, a vehicle enters a parking lot containing special parking spaces located on the curb, which are not flush with the road surface. When the vehicle activates its automatic parking function, it needs to determine whether these curb-side spaces are available for parking and accurately estimate the curb height to ensure a safe and accurate parking operation.

[0131] The vehicle's camera captures images of the parking lot from all angles, focusing on parking spaces along the curb. The image data is processed by a feature extraction module using a multi-scale feature extraction method. Different sized convolutional kernels are used to extract features such as parking space boundaries and curb contours from the image, obtaining multiple primary feature data. These primary feature data are then scaled to the same size and fused through methods such as stitching or weighted summation to obtain the final image feature data.

[0132] The 3D mesh sampling function is determined based on the feature matrix corresponding to the image feature data. The image feature data is then filled into the first voxel to generate the second voxel. The parameters of the first voxel are set according to the approximate conditions of the parking lot surface and curb; for example, the voxel size can accurately capture the height changes of the curb.

[0133] By fusing the first and second voxels, a first elevation feature vector is obtained using a pre-trained height feature extraction model. This vector is then fused with the corresponding voxel values ​​of the first voxel to generate an elevation feature map. This elevation feature map clearly displays the ground elevation information of the parking area, especially the height and location of the curb.

[0134] The vehicle control system analyzes the elevation feature map to first determine whether the curb height is within the range that the vehicle can safely cross. Different types of vehicles have different parameters such as minimum ground clearance and suspension travel. The system compares these parameters of the vehicle itself with the curb height displayed on the elevation feature map.

[0135] For example, for a typical sedan, the minimum ground clearance is 15 centimeters. If the elevation feature map shows a curb height of 10 centimeters, and the vehicle's suspension system can remain stable when crossing the curb, the system initially determines that the parking space is suitable for parking. However, the system will further consider other factors, such as whether the length and width of the parking space are sufficient for the vehicle to park, and whether there are other obstacles in the surrounding area that may affect the parking operation.

[0136] The system combines data measured by ultrasonic sensors to accurately calculate the actual available space in the parking space. If the length of the parking space is only 20 centimeters longer than the length of the vehicle, and the width is close to the width of the vehicle, while the curb is high enough that the vehicle may be scratched during parking, the system will determine that the parking space is unsuitable and prompt the driver to find another suitable parking space.

[0137] If the parking space is determined to be available, the system will accurately estimate the height of the curb to control the vehicle's speed and suspension during parking. The system will further analyze the elevation feature map and perform detailed calculations on the voxel data of the curb area to obtain the precise height value of the curb.

[0138] During parking, the vehicle adjusts its speed according to the curb height. As the vehicle approaches the curb, it gradually reduces its speed to ensure a smooth crossing. Simultaneously, the vehicle's suspension system adjusts appropriately based on the curb height, such as increasing suspension stiffness, to prevent the chassis from colliding with the curb when crossing it.

[0139] For example, when the curb height is accurately estimated to be 12 centimeters, the vehicle reduces its speed to 5 kilometers per hour when it is 50 centimeters away from the curb and slowly drives towards it. When crossing the curb, the suspension system automatically adjusts to increase stiffness, allowing the vehicle to smoothly enter the parking space and complete the parking operation.

[0140] In one possible embodiment, the three-dimensional mesh sampling function is:

[0141] These are the coordinates in the vehicle coordinate system. Let be the coordinates of the voxels in the first voxel volume, and let be the intrinsic parameter matrix of the camera. The rotation and translation matrices of the extrinsic parameters are respectively and The three-dimensional mesh sampling function can map the pixel values ​​in the image feature data to the corresponding vehicle coordinate system in the voxel.

[0142] Based on the above embodiments, Figure 4 A schematic diagram of another vehicle control system provided in this application embodiment is shown below. Figure 4 As shown, the entire system includes: an image feature data extraction module, a coarse estimation module for height estimation, and a fine estimation module for height estimation. The system operation flow is as follows: The original image is mapped to a feature space to obtain a 2D feature map. This module can simultaneously support monocular, monocular temporal, and stereo inputs. When the input is a monocular temporal or stereo camera, the image backbone network will extract multiple image feature data. When the input is a monocular camera, the input data consists of only a single frame image. When the input is a monocular time series or a stereo camera, the input typically consists of two frames of images, namely... and All input images are processed by a backbone network to extract multi-scale features (typically the smallest three scales). These multi-scale features are then scaled to a specified size and concatenated along the feature channel dimension to obtain the final multi-scale image feature data. .

[0143] The coarse height estimation module is used to estimate the initial height. First, based on the zero-height plane, that is The plane, within a fixed XY range, initializes an initial voxel. That is, the first voxel in the first round of iterations. A relatively large height range can be adopted to accommodate greater variations in ground elevation. The coarsely estimated voxel grid is set by the sensing range, with a lateral range defined. Longitudinal range Height range 0, then when the input camera is a pinhole camera model, its camera intrinsic parameter matrix is The rotation and translation matrices of the extrinsic parameters are respectively and From this, we can obtain the formula for sampling voxels from the vehicle coordinate system to the image space, which is also the 3D mesh sampling function as follows:

[0144] In the above formula .right Uniformly discrete sampling is performed on the value space, with different sampling intervals. (i.e., sampling resolution) can control the voxel grid. The size of the z-values. For example, if the height is 0 and the range is [-2, 2], sampling with a step size of 1 will yield a set of z-values ​​such as [-2, -1, 0, 1, 2]. This is what we call the sampled output. Generally, the height of a road surface will be distributed in... Therefore, the height sampling range is set to a range of [missing information]. The symmetrical interval centered on this. Furthermore, in this coarse height estimation stage, generally... The value range will be set relatively large to avoid insufficient height range, which could lead to the omission of higher or lower features.

[0145] After obtaining the initial voxel, it is then sampled using a 3D mesh function. From multi-scale image feature data Sampled BEV 3D voxel That is, the second body in the first round of iterations, further estimates the preliminary height through the height Transformer (2 in series) and the coarse regression head module. This refers to the elevation feature map during the first iteration. For monocular or binocular inputs, only voxel volumes are considered. The calculation methods are slightly different.

[0146] The height estimation module is primarily used to fine-tune the estimation results from the coarse estimation stage, achieving higher height estimation accuracy. It is based on the initial voxel volume from the previous iteration. The obtained voxel features were used to estimate the preliminary height results. Then according to It can be further optimized. The sampling points along the height direction yield the updated fine-tuned voxel volume during the second round of iterations. This refers to the first body in the second round of iterations.

[0147] for In other words, The value space and the "coarse" estimation stage Same, but The range of values ​​for will be smaller, that is... , Roughly estimated height value The error will not be too large, so it can generally be taken as... The sampling interval can also be taken as It can significantly improve The sampling resolution in the height space is improved, thus significantly enhancing the accuracy of height estimation. This results in a finely tuned voxel volume. Then, similar to the previous iteration, a 3D mesh sampling function is used. From multi-scale image feature data Sample the BEV 3D voxel during the second round of iterations. (That is, the second body in the second round of iteration), and further estimate the fine-tuned height through another height Transformer (2 in series) and the fine regression head module. This refers to the elevation feature map in the second round of iterations.

[0148] This embodiment demonstrates the process of iteratively optimizing the first body in two rounds to obtain an elevation feature map. In actual implementation, multiple rounds of iterative optimization of the first body can be performed to obtain a more accurate elevation feature map.

[0149] Optionally, the system in this embodiment has low computational resource requirements and can run in real time on edge devices (such as the vehicle's infotainment system). While ensuring high-precision estimation, it achieves higher operating efficiency and consumes fewer computational resources.

[0150] Optionally, the system in this embodiment can also be deployed in the cloud. The vehicle uploads the captured image data to the cloud during driving, and the system deployed in the cloud processes the data to obtain elevation data.

[0151] The system in this embodiment exhibits strong robustness and can handle scenarios with large variations in road surface height. This invention proposes a two-stage height estimation method, from coarse to fine. In the coarse estimation stage, although the estimation accuracy is not high, it can effectively handle large-scale road surface height variations. In the fine estimation stage, the height estimation results from the coarse stage are further refined, significantly improving the final height estimation accuracy. Furthermore, it is compatible with a wider range of hardware and has greater potential for practical applications. It can simultaneously support monocular, monocular temporal, and binocular image input. It also has stronger compatibility with hardware configurations of different cameras.

[0152] To achieve the above embodiments, this application also proposes a vehicle control device.

[0153] Figure 5 This is a schematic diagram of the structure of a vehicle control device provided in an embodiment of this application.

[0154] like Figure 5 As shown, the device may include: The feature extraction module 510 is used to collect image data around the vehicle and extract features from the image data to obtain image feature data. The elevation determination module 520 is used to fuse the image feature data into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of the ground protrusion or the depth of the depression; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. The control module 530 is used to control the vehicle to execute corresponding instructions based on the elevation feature map.

[0155] Optionally, the elevation determination module includes: A sampling submodule is used to sample the image feature data and incorporate it into the first voxel to obtain a second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

[0156] Optionally, the device further includes: An iteration module is used to sample the image feature data multiple times according to a preset number of iterations and sequentially fuse it into the first voxel of the current round to iteratively update the elevation feature map; wherein the first voxel of the current round is constructed based on the elevation feature map of the previous round.

[0157] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.

[0158] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.

[0159] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.

[0160] To implement the above embodiments, this application also proposes a vehicle, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.

[0161] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example, the electronic device 800 can be a mobile phone, a vehicle, an in-vehicle system, etc.

[0162] Reference Figure 6 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.

[0163] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0164] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0165] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0166] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display and a touch panel. If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0167] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0168] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0169] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0170] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0171] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0172] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0173] To implement the above embodiments, this application also proposes a chip, including: the chip includes a processing circuit configured to perform the methods provided in the foregoing embodiments.

[0174] Figure 7 This is a schematic diagram of the structure of a chip according to an embodiment of this application. See also... Figure 7 The diagram shown is a schematic representation of the structure of chip 1100, but it is not limited to this.

[0175] Chip 1100 includes processing circuitry 1101, which is configured to perform any of the above methods.

[0176] In some embodiments, chip 1100 further includes one or more interface circuits 1102. Optionally, the interface circuit 1102 is connected to memory 1103, and the interface circuit 1102 can be used to receive signals from memory 1103 or other devices, and the interface circuit 1102 can be used to send signals to memory 1103 or other devices. For example, the interface circuit 1102 can read instructions stored in memory 1103 and send the instructions to processing circuit 1101.

[0177] In some embodiments, the interface circuit 1102 performs at least one of the communication steps such as sending and / or receiving in the above method, while the processing circuit 1101 performs other steps.

[0178] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.

[0179] In some embodiments, chip 1100 further includes one or more memories 1103 for storing instructions. Optionally, all or part of the memories 1103 may be located outside of chip 1100.

[0180] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0181] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0182] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0183] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0184] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0185] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0186] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0187] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A vehicle control method, characterized in that, include: Collect image data of the area surrounding the vehicle, and extract features from the image data to obtain image feature data; The image feature data is fused into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of ground protrusions or the depth of ground depressions; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. According to a preset number of iterations, the image feature data is sampled multiple times and sequentially fused into the first voxel of the current round to iteratively update the elevation feature map; wherein, the first voxel of the current round is constructed based on the elevation feature map of the previous round; The vehicle is controlled to execute corresponding commands based on the updated elevation feature map. The process of repeatedly sampling the image feature data and sequentially fusing it into the first voxel of the current round to iteratively update the elevation feature map includes: The values ​​of each voxel in the reference plane in the updated first voxel volume are determined based on the elevation feature map; wherein the height range and height sampling interval in the updated first voxel volume are smaller than the height range and height sampling interval of the first voxel volume before the update.

2. The method according to claim 1, characterized in that, The step of fusing the image feature data into a first voxel to obtain an elevation feature map includes: The image feature data is sampled and incorporated into the first voxel to obtain the second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

3. The method according to claim 1, characterized in that, The image data includes single-frame image data or multi-frame image data, and the step of extracting features from the image data to obtain image feature data includes: The image data of one or more frames is input into the feature extraction module to perform multi-scale feature extraction to obtain multiple first feature data. Scale each of the first feature data to the same size to obtain the second feature data; The image feature data is obtained by fusing the various second feature data.

4. The method according to claim 3, characterized in that, The step of sampling the image feature data and incorporating it into the first voxel to obtain the second voxel includes: The three-dimensional mesh sampling function is determined based on the feature matrix corresponding to the image feature data; The image feature data is filled into each voxel of the first voxel according to the three-dimensional network sampling function to obtain the second voxel.

5. The method according to claim 4, characterized in that, The step of filling the image feature data into each voxel of the first voxel body according to the three-dimensional network sampling function includes any one of the following: For a single image feature data, the voxel corresponding to the feature value in the image feature data in the first voxel volume is determined according to the three-dimensional network sampling function, and the feature value is filled into the corresponding voxel. For multiple image feature data, the feature value in each image feature data is determined in the first voxel according to the three-dimensional network sampling function, and the feature value corresponding to the same voxel is fused and filled into the corresponding voxel.

6. The method according to claim 1, characterized in that, The process of fusing the first voxel and the second voxel to determine the elevation feature map includes: The second voxel is input into the pre-trained first model to obtain the first elevation feature vector of the output. The elements in the first elevation feature vector are fused with the values ​​in the corresponding voxels in the first voxel to obtain the elevation values ​​of each voxel in the first voxel. Elevation values ​​corresponding to voxels with the same horizontal and vertical coordinates in the first voxel are merged to obtain elevation data, and the various elevation data are combined to obtain the elevation feature map.

7. The method according to claim 1, characterized in that, The control of the vehicle to execute corresponding instructions based on the elevation feature map includes at least one of the following: Control the vehicle to execute vehicle avoidance control commands; Control the vehicle to execute a collision risk warning command; Control the vehicle to execute the door locking command to prevent opening.

8. The method according to any one of claims 1-7, characterized in that, The method further includes at least one of the following: The vehicle's behavior is displayed through an in-vehicle display device to inform the user of the instructions executed by the vehicle; The vehicle's actions are played through a speaker to inform the user of the commands executed by the vehicle; The information corresponding to the vehicle's behavior is sent to the target device to inform the user of the instructions executed by the vehicle.

9. A vehicle control device, characterized in that, include: The feature extraction module is used to collect image data around the vehicle and extract features from the image data to obtain image feature data; An elevation determination module is used to fuse the image feature data into a first voxel to obtain an elevation feature map, wherein the elevation feature map is used to characterize the height of ground protrusions or the depth of ground depressions; wherein the parameters in the first voxel are the parameters of the ground perceived by the vehicle. An iteration module is used to sample the image feature data multiple times according to a preset number of iterations and sequentially fuse it into the first voxel of the current round to iteratively update the elevation feature map; wherein the first voxel of the current round is constructed based on the elevation feature map of the previous round; The control module is used to control the vehicle to execute corresponding commands based on the updated elevation feature map; The iteration module is also used for: The values ​​of each voxel in the reference plane in the updated first voxel volume are determined based on the elevation feature map; wherein the height range and height sampling interval in the updated first voxel volume are smaller than the height range and height sampling interval of the first voxel volume before the update.

10. The apparatus according to claim 9, characterized in that, The elevation determination module includes: A sampling submodule is used to sample the image feature data and incorporate it into the first voxel to obtain a second voxel; The first voxel and the second voxel are fused to determine the elevation feature map, which includes the elevation data of each pixel in the image data.

11. A vehicle, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1-8.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of the preceding claims 1-8.

13. A chip, characterized in that, The chip includes processing circuitry configured to perform the method as described in any one of claims 1-8.

14. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • CN118898838A

  • CN120388230A