Stereovision system and method of application thereof

By installing a stereo vision system with multi-view stereo cameras on vehicles, the high cost and low efficiency problems caused by multi-sensor fusion are solved, enabling high-precision intelligent driving and parking functions while reducing computing resources and assembly costs.

CN120735760BActive Publication Date: 2026-01-23BEIJING SMARTER EYE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511254704.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-01-23
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In existing advanced intelligent driving solutions, multi-sensor fusion leads to a significant increase in data redundancy, high consumption of computing resources, low data integration efficiency, and high assembly and maintenance costs. Furthermore, the perception fusion between sensors exhibits a game-like phenomenon, resulting in omissions or false triggers during the execution phase.

Method used

A stereo vision system is adopted, which uses multi-view stereo cameras at different positions on the vehicle body, including front-view, rear-view and surround-view vision units. AI models are used to process distorted images and point cloud data to achieve all-round perception, and data fusion and filtering are performed to reduce computing power requirements and improve overall data efficiency.

Benefits of technology

It achieves comprehensive and high-precision perception of the vehicle's surroundings, reduces computing resources and assembly and maintenance costs, improves data integration efficiency, and supports L2 and above intelligent driving, intelligent parking, and intelligent chassis functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120735760B_ABST
    Figure CN120735760B_ABST
Patent Text Reader

Abstract

The application discloses a kind of stereoscopic vision systems and its application method, the stereoscopic vision system includes: the front vision unit of being arranged in the front position of vehicle body, including: one or more long-focus cameras, one or more wide-angle cameras, for detecting the object in the predetermined range of vehicle front view, obtaining horizontal and longitudinal distance measurement and relative speed measurement, and realizing real-time modeling to the road surface in front of vehicle;Rear vision unit is arranged in the rear position of vehicle body, including: one or more wide-angle cameras, for detecting the object in the predetermined range of vehicle rear view, and realizing real-time modeling to the road surface in rear of vehicle;Surrounding vision unit is arranged in the front and rear position of vehicle body, left and right position, including: multiple groups of multi-view cameras, for detecting the object in the predetermined range of vehicle surrounding view, and realizing local modeling.According to the above technical scheme provided by the application, the multi-view stereo camera is arranged at different positions on the vehicle body with visual sensor as the main sensor, so that the comprehensive efficiency of data is improved, the assembly cost and maintenance cost are reduced, and the comprehensive efficiency of data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent driving, in particular to a stereoscopic vision system and an application method thereof. BACKGROUND

[0002] With the rapid development of intelligent driving technology, the accurate perception and decision-making capability of intelligent driving systems have become the core requirement for realizing safe and stable driving. High-order intelligent driving generally refers to L2 and above (referring to GB / T 40429-2021 Automobile Driving Automation Classification) functions, which can be simply summarized as realizing joint control of horizontal and vertical directions, supporting automatic driving and automatic parking, and adapting to various common actual road conditions.

[0003] The existing high-order intelligent driving scheme mainly uses multi-sensor fusion, including camera, millimeter wave radar, laser radar, ultrasonic wave, etc. These sensors provide a large amount of data redundancy, which leads to a substantial increase in the computing resources of the solution, low comprehensive efficiency of data, and high assembly and maintenance costs of these sensors. In addition, since these sensors are sensors with different working principles, there will be inevitable game phenomena in the perception, fusion, and control stages, which ultimately leads to omissions or false touches in the execution stage. Therefore, the intelligent driving application scene inevitably faces difficult scenarios that cannot be solved. SUMMARY

[0004] The main purpose of the present application is to disclose a stereoscopic vision system and an application method thereof, which at least solves the problem that the high-order intelligent driving scheme in the related art mainly uses multi-sensor fusion, including camera, millimeter wave radar, laser radar, ultrasonic wave, etc. The above-mentioned sensors provide a large amount of data redundancy, which leads to a substantial increase in the computing resources of the solution, low comprehensive efficiency of data, and high assembly and maintenance costs of these sensors.

[0005] According to one aspect of the present application, a stereoscopic vision system is provided.

[0006] The stereoscopic vision system provided by the present application comprises: a front-view vision unit arranged at a front position of a vehicle body, comprising one or more long-focus cameras and one or more wide-angle cameras, for detecting objects within a predetermined range in front of the vehicle, acquiring horizontal and vertical distance measurements and relative speed measurements, and realizing real-time modeling of the road surface in front of the vehicle; a rear-view vision unit arranged at a rear position of the vehicle body, comprising one or more wide-angle cameras, for detecting objects within a predetermined range behind the vehicle, and realizing real-time modeling of the road surface behind the vehicle; and a surround-view vision unit arranged at front and rear positions and left and right positions of the vehicle body, comprising multiple groups of multi-view cameras, for detecting objects within a predetermined range around the vehicle, and realizing local modeling.

[0007] Further, the front vision unit is arranged in a black area in the middle upper part of the windshield in the cockpit, the rear vision unit is arranged in the shark fin part of the rear of the vehicle or the upper part outside the rear windshield, and the surround view vision unit is arranged in the upper side of the front and rear license plates of the vehicle and the vicinity of the left and right fenders or rearview mirrors.

[0008] Further, the optical axis of the camera in the front vision unit and the optical axis of the camera in the rear vision unit are both horizontal to the ground, in the XOY plane, the optical axis of the camera arranged in the front and rear positions of the vehicle body is parallel to the center line of the vehicle body, and the optical axis of the camera arranged in the left and right positions of the vehicle body is perpendicular to the center line of the vehicle body; in the YOZ plane, the optical axis of the camera arranged in the left and right positions of the vehicle body is inclined to the ground direction so that the optical axis point falls on the ground near the vehicle body, and in the XOZ plane, the optical axis of the camera arranged in the front and rear positions of the vehicle body is inclined to the ground direction so that the optical axis point falls on the ground near the vehicle body.

[0009] Further, the front vision unit comprises one long-focus camera arranged in the center and wide-angle cameras arranged on the left and right sides of the long-focus camera, the rear vision unit comprises two wide-angle cameras, and the surround view vision unit comprises four sets of binocular cameras with the same structure parameters, wherein one set of binocular cameras is arranged in the front, rear, left and right positions of the vehicle body.

[0010] According to another aspect of the present application, an application method of the stereovision system is provided.

[0011] The application method of the stereovision system according to the present application comprises: directly transmitting distorted image data collected based on the stereovision system to an image processing related AI model for processing to obtain first processing data; transmitting point cloud data generated based on corrected image data corrected according to the distorted image data to a point cloud processing related AI model for processing to obtain second processing data; performing data pre-fusion operation on the corrected image data and the point cloud data, and transmitting the fused data to a fused data related AI model for processing to obtain third processing data; performing data post-fusion operation on the first processing data and the second processing data, performing data filtering operation on the fused data and the third processing data, splicing and outputting the filtered data.

[0012] Furthermore, the image processing-related AI models include: a small-scale AI model with a backbone and multiple heads, used to perform tasks related to target detection, lane detection, image segmentation, and drivable area detection; and point cloud processing-related AI models, which are invoked and applied after obtaining point cloud data, including: a small-scale AI model with a backbone and multiple heads, used to perform tasks related to point cloud segmentation, target detection, road surface detection, and drivable area detection; and data fusion-related AI models, including: a large-scale model, used to perform target perception tasks in various directions in driving and parking scenarios based on the data obtained by fusing the above-mentioned corrected image data and the above-mentioned point cloud data.

[0013] According to another aspect of the present invention, a method for applying a stereo vision system is provided.

[0014] The application method of the stereo vision system according to the present invention includes: correcting distorted image data acquired based on the stereo vision system to obtain corrected image data, and generating point cloud data based on the corrected image data; fusing the corrected image data and the point cloud data to realize three-dimensional reconstruction of the spatial environment around the vehicle, and obtaining a semantic point cloud map including vector features through spatial semantic point cloud mapping, wherein the vector features include: three-dimensional features, relative ground position features, target category, target state, and color features; when the vehicle is parked, detecting and updating the target parking space based on the corrected image data and the point cloud data, adjusting the self-tested vehicle posture according to the ground parking space marking lines, and identifying spatial obstacles according to the real-time updated semantic point cloud map.

[0015] Furthermore, when obtaining the aforementioned semantic point cloud map through spatial semantic point cloud mapping, it also includes:

[0016] The system memorizes the route the vehicle has traveled and guides the vehicle to achieve low-speed navigation based on the memorized route. When low-speed navigation is activated, the starting point is set to the parking position, and the ending point is set to the origin position at the moment the valet parking function is activated.

[0017] According to another aspect of the present invention, a method for applying a stereo vision system is provided.

[0018] The application method of the stereo vision system according to the present invention includes: correcting distorted image data acquired based on the stereo vision system to obtain corrected image data, and generating point cloud data based on the corrected image data; performing real-time modeling of the road surface ahead based on the corrected image data and the point cloud data to achieve terrain classification; identifying and measuring the location information of road surface protrusions or depressions on structured hard roads; detecting changes in road surface height and road surface condition; detecting dynamic targets around the vehicle; and dynamically monitoring dynamic targets with a collision probability higher than a predetermined threshold.

[0019] Furthermore, the detection of road surface condition changes based on the aforementioned corrected image data and point cloud data includes: using the forward vision unit of the aforementioned stereo vision system to identify and perceive the road surface condition; based on the acquired corrected image data and point cloud data, obtaining the pattern and time point of the road surface condition change ahead; pre-notifying the control mechanism to execute a response strategy; and adjusting the vehicle according to a new dynamic model when driving on a changed road surface. The detection of dynamic targets around the vehicle and dynamic monitoring of dynamic targets with a collision probability higher than a predetermined threshold includes: using the surround vision unit and rear vision unit of the aforementioned stereo vision system to detect dynamic targets around the vehicle in real time; dynamically monitoring dynamic targets with a collision probability higher than a predetermined threshold based on the acquired corrected image data and point cloud data; and when a warning standard is reached, calculating the vehicle state after the collision based on the predicted collision state and correcting the vehicle power distribution based on the vehicle state.

[0020] This invention proposes a stereo vision system and its application method. Using a vision sensor as the main sensor, multi-view stereo cameras are set at different positions on the vehicle body to achieve all-round perception of the surrounding environment of the vehicle, thereby reducing computing resources, improving the overall efficiency of data, and reducing assembly and maintenance costs. Attached Figure Description

[0021] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0022] Figure 1 This is a structural block diagram of a stereo vision system according to an embodiment of the present invention;

[0023] Figure 2 This is a layout schematic diagram of a stereo vision system according to a preferred embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the optical axes of two binocular cameras on the left and right sides of the YOZ plane tilted towards the ground direction according to a preferred embodiment of the present invention.

[0025] Figure 4 This is a schematic diagram of two binocular cameras tilted towards the ground on the XOZ plane according to a preferred embodiment of the present invention.

[0026] Figure 5 This is a schematic diagram of the overall application scheme of the stereo vision system according to an embodiment of the present invention;

[0027] Figure 6 This is a flowchart of the application method of the stereo vision system according to Embodiment 1 of the present invention;

[0028] Figure 7 This is a flowchart illustrating the application method of the stereo vision system according to a preferred embodiment of the present invention.

[0029] Figure 8 This is a flowchart of the application method of the stereo vision system according to Embodiment 2 of the present invention;

[0030] Figure 9 This is a flowchart illustrating the application method of the stereo vision system according to a preferred embodiment of the present invention.

[0031] Figure 10 This is a flowchart of the application method of the stereo vision system according to Embodiment 3 of the present invention;

[0032] Figure 11 This is a flowchart illustrating the application method of the stereo vision system according to a preferred embodiment of the present invention. Detailed Implementation

[0033] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] The specific implementation of the present invention will now be described in detail with reference to the accompanying drawings.

[0035] According to an embodiment of the present invention, a stereo vision system is provided.

[0036] Figure 1 This is a structural block diagram of a stereo vision system according to an embodiment of the present invention. Figure 1As shown, the stereo vision system includes: a forward vision unit 10 positioned at the front of the vehicle body, comprising one or more telephoto cameras and one or more wide-angle cameras, used to detect objects within a predetermined range of forward vision of the vehicle, acquire lateral and longitudinal distance measurements and relative speed measurements, and perform real-time modeling of the road surface in front of the vehicle; a rear vision unit 12 positioned at the rear of the vehicle body, comprising one or more wide-angle cameras, used to detect objects within a predetermined range of rear vision of the vehicle, and perform real-time modeling of the road surface behind the vehicle; and surround vision units 14 positioned at the front and rear, and left and right sides of the vehicle body, comprising multiple sets of multi-view cameras, used to detect objects within a predetermined range of surround vision of the vehicle, and perform local modeling.

[0037] In related technologies, advanced intelligent driving solutions mainly rely on the fusion of multiple sensors, including cameras, millimeter-wave radar, lidar, and ultrasonic sensors. These sensors provide a large amount of data redundancy, which leads to a significant increase in the computing resources required for the solution. The overall efficiency of the data is relatively low, and the assembly and maintenance costs of these sensors are also high. Figure 1 The system shown uses a vision sensor as the main sensor and multiple stereo cameras are set at different positions on the vehicle body to achieve all-round perception of the surrounding environment of the vehicle. This reduces computing resources, improves the overall efficiency of data, and reduces assembly and maintenance costs.

[0038] For example, the aforementioned forward vision unit may include, but is not limited to: a telephoto camera centrally located, and wide-angle cameras respectively located on the left and right sides of the telephoto camera; the aforementioned rear vision unit may include, but is not limited to: two wide-angle cameras; the aforementioned surround vision unit may include, but is not limited to: four sets of binocular cameras with identical structural parameters, wherein one set of the aforementioned binocular cameras is located at the front, rear, left, and right of the vehicle body.

[0039] In the preferred implementation, the aforementioned forward-looking vision unit includes one or more telephoto cameras and one or more wide-angle cameras, primarily supporting forward driving assistance functions. This means it can perceive both distant targets and lateral width, accurately detecting objects within a predetermined range (e.g., within a 250-meter distance) in front of the vehicle. This forward-looking vision unit can acquire both image semantic information and point cloud structure data; it can detect and perceive various types of obstacles, obtaining accurate lateral and longitudinal distance measurements and relative speed measurements. Simultaneously, this forward-looking vision unit can also perform real-time modeling of the road surface ahead, supporting the calibration of autonomous driving and intelligent chassis functions.

[0040] The aforementioned rear-view vision unit includes one or more wide-angle cameras, primarily supporting functions such as rear warning, automatic navigation, and automatic parking. This rear-view vision unit can provide perception and measurement results of obstacles within a predetermined range behind the vehicle (e.g., within a 100-meter distance), and can also detect suspended targets and non-standard targets, as well as perform real-time modeling of the road environment.

[0041] The aforementioned surround-view vision units are installed at the front, rear, left, and right of the vehicle, and include multiple sets of multi-view cameras. These surround-view vision units primarily support driving functions such as automatic navigation, side assist, automatic parking, and local mapping. The surround-view vision units can provide obstacle target perception within a predetermined range (e.g., within a 100-meter distance) and high-precision local mapping within a predetermined range (e.g., within a 30-meter distance).

[0042] The aforementioned forward vision unit can be located in the upper center of the windshield inside the driver's cabin; the aforementioned rear vision unit can be located at the shark fin area at the rear of the vehicle, or on the upper outer side of the rear windshield; the aforementioned surround vision unit can be located above the front and rear license plates of the vehicle, as well as near the left and right fenders or rearview mirrors.

[0043] In the preferred implementation process, in the overall vehicle layout, the aforementioned forward vision unit can be installed in the black area in the upper middle of the windshield inside the driver's cabin (usually located near the rearview mirror); the rear vision unit can be installed at the shark fin area at the rear of the vehicle, or on the upper outer side of the rear windshield; the surround vision unit can be installed above the front and rear license plates, as close as possible to the center line of the vehicle body, and in the area near the left and right fenders or rearview mirrors, generally at the same height as the surround vision binocular camera.

[0044] Specifically, the optical axes of the cameras in the aforementioned forward vision unit and the rear vision unit can both be set horizontally to the ground. In the aforementioned surround vision unit, on the XOY plane, the optical axes of the cameras positioned at the front and rear of the vehicle are parallel to the center line of the vehicle, and the optical axes of the cameras positioned on the left and right sides of the vehicle are perpendicular to the center line of the vehicle. On the YOZ plane, the optical axes of the cameras positioned on the left and right sides of the vehicle are tilted towards the ground so that the optical axis point falls on the ground near the vehicle. On the XOZ plane, the optical axes of the cameras positioned at the front and rear of the vehicle are tilted towards the ground so that the optical axis point falls on the ground near the vehicle.

[0045] In the process of optimal implementation, such as Figure 2As shown, the optical axes of the cameras in the forward-looking vision unit (e.g., a forward-looking trinocular camera, wherein a telephoto camera is centrally located, and wide-angle cameras are respectively located on the left and right sides of the telephoto camera) and the optical axes of the cameras in the rear-looking vision unit (e.g., two wide-angle cameras) are both horizontal to the ground. However, in actual implementation, a certain degree of fluctuation is allowed within the installation error range. The installation positions of the surround-view vision unit (e.g., four sets of binocular cameras with identical structural parameters, wherein one set of binocular cameras is located at the front, rear, left, and right of the vehicle body) are as follows: Figure 2 As shown, on the XOY plane, the optical axes of the binocular cameras set at the front and rear are parallel to the center line of the vehicle body, while the optical axes of the binocular cameras set on the left and right sides are perpendicular to the center line of the vehicle body. Figure 3 On the YOZ plane shown, the optical axes of each set of binocular cameras (for example, a set of binocular cameras on the left side of the vehicle body and a set of binocular cameras on the right side of the vehicle body) are tilted towards the ground to ensure that the optical axis point falls on the ground near the vehicle body. Figure 4 On the XOZ plane shown, the optical axes of the binocular cameras positioned at the front and rear (e.g., one set at the front of the vehicle and one set at the rear) are tilted towards the ground, ensuring that the optical axis falls on the ground near the vehicle. This binocular camera setup ensures that the area near the optical axis in the panoramic camera image is the primary focus for the user, while the surrounding areas, such as the vehicle image, can be cropped out.

[0046] In the aforementioned stereo vision system, stereo vision is used as the primary sensor, and multi-view stereo cameras are installed at different locations on the vehicle body, such as... Figure 5 As shown, taking forward-facing three-eye, rear-facing two-eye, and surround-view two-eye as examples, it realizes all-round perception of the vehicle's surrounding environment. The perception solution provided can support L2 and above intelligent driving solutions, intelligent parking solutions, and intelligent chassis solutions, providing more precise, robust, and low-cost intelligent perception solutions for the three intelligent directions of intelligent driving, intelligent parking, and intelligent chassis.

[0047] In intelligent driving mode, such as Figure 5 As shown, the main functions include: forward assist, lateral assist, side assist, rear assist, and autopilot. These functions activate under different conditions, with forward assist being the most basic and the foundation for activating other intelligent driving functions. Autopilot is mutually exclusive with lateral and rear assist and has the highest priority; when autopilot is activated, lateral and rear assist are suppressed.

[0048] In intelligent parking mode, such as Figure 5As shown, the main functions include: 360-degree surround view, automatic parking, memory parking, and valet parking. The 360-degree surround view is the basic function, primarily responsible for the interaction with the vehicle's infotainment system and its effect on the real-world environment, and it's also the foundation for activating other intelligent driving functions. Automatic parking, memory parking, and valet parking are three mutually exclusive functions, with their priority increasing in that order.

[0049] In intelligent chassis mode, such as Figure 5 As shown, it mainly includes: power adjustment function and suspension adjustment function. Among them, power adjustment and suspension adjustment are mutually exclusive functions, and power adjustment has the highest priority.

[0050] It should be noted that intelligent driving and intelligent parking are two mutually exclusive functions. While both can be active simultaneously, intelligent driving has the highest priority. Similarly, intelligent parking has a higher priority than intelligent chassis functionality.

[0051] According to an embodiment of the present invention, a method for applying a stereo vision system is also provided, namely, an intelligent driving scheme based on the above-mentioned stereo vision system.

[0052] Figure 6 This is a flowchart of an application method for a stereo vision system according to Embodiment 1 of the present invention. As shown in the figure, the application method of the stereo vision system includes:

[0053] Step S601: The distorted image data acquired based on the stereo vision system is directly transmitted to the image processing-related AI model for processing to obtain the first processed data;

[0054] Step S602: The point cloud data generated from the distorted image data after distortion correction is transmitted to the point cloud processing-related AI model for processing to obtain the second processed data;

[0055] Step S603: Perform a data pre-fusion operation on the above-mentioned corrected image data and the above-mentioned point cloud data, and transmit the fused data to the AI ​​model related to the fused data for processing to obtain the third processed data;

[0056] Step S604: Perform a data fusion operation on the first and second processed data, perform a data filtering operation on the fused data and the third processed data, and then concatenate and output the filtered data.

[0057] Among them, the aforementioned image processing-related AI models may include, but are not limited to: small-scale AI models with a backbone and multiple heads, used to achieve tasks related to object detection, lane detection, image segmentation, and drivable area;

[0058] Among them, the aforementioned AI models related to point cloud processing, after obtaining point cloud data, are invoked and applied, and may include, but are not limited to: small-scale AI models with a backbone and multiple heads, used to achieve tasks related to point cloud segmentation, target detection, road surface detection, and drivable area detection.

[0059] Among them, the AI ​​models related to the aforementioned fused data may include, but are not limited to: large-scale models used to achieve the task of perceiving targets in various directions in driving and parking scenarios based on the data after the aforementioned fused corrected image data and the aforementioned point cloud data.

[0060] The following description uses examples of forward trinocular (i.e., the forward vision unit includes: a trinocular camera, wherein a telephoto camera is centrally located and wide-angle cameras are respectively located on the left and right sides of the telephoto camera), rear binocular (i.e., the rear vision unit includes: a binocular camera, wherein the binocular camera consists of two wide-angle cameras), and surround binocular (i.e., the surround vision unit includes: four sets of binocular cameras with the same structural parameters respectively located at the front, rear, left, and right of the vehicle body).

[0061] Intelligent driving is mainly divided into five modules: forward assist, lateral assist, side assist, rear assist, and automatic navigation. The status requirements of the sensors are shown in Table 1.

[0062]

[0063] Among them, the forward-looking tri-camera and the rear-looking binoculars use raw data to access the controller, while the surround-view binoculars use YUV data to access the controller. In other words, the surround-view binoculars come with ISP debugging, while the forward-looking tri-camera and the rear-looking binoculars do not.

[0064] When traveling forward, the forward-looking three-lens camera outputs a telephoto image, a wide-angle image, and point cloud data, respectively. The telephoto image primarily focuses on distant targets within the current lane and adjacent lanes to the left and right; the wide-angle camera primarily focuses on targets within the lateral field of view at medium to close distances. The surround-view binocular camera primarily focuses on targets at medium to close distances around the vehicle, supplementing the blind spots near the ground near the vehicle body of the forward and rear-view stereo cameras, while also covering the front and rear sides. The rear-view binocular camera is mainly used for detecting medium to close distance targets in the rear and rear-side directions.

[0065] In a stereo vision-based intelligent driving solution, the software process is as follows: Figure 7 As shown.

[0066] Figure 7 This is a flowchart illustrating the application method of the stereoscopic vision system according to a preferred embodiment of the present invention. Figure 7 As shown, the application method of this stereo vision system includes the following process:

[0067] Based on the distorted image data acquired by the front-view trinocular, surround-view binocular, and rear-view binocular systems described above, the distorted image, the corrected image, and the point cloud data are generated sequentially from front to back. That is, the distorted image data is distorted to obtain the corrected image data, and the point cloud data is generated based on the corrected image data.

[0068] The distorted image data is directly submitted to image processing-related AI models, applied to a small-scale AI model with a backbone and multiple heads, to achieve basic small-scale model-related tasks such as target detection, lane detection, image segmentation, and drivable area detection. This small-scale AI model processing scheme boasts high operating efficiency, short system latency, and fast control and execution response speeds. It is suitable for emergency situations requiring rapid decision-making and response (such as AEB), covering common pedestrians, vehicles, and other obstacles in a limited number of scenarios. Furthermore, it can provide safety redundancy in case of system anomalies, such as for pulling over to the side of the road.

[0069] Point cloud processing-related AI models require point cloud data before they can be activated and applied to small-scale AI models with a backbone and multiple heads. This enables basic small-scale model tasks such as point cloud segmentation, object detection, road surface detection, and drivable area detection. This stage is highly efficient, but has a longer system latency compared to image-related AI models, primarily because it requires waiting for two preceding data preparation stages: image distortion correction and point cloud generation. This stage serves for data redundancy, mainly through data fusion with the preceding image-based detection results. Its advantage is that perceptual redundancy can be achieved based on different data sources, thus validating the perceptual results.

[0070] The AI ​​model based on fused data relies on the fusion of image and point cloud data. This can be understood as follows: an image is a three-channel data structure, with each pixel containing three color components (R, G, B); a parallax point cloud is another three-channel data structure, corresponding one-to-one with the pixel units of the corrected image, with each point cloud data containing three coordinate components (x, y, z). Therefore, pre-fusion of data involves concatenating the two three-channel data sets into a six-channel data structure (R, G, B, x, y, z) according to their one-to-one correspondence, and then inputting this six-channel data structure into the AI ​​model. Therefore, compared to previous image or point cloud AI models, this fused data AI model has the longest data preparation chain; the AI ​​selection also tends to favor larger-scale models with better performance (e.g., BEV+OCC). This solution can provide an end-to-end perception solution based on the AI ​​model, achieving omnidirectional target perception in driving and parking scenarios based on the pre-fusion data structure of images and point clouds.

[0071] According to an embodiment of the present invention, an application method of a stereo vision system is also provided, namely, an intelligent parking solution based on the above-mentioned stereo vision system.

[0072] Figure 8 This is a flowchart of the application method of the stereo vision system according to Embodiment 2 of the present invention. Figure 8 As shown, the application methods of this stereo vision system include:

[0073] Step S801: Correct the distorted image data acquired based on the stereo vision system to obtain corrected image data, and generate point cloud data based on the corrected image data;

[0074] Step S802: The above-mentioned corrected image data and the above-mentioned point cloud data are fused to realize the three-dimensional reconstruction of the spatial environment around the vehicle, and a semantic point cloud map including vector features is obtained by spatial semantic point cloud mapping. The above-mentioned vector features include: three-dimensional features, relative ground position features, target category, target state, and color features.

[0075] Step S803: When the vehicle is parked, based on the above-mentioned corrected image data and the above-mentioned point cloud data, the target parking space is detected and updated, the self-tested vehicle body posture is adjusted according to the ground parking space marking lines, and spatial obstacles are identified according to the above-mentioned semantic point cloud map updated in real time.

[0076] When obtaining the semantic point cloud map through spatial semantic point cloud mapping, the following processing may also be included: memorizing the path traveled by the vehicle; guiding the vehicle to achieve low-speed navigation through the memorized path, wherein when low-speed navigation is activated, the starting point is set to the parking position and the ending point is set to the origin position at the moment when the valet parking function is activated.

[0077] The following examples illustrate forward trinocular, backward binocular, and panoramic binocular vision.

[0078] Intelligent parking mainly includes four stages: panoramic display, assisted parking, automatic parking, and advanced parking. The status requirements for sensors are shown in Table 2.

[0079]

[0080] The functional requirements of the four stages described above are progressively more advanced. Panoramic display primarily provides drivers with images and point cloud maps, using a surround-view dual-eye system as its main function to assist with parking. Assisted parking mainly involves real-time perception of the surrounding environment while the driver is parking, issuing warnings when there is a collision risk. Automatic parking refers to the vehicle automatically adjusting its position and parking near the target parking space. Advanced parking involves the vehicle starting far from a parking space, automatically finding and parking, and, upon receiving an activation signal, automatically moving from the parking space to the destination to pick up passengers.

[0081] In a stereo vision-based intelligent parking solution, the software process is as follows: Figure 8 As shown.

[0082] Figure 9 This is a flowchart illustrating the application method of the stereoscopic vision system according to a preferred embodiment of the present invention. Figure 9 As shown, the main application methods of this stereo vision system include:

[0083] Based on the distorted image data acquired by the front-view trinocular, surround-view binocular, and rear-view binocular systems described above, the distorted image, the corrected image, and the point cloud data are generated sequentially from front to back. That is, the distorted image data is distorted to obtain the corrected image data, and the point cloud data is generated based on the corrected image data.

[0084] For assisted parking and automatic parking functions, it's not enough to simply identify standard parking spaces; accessibility in three-dimensional space also needs to be assessed. Spatial semantic point cloud mapping can memorize the vehicle's travel path, guiding it point-to-point low-speed navigation within the memorized scene. First, based on corrected image data and point cloud data, a three-dimensional reconstruction of the vehicle's surrounding environment is achieved, obtaining a local semantic point cloud map, primarily composed of vector features. When low-speed navigation is activated, the starting point is designed as the parking position, and the ending point is set to the origin position at the moment the valet parking function is activated.

[0085] Therefore, the scope of point cloud mapping extends from the activation of the advanced parking function to the completion of vehicle parking; the entire spatial range needs to be recorded. The process of returning from the parking space to the original starting point involves a loop update of the recorded point cloud map, primarily achieved through vector feature matching between the local point cloud map and the remembered map to realize global positioning.

[0086] Vector feature encoding primarily records static and temporary targets. The vector features are defined as follows: (3D features, relative ground position features, target category, target state, and color features). 3D features mainly include the target object's 3D dimensions and spatial envelope; relative ground position features include orientation, pitch, and offset; target categories include vehicles, signs, typical roadblocks (such as water-filled barriers, metal barriers, and cones), and other non-standard obstacles; target states include moving, stationary, and temporary; color features refer to the statistical encoding of the RGB space of the corresponding image pixels. These information are assigned different weights during matching; the weights are influenced by the target state, with stationary states receiving the highest weight, temporary states the next highest, and moving states the lowest. Matching is defined as similarity; the feature vector is subtracted element-wise and its L2 norm is calculated to obtain a scalar value. The smaller this scalar value, the higher the degree of matching.

[0087] When parking, the target parking space is detected and updated in real time. In addition to adjusting the vehicle's posture according to the ground parking space markings, it also needs to identify spatial obstacles based on the real-time updated semantic point cloud map. On one hand, from the BEV perspective, the vehicle is parked into the corresponding parking space as a two-dimensional target. Based on the positioning and perception information from the BEV perspective, the vehicle's posture is adjusted in real time and dynamic path planning is performed. At the same time, obstacle perception is performed on both ground and non-ground targets. The vehicle is treated as a three-dimensional target, and its passability is detected in real time. When there is a collision risk, a secondary warning and braking control are performed based on the TTC threshold. If the automatic parking cannot avoid the target, the current parking space is abandoned, the vehicle is restored to its pre-parking configuration, and a new target parking space is selected.

[0088] According to an embodiment of the present invention, an application method for a stereo vision system is also provided, namely, an intelligent chassis solution based on the above-mentioned stereo vision system.

[0089] Figure 10 This is a flowchart of the application method of the stereo vision system according to Embodiment 3 of the present invention. Figure 10 As shown, the application methods of this stereo vision system include:

[0090] Step S1001: Correct the distorted image data acquired based on the stereo vision system to obtain corrected image data, and generate point cloud data based on the corrected image data;

[0091] Step S1002: Based on the above-mentioned corrected image data and point cloud data, the road surface ahead is modeled in real time to achieve terrain classification. On structured hard roads, the location information of road surface protrusions or depressions is identified and measured, changes in road surface height and road surface condition are detected, dynamic targets around the vehicle are detected, and dynamic targets with a collision probability higher than a predetermined threshold are dynamically monitored.

[0092] The detection of road surface condition changes based on the aforementioned corrected image data and point cloud data may further include: using the forward vision unit of the aforementioned stereo vision system to identify and perceive the road surface condition; obtaining the pattern and time point of the road surface condition change based on the acquired corrected image data and point cloud data; notifying the control mechanism in advance to execute a response strategy; and adjusting the vehicle according to the new dynamic model when driving on a road surface that has changed.

[0093] The aforementioned detection of dynamic targets around the vehicle and dynamic monitoring of dynamic targets with a collision probability higher than a predetermined threshold may further include: using the surround-view vision unit and rear-view vision unit of the aforementioned stereo vision system to detect dynamic targets around the vehicle in real time; based on the acquired corrected image data and point cloud data, dynamically monitoring dynamic targets with a collision probability higher than a predetermined threshold; when a warning standard is reached, calculating the vehicle state after the collision based on the predicted collision state, and correcting the vehicle power distribution based on the aforementioned vehicle state.

[0094] The following examples illustrate forward trinocular, backward binocular, and panoramic binocular vision.

[0095] The intelligent chassis mainly includes three functions: power adjustment, suspension adjustment, and stability adjustment. The status requirements of the sensors are shown in Table 3.

[0096]

[0097] Powertrain adjustment primarily involves assessing road conditions during forward driving. If road conditions change, the power distribution to all four axles needs to be optimized to ensure optimal vehicle performance. Suspension adjustment ensures timely response to impacts on the road surface during forward and rearward driving, adjusting suspension damping or providing compensation to reduce rollover risk and improve ride comfort. Stability adjustment assesses potential lateral and rearward collision risks during driving, prioritizing collision avoidance. If avoidance is not possible, the impact of a collision needs to be predicted based on the vehicle's dynamics model, and appropriate compensatory measures should be taken at the moment of impact to prevent vehicle instability.

[0098] In the intelligent chassis solution based on stereo vision, the software process is as follows: Figure 10As shown.

[0099] Figure 11 This is a flowchart illustrating the application method of the stereoscopic vision system according to a preferred embodiment of the present invention. Figure 11 As shown, the main application methods of this stereo vision system include:

[0100] Based on the distorted image data acquired by the front-view trinocular, surround-view binocular, and rear-view binocular systems described above, the distorted image, the corrected image, and the point cloud data are generated sequentially from front to back. That is, the distorted image data is distorted to obtain the corrected image data, and the point cloud data is generated based on the corrected image data.

[0101] For power distribution, a combination of vehicle dynamics and visual recognition is generally used for judgment, and the fusion scheme between the two is dynamically adjusted according to different operating conditions. Suspension adjustment generally prioritizes damping to ensure smooth driving. Based on the aforementioned corrected image data and point cloud data, real-time modeling of the road surface ahead is required. On a structured hard road surface, sudden bumps or depressions need to be identified and measured. Another typical operating condition is when there is a height jump in the ground ahead of the vehicle. The suspension compensates for this height jump in a timely manner according to the vehicle's operating status, which can significantly reduce the impact and collision sensation during driving, improving the driving and riding experience.

[0102] Vehicle stability primarily addresses scenarios involving sudden events such as tire slippage, tire blowouts, side collisions, and crosswinds. Tire slippage often occurs when road conditions change, such as icy surfaces. In these situations, a forward-facing tri-lens camera identifies and senses the road conditions, capturing patterns and timing of changes to proactively notify the control system to implement countermeasures. Immediately after approaching a changed road surface, the vehicle adjusts based on a new dynamics model. In passive instability scenarios like side collisions, surround-view and rear-view binocular cameras continuously monitor dynamic targets around the vehicle, dynamically tracking those posing a collision risk. Once relevant warning criteria are met, the system calculates the vehicle's post-collision state based on predicted collision conditions, such as relative speed, collision location, and collision force. This state is then used to adjust the vehicle's power distribution, helping the vehicle achieve rapid stabilization.

[0103] In summary, by utilizing the embodiments provided by this invention, and employing stereo vision as the primary sensor, multi-view stereo cameras are mounted at different locations on the vehicle body to achieve environmental perception and precise local mapping around the vehicle. This enables omnidirectional perception of the vehicle's surroundings. The provided perception solution supports L2 and higher-level intelligent driving solutions, intelligent parking solutions, and intelligent chassis solutions. It overcomes the challenges of the long-tail effect in intelligent driving scenarios with a low-cost solution, compensates for functional shortcomings in high-level intelligent driving, supports and develops new business areas such as chassis intelligence, and provides a more precise, robust, and low-cost intelligent perception solution.

[0104] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for applying a stereo vision system, characterized in that, include: In the intelligent driving solution, the distorted image data acquired by the stereo vision system is distorted to obtain rectified image data. Point cloud data is generated based on the rectified image data. The distorted image data is directly transmitted to a small-scale AI model with a first backbone and multiple heads for processing to obtain the first processed data. The point cloud data is transmitted to a small-scale AI model with a second backbone and multiple heads for processing to obtain the second processed data. Based on the two three-channel data, namely the corrected image data and the point cloud data, they are spliced ​​together into a six-channel data structure according to a one-to-one correspondence. The six-channel data structure is transmitted to a large-scale model for processing to obtain the third processed data, thereby realizing the task of all-round target perception in driving and parking scenarios. The first and second processed data are fused together to achieve perceptual redundancy based on different data sources. The fused data is then filtered together with the third processed data, and the filtered data is concatenated and output.

2. The application method according to claim 1, characterized in that, The first main body plus a multi-head small-scale AI model is used to achieve tasks related to target detection, lane line detection, image segmentation, and drivable area; The second backbone plus a multi-head small-scale AI model is used to realize tasks related to point cloud segmentation, target detection, road surface detection, and drivable area detection; The large-scale model is used to achieve the task of perceiving targets in various directions in driving and parking scenarios based on the data obtained by fusing the corrected image data and the point cloud data.

3. A stereo vision system, performing the method as described in claim 1 or 2, characterized in that, include: The forward vision unit, located at the front of the vehicle, includes one or more telephoto cameras and one or more wide-angle cameras, for detecting objects within a predetermined range of the vehicle's forward vision, acquiring lateral and longitudinal distance measurements and relative speed measurements, and for real-time modeling of the road surface in front of the vehicle. The rear-view vision unit, located at the rear of the vehicle, includes one or more wide-angle cameras for detecting objects within a predetermined range of the vehicle's rear view and for real-time modeling of the road surface behind the vehicle. The surround-view vision unit, located at the front and rear, and left and right sides of the vehicle, includes multiple sets of multi-view cameras for detecting objects within a predetermined surround-view range of the vehicle and for local modeling.

4. The stereo vision system according to claim 3, characterized in that, The forward-looking vision unit is located in the black area in the upper middle part of the windshield inside the cockpit. The rear-view vision unit is located at the shark fin area at the rear of the vehicle, or on the upper outer side of the rear windshield. The surround view unit is located above the front and rear license plates of the vehicle, as well as in the vicinity of the left and right fenders or rearview mirrors.

5. The stereo vision system according to claim 3, characterized in that, The optical axis of the camera in the forward vision unit and the optical axis of the camera in the rear vision unit are both horizontal with respect to the ground. In the surround-view vision unit, on the XOY plane, the optical axes of cameras positioned at the front and rear of the vehicle body are parallel to the vehicle body's centerline, while the optical axes of cameras positioned on the left and right sides of the vehicle body are perpendicular to the vehicle body's centerline. On the YOZ plane, the optical axes of cameras positioned to the left and right of the vehicle body are tilted towards the ground so that the optical axis point falls on the ground near the vehicle body. On the XOZ plane, the optical axes of cameras positioned to the front and rear of the vehicle body are tilted towards the ground so that the optical axis point falls on the ground near the vehicle body.

6. The stereoscopic vision system according to any one of claims 3 to 5, characterized in that, The forward-looking vision unit includes: a telephoto camera centrally located, and wide-angle cameras respectively located on the left and right sides of the telephoto camera; The rear-view vision unit includes: two wide-angle cameras; The surround vision unit includes four sets of binocular cameras with identical structural parameters, wherein one set of binocular cameras is respectively located at the front, rear, left, and right of the vehicle body.

Citation Information

Patent Citations

  • Alarm system and alarm method for vehicle

    CN101443830A

  • A multiocular camera system, an intelligent driving system, an automobile, a method and a storage medium

    CN107122770A

  • Multi-sensor fused intelligent parking system and method

    CN112180373A

  • Method and system for controlling vehicle driving through road surface perception

    CN118082850A

  • Basement scene automatic vector map construction method

    CN120279527A