Automatic driving multi-view fusion elasticity enhancement method and system based on aerial view
Through a multi-view fusion method based on bird's eye view, multi-camera data preprocessing and compensation strategies are used to solve the problem of degradation of environmental perception capabilities of the autonomous driving system when sensor abnormalities are abnormal, efficient obstacle detection and safe driving are achieved, and the robustness and adaptability of the system are improved.
Patent Information
- Application Number
- CN202510157246.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-08
AI Technical Summary
When a single or multiple sensors fail, the environmental perception capability of existing autonomous driving systems will be reduced, data fusion in multiple visual field overlapping areas is insufficient, and the effective elastic enhancement mechanism is lacking, resulting in a decrease in system robustness and safety.
Using a multi-view fusion method based on bird's eye view, image data is acquired through multiple cameras installed in front of the vehicle, pre-processing, timestamp alignment and abnormal state detection is performed, obstacle information in abnormal areas is calculated from normal camera data using compensation strategies, and weighted scores are performed to ensure the safe driving of the vehicle.
It realizes efficient integration of multi-view data in the case of sensor abnormalities, improves the accuracy of obstacle detection and tracking, ensures the real-time and adaptability of the autonomous driving system, improves the elasticity and robustness of the system, and avoids technical difficulties in the image stitching process.
Smart Images

Figure CN120278893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving environment perception, and particularly relates to a method and system for elastic enhancement of multi-view fusion in autonomous driving based on a Bird's Eye View (BEV). Background Art
[0002] Since the 1980s, with the progress of computer vision and artificial intelligence technologies, the concept of autonomous driving vehicles has gradually been proposed. Early research mainly focused on rule-based methods, such as achieving basic vehicle control through preset path planning and simple environment perception. After entering the 21st century, with the development of sensor technologies (cameras, radars, lidars, etc.) and the breakthrough of machine learning algorithms, autonomous driving technology has started to develop rapidly. Especially in recent years, the application of deep learning in image recognition and processing has greatly improved the vehicle's ability to understand complex environments.
[0003] With the continuous progress of autonomous driving technology, the vehicle's precise perception ability of the surrounding environment has become one of the key factors to ensure safe driving. In recent years, pure vision solutions have made remarkable progress in the field of autonomous driving. With the progress of sensor technology, high-resolution and high-frame-rate cameras have become standard configurations, capable of capturing richer image information. The progress of deep learning and computer vision technologies has greatly improved the image processing ability. The application of technologies such as convolutional neural networks (CNNs), object detection algorithms (such as YOLO, SSD), and semantic segmentation has significantly improved the perception accuracy and robustness of the system. To achieve this goal, modern autonomous driving systems widely adopt a combination of multiple sensors, such as cameras, radars, and lidars. These devices can capture environmental information from different angles and positions. Especially for the monitoring of the vehicle's front field of view, usually three main cameras are installed on the left, middle, and right to obtain a wide-angle view, thus providing continuous and non-blind-spot environmental coverage. Through multi-sensor data fusion technology, data from different sensors can be comprehensively processed to improve the robustness and accuracy of the system and ensure reliable operation under various conditions.
[0004] However, although the multi-sensor configuration improves the overall perception performance of the system, the potential failure problem of a single sensor still exists, which may be caused by physical damage, occlusion, or malicious attacks. Once a key camera fails, it will directly lead to the loss of some important fields of view, seriously affecting the decision-making ability and safety of the autonomous driving system. Therefore, how to maintain a highly reliable environment perception ability when one or more sensors are abnormal has become an urgent problem to be solved in current autonomous driving research. Summary of the Invention
[0005] To this end, the present invention provides a method and system for elastic enhancement of multi - perspective fusion in autonomous driving based on bird's - eye view, which solves the problems existing in the environmental perception of existing autonomous driving, such as the reduction of the vehicle's comprehensive environmental perception ability when a single sensor fails, insufficient data fusion in the overlapping areas of multiple fields of view, and the lack of an effective elastic enhancement mechanism in the case of an attack or other forms of failure on a certain sensor.
[0006] According to the design scheme provided by the present invention, on the one hand, a method for elastic enhancement of multi - perspective fusion in autonomous driving based on bird's - eye view is provided, including:
[0007] Obtain image data of different perspectives by using multiple cameras installed on the vehicle, and pre - process the image data. The pre - processing includes image data parsing, image coordinate system normalization, and image data timestamp alignment.
[0008] Judge the abnormal state of each camera according to the obstacle information at the front and rear timestamps in the image data, and use a compensation strategy to detect and output the obstacles in the perspective of the abnormal - state camera. The compensation strategy is used to calculate and output the obstacle information in the perspective area of the abnormal - state camera by using the image data of the normal cameras.
[0009] As the method for elastic enhancement of multi - perspective fusion in autonomous driving based on bird's - eye view of the present invention, further, the pre - processing of the image data includes:
[0010] Screen the image data from the front - left camera, front - middle camera, and front - rear camera fields of view of the vehicle according to the driving scene and the number of obstacles.
[0011] Establish a vehicle global coordinate system with the vehicle's own position as the reference point, and use the rotation matrix and translation vector to adjust the relative positions and directions of each camera to integrate the image data of each camera.
[0012] Align the image data of each camera according to the timestamp of image data acquisition of each camera, and make the image data of each camera have corresponding points on the same time axis through interpolation operation.
[0013] As the method for elastic enhancement of multi - perspective fusion in autonomous driving based on bird's - eye view of the present invention, further, judging the abnormal state of each camera according to the obstacle information at the front and rear timestamps in the image data includes:
[0014] Compare the obstacles and coordinates at the front and rear timestamps in the image data of each camera. If the obstacle coordinates and / or the front and rear timestamps are abnormal, it is determined that the camera state is abnormal.
[0015] Or, if there is an abnormality in the obstacle information in the image data of two or more cameras with overlapping areas under multiple consecutive identical timestamps, it is determined that the corresponding camera state is abnormal.
[0016] As the multi - perspective fusion elastic enhancement method for autonomous driving based on the bird's - eye view of the present invention, further, a compensation strategy is used to detect and output obstacles in the image data of the abnormal - state camera, including:
[0017] In the case where the camera is in an abnormal state, the image data obtained by other cameras within the same time period is used as backup data;
[0018] The obstacle information in the perspective area of the abnormal - state camera is recalculated using the backup data and output.
[0019] As the multi - perspective fusion elastic enhancement method for autonomous driving based on the bird's - eye view of the present invention, further, the process of recalculating and outputting the obstacle information in the perspective area of the abnormal - state camera using the backup data includes:
[0020] Extract the overlapping area formed by the camera perspective in the backup data and the abnormal - state camera;
[0021] Set the weights of each camera according to the camera reference factors, and assign different weight values to the obstacles in the overlapping area according to the obstacle types. The camera reference factors include the historical performance of the camera, the current working state of the camera, the camera perspective, and the camera distance;
[0022] Divide the overlapping area into a dangerous area, an observation area, and a safe area according to the preset minimum safety distance and maximum safety distance, and perform weighted scoring on the obstacles in the overlapping area according to the camera weights and the obstacle weight values;
[0023] Determine whether there are obstacles in the abnormal - state perspective area according to the weighted scoring result, and when there are obstacles, take specified safety measures to ensure the safe driving of the vehicle. The specified safety measures include deceleration warning and / or steering warning.
[0024] As the multi - perspective fusion elastic enhancement method for autonomous driving based on the bird's - eye view of the present invention, further, the process of performing weighted scoring on the obstacles in the overlapping area according to the camera weights and the obstacle weight values is expressed as:
[0025] a, b, and c respectively represent three cameras, A and B respectively represent the overlapping areas formed by the perspectives of the other two cameras and the abnormal - state camera, Score A 、Score B 、Score a 、Score b 、Score crespectively represent the obstacle scores of the overlapping area A, the overlapping area B, the camera a, the camera b, and the camera c. x1, x2, and x3 respectively represent the weights assigned to the danger zone, the observation zone, and the safe zone. i represents the type of obstacle, and y i represents the weight value assigned to the corresponding type of obstacle, and n i represents the number of obstacles of the corresponding type in the corresponding area. m1, m2, and m3 respectively represent the total number of obstacles detected in the danger zone, the observation zone, and the safe zone. z a 、z b 、z c respectively represent the weights corresponding to different cameras.
[0026] As the multi-view fusion elastic enhancement method for autonomous driving based on the bird's-eye view of the present invention, further, it is determined whether there are obstacles in the abnormal state view area according to the weighted score result, including:
[0027] If the weighted score result exceeds the preset threshold, it is determined that there are obstacles in the corresponding overlapping area between the abnormal state view and other cameras.
[0028] On the other hand, the present invention also provides a multi-view fusion elastic enhancement system for autonomous driving based on the bird's-eye view, including: a data processing module and a fusion output module, wherein,
[0029] The data processing module is used to obtain different perspective image data by using multiple cameras installed on the vehicle, and preprocess the image data. The preprocessing includes image data parsing, image coordinate system normalization, and image data timestamp alignment;
[0030] The fusion output module is used to judge the abnormal state of each camera according to the obstacle information of the front and rear timestamps in the image data, and use the compensation strategy to detect and output the obstacles in the abnormal state camera view. The compensation strategy is used to calculate and output the obstacle information of the abnormal state camera view area by using the image data of the normal camera.
[0031] The beneficial effects of the present invention:
[0032] The present invention can utilize the image information collected by three heterogeneous wide-angle cameras installed on the front of the vehicle to achieve efficient fusion of the original images obtained from different perspectives into a unified coordinate system. This not only simplifies the subsequent processing flow but also enables data from different cameras to be compared and analyzed within the same space, greatly improving the accuracy of obstacle detection and tracking. At the same time, image stitching is no longer performed, avoiding technical problems during the image stitching process. When a certain camera fails or is attacked, the solution in this case can immediately detect the abnormal situation and activate a pre-designed compensation strategy. Based on the data provided by the unaffected cameras, important parameters such as the position and speed of obstacles in the affected area are recalculated to ensure that the entire perception system does not lose its function due to problems with a single sensor. This can ensure that the vehicle can achieve real-time data processing and decision support even in a traffic environment with high dynamic changes, having good real-time performance and adaptability for autonomous driving environment perception, and being able to enhance the flexibility and robustness of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Schematic diagram of the process for enhancing the flexibility of multi-view fusion in autonomous driving based on a bird's-eye view in the embodiment;
[0034] Figure 2 Schematic diagram of the algorithm principle framework in the embodiment;
[0035] Figure 3 Schematic diagram of the division of the overlapping area in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.
[0037] Aiming at the problems in autonomous driving environment perception, such as the risk of failure of a single sensor, insufficient data fusion in the overlapping area of multi-sensor perspectives, and the lack of an effective flexibility enhancement mechanism when the sensor is attacked or fails, the embodiment of the present invention, as shown in Figure 1 provides a method for enhancing the flexibility of multi-view fusion in autonomous driving based on a bird's-eye view, including:
[0038] S101. Obtain image data from different perspectives using multiple cameras installed on the vehicle, and preprocess the image data. The preprocessing includes image data parsing, image coordinate system normalization, and image data timestamp alignment.
[0039] Among them, the preprocessing of the image data can be designed to include:
[0040] Screen the image data from the fields of view of the front left camera, front middle camera, and front and rear cameras of the vehicle according to the driving scenario and the number of obstacles;
[0041] A vehicle global coordinate system is established with the vehicle's own position as the reference point, and the relative positions and directions of each camera are adjusted using rotation matrices and translation vectors to integrate the image data of each camera.
[0042] The image data of each camera is time-aligned according to the time stamps of the image data acquisition of each camera, and interpolation operations are performed to make the image data of each camera have corresponding points on the same time axis.
[0043] Data preprocessing is a key step to ensure the accuracy and efficiency of subsequent analysis. It mainly includes three main links: dataset parsing, coordinate normalization, and time stamp alignment, ensuring that the data input into the system is not only of high quality but also has good consistency, thus providing reliable support for subsequent multi-view fusion and elastic enhancement mechanisms.
[0044] According to the research objectives and application scenarios, the most relevant samples are selected from a large amount of raw data through dataset parsing. For example, in an autonomous driving environment, samples containing typical traffic participants such as vehicles, pedestrians, and bicycles need to be selected. These samples should cover various typical driving scenarios, such as urban roads, highways, intersections, etc., to ensure the generalization ability of the model. When selecting samples, the number of obstacles in each sample also needs to be considered. Usually, scenarios with a moderate number of obstacles are selected, avoiding overly crowded or empty situations, which can ensure the effectiveness and representativeness of model training. The data mainly screened comes from the front left camera, the front middle camera, and the front right camera because these cameras provide the key views in front of the vehicle.
[0045] To facilitate the integration and processing of data from different sensors, a unified global coordinate system needs to be determined. Usually, this global coordinate system is established with the vehicle's own position as the reference point, such as using the vehicle's GPS position as the origin. The data of all sensors is converted into this unified global coordinate system. This mainly involves using rotation matrices and translation vectors to adjust the relative positions and directions of each sensor to ensure that they can be compared and fused within the same spatial framework. For camera data, in addition to coordinate transformation, the internal parameters of the camera (such as focal length, principal point position) and external parameters (describing the position and orientation of the camera relative to the vehicle) are also needed for more accurate calibration. By applying the projection transformation formula, the image pixel coordinates are converted into world coordinates. This helps to eliminate information distortion caused by perspective differences and improve the consistency of multi-source data.
[0046] Since the time points of data collection by different sensors may not be exactly the same, it is necessary to align the data of all sensors in terms of time. This can be achieved by recording the specific timestamps of data collection for each sensor. For sensors with different sampling frequencies, linear interpolation or other advanced interpolation methods can be used to make all data have corresponding points on the same time axis. For example, if the data collection interval of a certain sensor is long, while the data collection of another sensor is more frequent, intermediate values can be inserted between two time points to achieve a smooth transition in time.
[0047] S102. Determine the abnormal states of each camera based on the obstacle information at the previous and subsequent timestamps in the image data, and use a compensation strategy to detect and output the obstacles in the view of the abnormally - state camera. The compensation strategy is used to calculate and output the obstacle information in the view area of the abnormally - state camera by using the image data of the normal cameras.
[0048] Among them, determining the abnormal states of each camera based on the obstacle information at the previous and subsequent timestamps in the image data may include:
[0049] Compare the obstacles and coordinates at the previous and subsequent timestamps in the image data of each camera. If the obstacle coordinates and / or the previous and subsequent timestamps are abnormal, it is determined that the camera state is abnormal;
[0050] Or, if there are abnormalities in the obstacle information in the image data of two or more cameras with overlapping areas at multiple consecutive identical timestamps, it is determined that the corresponding camera state is abnormal.
[0051] And for the case where a camera has an abnormal state, use the image data obtained by other cameras within the same time period as backup data; recalculate and output the obstacle information in the view area of the abnormally - state camera by using the backup data.
[0052] Specifically, recalculating and outputting the obstacle information in the view area of the abnormally - state camera by using the backup data can be designed to include:
[0053] Extract the overlapping area formed by the camera view in the backup data and the abnormally - state camera;
[0054] Set the weights of each camera according to the camera reference factors, and assign different weight values to the obstacles in the overlapping area according to the types of obstacles. The camera reference factors include the historical performance of the camera, the current working state of the camera, the camera view, and the camera distance;
[0055] Divide the overlapping area into a danger area, an observation area, and a safe area according to the preset minimum safety distance and maximum safety distance, and perform weighted scoring on the obstacles in the overlapping area according to the camera weights and the obstacle weight values;
[0056] Determine whether there are obstacles in the abnormal state perspective area based on the weighted scoring result, and when there are obstacles, take specified safety measures to ensure the safe driving of the vehicle. The specified safety measures include deceleration warning and / or steering warning.
[0057] As Figure 2 shown in the algorithm framework, through data preprocessing of three links: dataset parsing, coordinate system normalization, and timestamp alignment, based on multi-view data fusion of the overlapping field of view area of the camera, using the image information collected by the left, middle, and right three wide-angle cameras installed in front of the vehicle, combined with BEV conversion technology, to achieve efficient fusion of the original images obtained from different perspectives into a unified coordinate system; using the elastic enhancement mechanism, when a certain camera fails or is attacked, it can immediately detect abnormal situations and activate a pre-designed compensation strategy, and recalculate important parameters such as the position and speed of obstacles in the affected area according to the data provided by the unaffected cameras, ensuring that the entire perception system will not lose its function due to problems with a single sensor. By making full use of the overlapping field of view area between multiple cameras to improve the environmental perception ability of autonomous vehicles, it is especially optimized for the situation when a key camera is attacked or fails, ensuring that the system can continue to maintain effective perception of the environment through the data of other cameras, thus significantly enhancing the elasticity and robustness of the system.
[0058] Among them, determine the overlapping areas between different sensors or different perspectives, and these areas are: the overlapping area A between the front left camera and the front middle camera, and the overlapping area B between the front middle camera and the front right camera.
[0059] Collect the detection data of obstacles from each sensor. After the dataset is preprocessed, it is sent to the corresponding camera module. After algorithm processing, each camera will obtain relevant information such as the number, type, coordinates, and speed of the detected obstacles after processing.
[0060] Obtain the minimum safety distance d according to the RSS (Responsibility Sensitive Safety) model min . The minimum safety distance refers to the distance that can still avoid collisions under the worst conditions. The worst condition means that the leading vehicle starts braking with the maximum braking acceleration, the following vehicle discovers it and has a certain reaction time, and still advances with the maximum acceleration during the reaction time, and then changes to brake with the minimum braking acceleration until the danger is lifted. Therefore, according to the RSS model, the calculation formula of the minimum safety distance d min is as follows:
[0061]
[0062] At the same time, set a maximum safety distance d max as a threshold.
[0063] According to the minimum safety distance d min and the maximum safety distance d max The overlapping area is divided into three parts: the dangerous area, the observation area, and the safe area. The dangerous area refers to the area where the distance is less than or equal to d min When the obstacle is in the dangerous area, the weight it occupies is relatively large; the observation area is the area where the distance is greater than d min but at the same time less than d max In this area, the obstacle generally will not interfere with the normal driving of the vehicle, but still requires certain observation. The obstacle has the risk of entering the dangerous area. Therefore, the weight of the obstacle in this area is relatively small; the safe area is the area where the distance is greater than or equal to d max In this area, the obstacle is far from the vehicle position and will not interfere with the normal driving of the vehicle in a short time. The weight of the obstacle in this area is 0. Draw the overlapping area according to the divided areas as Figure 3 shown, and divide the overlapping area into the dangerous area, the observation area, and the safe area.
[0064] As Figure 2 shown in the algorithm, attacking a camera module and verifying whether the attack is successful mainly includes two parts: the attack method and the attack verification.
[0065] Among them, the attack method mainly attacks a certain camera by selecting one of the following three methods:
[0066] Physical attack involves directly physically operating on the camera to interfere with or damage its normal operation. In the test environment, a similar scenario can be built to simulate a real-world physical attack, such as covering the lens so that the camera cannot capture clear images and is difficult to accurately identify target objects.
[0067] Network attack refers to a remote attack on the camera by exploiting network protocol vulnerabilities or weaknesses. This may involve the following activities: using a network sniffer tool to capture the communication traffic between the camera and the server and attempting to modify the data packets in transit, such as changing the video frame content or command instructions. Inserting oneself as an intermediate node to intercept and manipulate all communications between the camera and the control end. In the test environment, an isolated virtual network can be constructed to reproduce the above attack scenario while ensuring that no risk is posed to the actual operating environment.
[0068] Denial of Service (DoS / DDoS) attacks are carried out by consuming the resources of the camera and its associated systems, rendering them unable to provide normal monitoring services. Common practices include: sending a large number of requests to the camera until its processing capacity is saturated, preventing legitimate users from accessing; taking advantage of the characteristics of certain protocols to send a very small amount of data while keeping the connection open, gradually exhausting the server resources; and blocking other devices from connecting to the camera network normally by occupying too much bandwidth. In a test environment, appropriate tools and parameters should be carefully selected to avoid unnecessary impacts on non-target systems, while recording the comparison of the system's performance metrics before and after each attack.
[0069] Attack verification is to verify whether the attack is successful, and it is verified in two ways: One is to collect the timestamps and coordinate data of the obstacles detected by the camera module for the timestamps before and after the attack, compare the data collected before and after the attack, and analyze the changes in the obstacle coordinates. If the following situations are observed, such as abnormal or inconsistent changes in the obstacle coordinates, inconsistent data acquisition intervals shown in the timestamp records, a significant decrease in the confidence level of obstacle detection, etc., based on the comparison results, it can be evaluated whether the attack is successful. If the camera cannot correctly detect obstacles or provides incorrect information after being attacked, the attack can be considered successful. The other is to compare the obstacle data collected by two cameras with overlapping areas at the same timestamp after the attack. If the camera under attack has missing or completely lost obstacle information detected in the overlapping area compared with the normal camera, and this situation is observed at multiple consecutive timestamps after the attack, then it can be judged that the attack is successful.
[0070] When a certain camera of a vehicle is attacked, it may cause it to be unable to correctly detect obstacles within a specific time or duration, resulting in a certain degree of functional failure. This failure may manifest as a complete lack of image output, a serious decline in image quality (such as blurring, distortion), or incorrect recognition of the surrounding environment (for example, mistaking an obstacle for no obstacle). This will directly threaten the safe driving of the vehicle, especially in the autonomous driving mode.
[0071] For example, if the front left camera is attacked and its function is damaged, and it cannot display the presence of any obstacles for a period of time, the system will enable the complementary mechanism in the redundant design. It will check the data captured by other cameras during the same period and can rely on the front middle side camera with an overlapping field of view to obtain information about the same area as a backup data source to obtain obstacle information in the overlapping area A.
[0072] In order to make the optimal judgment from the data of multiple cameras, the system can adopt a weighted calculation method to evaluate the importance of the information provided by each camera. The specific steps of the weighted calculation can be summarized as follows:
[0073] 1). Weight assignment: Different weight values are assigned to each camera based on factors such as its historical performance, current working status (whether it has been attacked), field of view angle, distance, etc. Cameras that are not affected and perform stably are given higher weights; while for cameras known to have problems, their weights are appropriately reduced. At the same time, according to the different types of obstacles, different weights are assigned to the obstacles in the overlapping area. Pedestrians, electric vehicles, bicycles, cars, etc. should occupy different weights.
[0074] 2). Data fusion: Collect the obstacle detection results of all relevant cameras within the overlapping area A, including parameters such as position coordinates, speed, confidence level, etc., and integrate them together to form a comprehensive data set.
[0075] 3). Weighted scoring: Based on the above weights, score the obstacles reported by each camera. This score reflects the likelihood of the existence of the obstacle and its characteristics. The formula is as follows:
[0076]
[0077] Where: a, b, and c respectively represent three cameras, x1, x2, and x3 respectively represent the weights assigned to the dangerous area, observation area, and safe area, i represents the type of obstacle, corresponding to one of pedestrians, bicycles, electric vehicles, and cars, y i represents the weight assigned to the corresponding type of obstacle, n i represents the number of a certain type of obstacle in the corresponding area, and m1, m2, and m3 respectively represent the total number of obstacles detected in the dangerous area, observation area, and safe area.
[0078] 4). Final decision: Through the weighted scores, the scores are respectively obtained according to the following formula:
[0079] Score A = z a × Score a + z b × Score b
[0080] Score B = z b × Score b + z c × Score c
[0081] Where: A and B respectively represent the overlapping areas of the front left and front middle cameras, and the front right and front middle cameras, z a , z b , z cThey respectively represent the weights corresponding to different cameras. At the same time, when a certain camera is verified to be under attack, the corresponding weight should be 0.
[0082] If the total score exceeds the preset threshold, it is confirmed that there are obstacles in the corresponding area, and corresponding safety measures (such as deceleration, steering warning, etc.) are taken to ensure the safe driving of the vehicle, ensure that the attacked camera still performs some functions, and the damaged camera recovers elastically in part of its functions.
[0083] Furthermore, based on the above method, an embodiment of the present invention further provides a multi-view fusion elastic enhancement system for autonomous driving based on a bird's-eye view, including: a data processing module and a fusion output module, where,
[0084] The data processing module is used to obtain image data from different perspectives by using multiple cameras installed on the vehicle, and preprocess the image data. The preprocessing includes image data parsing, image coordinate system normalization, and image data timestamp alignment;
[0085] The fusion output module is used to judge the abnormal state of each camera according to the obstacle information at the front and rear timestamps in the image data, and detect and output the obstacles in the perspective of the abnormal state camera by using a compensation strategy. The compensation strategy is used to calculate and output the obstacle information in the perspective area of the abnormal state camera by using the image data of the normal camera.
[0086] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0087] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0088] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0089] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.
[0090] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An elastic enhancement method for multi - perspective fusion of autonomous driving based on bird's - eye view, characterized in that Including: Obtaining image data from multiple cameras installed on a vehicle at different perspectives and preprocessing the image data, where the preprocessing includes image data parsing, image coordinate system normalization, and image data timestamp alignment; Judging the abnormal states of each camera according to the obstacle information at the front and rear timestamps in the image data, and using a compensation strategy to detect and output the obstacles in the perspective of the abnormal state camera, where the compensation strategy is used to calculate and output the obstacle information in the perspective area of the abnormal state camera by using the image data of the normal cameras.
2. The method for flexibly enhancing multi-view fusion of autonomous driving based on an aerial view according to claim 1, wherein Preprocessing the image data includes: Screening the image data from the fields of view of the front left camera, front middle camera, and front and rear cameras of the vehicle according to the driving scenario and the number of obstacles; Establishing a vehicle global coordinate system with the vehicle's own position as a reference point, and using a rotation matrix and a translation vector to adjust the relative positions and directions of each camera to integrate the image data of each camera; Performing time alignment on the image data of each camera according to the image data acquisition timestamps of each camera, and making the image data of each camera have corresponding points on the same time axis through interpolation operations.
3. The method for flexibly enhancing multi - perspective fusion of autonomous driving based on an aerial view according to claim 1, wherein, Judging the abnormal states of each camera according to the obstacle information at the front and rear timestamps in the image data, including: Comparing the obstacles and coordinates at the front and rear timestamps in the image data of each camera, and if the obstacle coordinates and / or the front and rear timestamps are abnormal, determining that the camera state is abnormal; Or, if there is an abnormality in the obstacle information in the image data of two or more cameras with overlapping areas at multiple consecutive identical timestamps, determining that the corresponding camera state is abnormal.
4. The method for flexibly enhancing multi - perspective fusion of autonomous driving based on an aerial view according to claim 1, wherein Using a compensation strategy to detect and output the obstacles in the image data of the abnormal state camera, including: For the case where the camera has an abnormal state, taking the image data obtained by other cameras in the same time period as backup data; Recalculating and outputting the obstacle information in the perspective area of the abnormal state camera by using the backup data.
5. The method for flexibly enhancing multi-view fusion of autonomous driving based on an aerial view according to claim 1 or 4, characterized in that, Recalculating and outputting the obstacle information in the perspective area of the abnormal state camera by using the backup data, including: Extracting the overlapping area formed by the camera perspective in the backup data and the abnormal state camera; Setting the weights of each camera according to the camera reference factors, and assigning different weight values to the obstacles in the overlapping area according to the obstacle types, where the camera reference factors include the historical performance of the camera, the current working state of the camera, the camera perspective, and the camera distance; Dividing the overlapping area into a danger area, an observation area, and a safe area according to the preset minimum safety distance and maximum safety distance, and performing weighted scoring on the obstacles in the overlapping area according to the camera weights and the obstacle weight values; Judging whether there are obstacles in the abnormal state perspective area according to the weighted scoring result, and when there are obstacles, taking specified safety measures to ensure the safe driving of the vehicle, where the specified safety measures include deceleration warning and / or steering warning.
6. The method for flexibly enhancing multi - perspective fusion of autonomous driving based on an aerial view according to claim 5, characterized in that The process of weighted scoring of obstacles in the overlapping area according to the camera weight and obstacle weight values is expressed as: a, b, and c respectively represent three cameras, and A and B respectively represent the overlapping areas formed by the perspectives of the other two cameras and the abnormal state camera. Score A , Score B , Score a , Score b , Score c respectively represent the obstacle scores of the overlapping area A, the overlapping area B, the camera a, the camera b, and the camera c. x1, x2, and x3 respectively represent the weights assigned to the dangerous area, the observation area, and the safe area. i represents the type of obstacle. y i represents the weight value assigned to the corresponding type of obstacle. n i represents the number of obstacles of the corresponding type in the corresponding area. m1, m2, and m3 respectively represent the total number of obstacles detected in the dangerous area, the observation area, and the safe area. z a , z b , z c respectively represent the weights corresponding to different cameras.
7. The method for flexibly enhancing multi-view fusion of autonomous driving based on an aerial view according to claim 5, wherein Judging whether there are obstacles in the abnormal state perspective area according to the weighted scoring result, including: If the weighted scoring result exceeds the preset threshold, determining that there are obstacles in the corresponding overlapping area between the abnormal state perspective and other cameras.
8. An elastic enhancement system for multi-view fusion of autonomous driving based on a bird's-eye view, characterized in that, Including: a data processing module and a fusion output module, where, A data processing module, configured to obtain image data from different perspectives by using a plurality of cameras installed on a vehicle, and preprocess the image data, where the preprocessing includes image data parsing, image coordinate system normalization, and image data timestamp alignment; A fusion output module, configured to determine the abnormal states of each camera according to the obstacle information at the front and rear timestamps in the image data, and detect and output the obstacles in the perspective of the abnormal state camera by using a compensation strategy, where the compensation strategy is used to calculate and output the obstacle information in the perspective area of the abnormal state camera by using the image data of the normal camera.
9. An electronic device, characterized in that, Comprising: At least one processor, and a memory coupled to the at least one processor; Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.