Vehicle automatic driving control method and related device

CN122808773APending Publication Date: 2026-09-25VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610875106.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,在实际道路行驶过程中,车辆经常会遇到建筑、山体、植被等障碍物遮挡视线的场景,尤其是山区弯道、狭窄巷道等典型盲区路段,视觉与雷达类传感器的感知范围会被完全阻断,无法提前探测到盲区内的对向来车

Benefits of technology

[0018]根据本申请实施例的第四方面,提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有至少一条计算机程序指令,所述至少一条计算机程序指令由处理器加载并执行以实现如上述第一方面任一项所述的方法所执行的操作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808773A_ABST
    Figure CN122808773A_ABST
Patent Text Reader

Abstract

The application discloses a vehicle automatic driving control method and related equipment, the method comprises the following steps: acquiring the audio data outside the vehicle collected in the driving process, visual sensing data and laser radar data; preprocessing the audio data outside the vehicle to obtain audio spatial features; and converting the visual sensing data into visual bird's eye view features, and converting the laser radar data into laser bird's eye view features; fusing the audio spatial features, the visual bird's eye view features and the laser bird's eye view features to obtain multi-modal fusion features; identifying the blind area to the vehicle based on the multi-modal fusion features, and controlling the vehicle driving according to the identification result. The technical scheme provided by the application can improve the safety of vehicle automatic driving control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of vehicle autonomous driving technology, and in particular relates to a vehicle autonomous driving control method and related equipment. Background Technology

[0002] With the rapid iteration of automotive intelligence technology, autonomous driving has gradually evolved from Level 2 assisted driving to Level 3 and above advanced autonomous driving. Vehicle autonomous perception, decision-making, and control capabilities have become the core direction of industry development. In autonomous driving systems, environmental perception is fundamental to ensuring driving safety. Existing technologies mainly achieve perception of the surrounding environment through the fusion of visual and radar sensors such as cameras, millimeter-wave radar, and lidar. This enables effective identification of obstacles such as vehicles and pedestrians in unobstructed open environments, providing data support for autonomous driving decisions.

[0003] However, in actual road driving, vehicles often encounter situations where obstacles such as buildings, mountains, and vegetation obstruct their view, especially in typical blind spot sections such as mountain curves and narrow alleys. The perception range of visual and radar sensors is completely blocked, making it impossible to detect oncoming vehicles in the blind spot in advance. When an oncoming vehicle suddenly leaves the blind spot, the autonomous driving system often cannot react effectively in time. If the vehicle is traveling at too high a speed or the oncoming vehicle is a large vehicle with a long braking distance and a large body size, a collision is highly likely, seriously threatening the lives and property of the occupants.

[0004] Therefore, improving the safety of autonomous driving control of vehicles has become an urgent technical problem to be solved. Summary of the Invention

[0005] The embodiments of this application provide a vehicle autonomous driving control method, apparatus, computer program product, computer-readable storage medium, and vehicle, which can at least to some extent improve the safety of vehicle autonomous driving control.

[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0007] According to a first aspect of the embodiments of this application, a method for controlling autonomous driving of a vehicle is provided. The method includes: acquiring external audio data, visual sensing data, and lidar data collected during vehicle operation; preprocessing the external audio data to obtain audio spatial features; converting the visual sensing data into visual bird's-eye view features and converting the lidar data into lidar bird's-eye view features; fusing the audio spatial features, the visual bird's-eye view features, and the lidar bird's-eye view features to obtain multimodal fusion features; identifying vehicles approaching in blind spots based on the multimodal fusion features, and controlling the vehicle's operation according to the identification results.

[0008] In some embodiments of this application, based on the foregoing scheme, the preprocessing of the external audio data to obtain audio spatial features includes: performing adaptive filtering on the external audio data to remove vehicle noise and obtain a preliminary denoised frequency; performing weighted denoising on the preliminary denoised frequency to remove reverberation caused by sound reflection and obtain a direct sound frequency; and performing sound source localization on the direct sound frequency to determine the incident angle of the sound and obtain the audio spatial features.

[0009] In some embodiments of this application, based on the foregoing scheme, fusing the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features to obtain multimodal fusion features includes: converting the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features into feature tokens of a unified format, and adding spatiotemporal location encoding to each feature token; and performing weighted fusion of the feature tokens through an attention mechanism to obtain the multimodal fusion features.

[0010] In some embodiments of this application, based on the foregoing scheme, the weighted fusion of the feature tokens through an attention mechanism includes: assigning different weights to the corresponding feature tokens according to the confidence level of each modality data, so that the feature tokens corresponding to the laser bird's-eye view features pay attention to the texture information of the feature tokens corresponding to the visual bird's-eye view features, the feature tokens corresponding to the visual bird's-eye view features pay attention to color and texture information, and the feature tokens corresponding to the audio spatial features pay attention to the sound information in the blind zone direction; and fusing the multimodal fused features based on the assigned weights and the attention results of each feature token.

[0011] In some embodiments of this application, based on the aforementioned scheme, the identification of vehicles approaching from blind spots based on the multimodal fusion features includes: performing three-dimensional target detection based on the multimodal fusion features; when a vehicle approaching from a blind spot is detected in a blind spot not covered by visual sensing data and lidar data, enhancing the identification of the vehicle approaching from the blind spot by combining audio spatial features; tracking the movement trajectory of the vehicle approaching from the blind spot, and predicting the driving direction and speed of the vehicle approaching from the blind spot.

[0012] In some embodiments of this application, based on the aforementioned scheme, controlling the vehicle's movement according to the identification result includes: determining the vehicle type of the vehicle approaching from the blind spot and the collision time between the vehicle and the vehicle approaching from the blind spot; when the vehicle is in the blind spot before the curve at a distance greater than a first preset distance from the curve center, determining a first vehicle control strategy based on the vehicle type and the collision time; when the vehicle is in the blind spot in the curve at a distance less than or equal to the first preset distance from the curve center, determining a second vehicle control strategy based on the vehicle type and the collision time; and controlling the vehicle's speed and lateral position according to the first vehicle control strategy or the second vehicle control strategy.

[0013] In some embodiments of this application, based on the foregoing scheme, determining the first vehicle control strategy according to the vehicle type and the collision time includes: when the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than a first time threshold, controlling the vehicle to smoothly decelerate to a first speed threshold and controlling the vehicle to laterally deviate a first distance within the lane; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold and greater than or equal to a second time threshold, controlling the vehicle to smoothly decelerate to a second speed threshold and controlling the vehicle to laterally deviate a second distance within the lane, wherein the second speed threshold is less than the first speed threshold and the second distance is greater than the first distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, controlling the vehicle to smoothly decelerate to a third speed threshold and controlling the vehicle to laterally deviate a third distance within the lane, wherein the third speed threshold is less than the second speed threshold and the third distance is greater than the second distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the third time threshold, controlling the vehicle to brake to a stop until the vehicle in the blind spot has left before starting again, wherein the third time threshold is less than the second time threshold. The system implements the following parameters: When a large vehicle approaches from the blind spot and the collision time exceeds a first time threshold, the vehicle is controlled to smoothly decelerate to a fourth speed threshold and then laterally deviate a fourth distance within the lane. The fourth speed threshold is less than the first speed threshold, and the fourth distance is greater than the first distance. When a large vehicle approaches from the blind spot and the collision time is less than or equal to the first time threshold and greater than or equal to the second time threshold, the vehicle is controlled to smoothly decelerate to a fifth speed threshold and then laterally deviate a fifth distance within the lane. The fifth speed threshold is less than the second speed threshold and less than the fourth speed threshold, and the fifth distance is greater than the second distance and greater than the fourth distance. When a large vehicle approaches from the blind spot and the collision time is less than the second time threshold, the vehicle is controlled to smoothly decelerate to a sixth speed threshold and then laterally deviate a third distance within the lane. The sixth speed threshold is less than the third speed threshold and less than the fifth speed threshold, and the third distance is greater than the fifth distance. When a large vehicle approaches from the blind spot and the collision time is less than the third time threshold, the vehicle is controlled to brake to a stop until the vehicle in the blind spot has left before starting again.

[0014] In some embodiments of this application, based on the foregoing scheme, determining the second vehicle control strategy according to the vehicle type and the collision time includes: when the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than a first time threshold, controlling the vehicle to smoothly decelerate to a first speed threshold, and controlling the vehicle to laterally deviate a fourth distance within the lane; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold, and greater than or equal to a second time threshold, controlling the vehicle to smoothly decelerate to a second speed threshold, and controlling the vehicle to laterally deviate a fifth distance within the lane, wherein the second speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, controlling the vehicle to smoothly decelerate to a third speed threshold, and controlling the vehicle to laterally deviate a third distance within the lane, wherein the third speed threshold is less than the second speed threshold, and the third distance is greater than the fifth distance. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than a third time threshold, the vehicle is controlled to brake to a stop until the vehicle leaves the blind spot before starting again. The third time threshold is less than the second time threshold. When the vehicle approaching from the blind spot is a large vehicle and the collision time is greater than a first time threshold, the vehicle is controlled to smoothly decelerate to a seventh speed threshold and to laterally deviate a fifth distance within the lane. The seventh speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. When the vehicle approaching from the blind spot is a large vehicle and the collision time is less than or equal to the first time threshold, the vehicle is controlled to brake to a stop and to laterally deviate a sixth distance within the lane. The sixth distance is greater than the third distance. When the vehicle approaching from the blind spot is a large vehicle and the collision time is less than a fourth time threshold, the vehicle is controlled to brake to a stop and then reverse a preset distance until the vehicle leaves the blind spot before starting again. The fourth time threshold is less than the second time threshold and greater than the third time threshold.

[0015] In some embodiments of this application, based on the foregoing scheme, the method further includes: updating external audio data, visual sensing data, and lidar data in real time during the process of controlling the vehicle's movement; updating multimodal fusion features based on the updated external audio data, visual sensing data, and lidar data; re-identifying vehicles approaching from blind spots based on the updated multimodal fusion features; and dynamically adjusting the vehicle control strategy according to the re-identification results.

[0016] According to a second aspect of the embodiments of this application, a vehicle autonomous driving control device is provided. The device includes: an acquisition unit, configured to acquire external audio data, visual sensing data, and lidar data collected during vehicle operation; a processing unit, configured to preprocess the external audio data to obtain audio spatial features; and to convert the visual sensing data into visual bird's-eye view features and the lidar data into lidar bird's-eye view features; a fusion unit, configured to fuse the audio spatial features, the visual bird's-eye view features, and the lidar bird's-eye view features to obtain multimodal fusion features; and a control unit, configured to identify vehicles approaching in blind spots based on the multimodal fusion features and control the vehicle's operation according to the identification results.

[0017] According to a third aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor to cause a computer device having the processor to perform an operation as described in any of the first aspects above.

[0018] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one computer program instruction, the at least one computer program instruction being loaded and executed by a processor to perform the operation as described in any of the first aspects above.

[0019] According to a fifth aspect of the embodiments of this application, a vehicle is provided, the vehicle including one or more processors and one or more memories, the one or more memories storing at least one computer program instruction, the at least one computer program instruction being loaded and executed by the one or more processors to perform the operation as described in any of the first aspects above.

[0020] Based on the technical solution proposed in this application, by simultaneously collecting three types of data—external audio, visual sensing, and LiDAR—and extracting audio spatial features and visual and LiDAR bird's-eye view features respectively, multimodal fusion is performed. This effectively compensates for the deficiencies of single sensors, solving the blind spot perception problems of visual perception being easily obstructed by obstacles and LiDAR having detection blind spots. Audio signals can penetrate obstructions to capture the sound source information of vehicles approaching in blind spots and locate their direction. The three types of features are mutually verified and complementary, which can significantly improve the accuracy of blind spot vehicle recognition and reduce the occurrence of missed and false detections. Relying on high-precision multimodal fusion features to complete blind spot vehicle recognition and adjust the vehicle's driving status accordingly, it can promptly perform safety operations such as deceleration, lane change prohibition, and avoidance, making up for the shortcomings of lateral and forward blind spot perception in autonomous driving, improving the stability and reliability of autonomous driving decisions, effectively avoiding collision risks in scenarios such as turning and lane changing, and significantly improving the driving safety of autonomous vehicles. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 A flowchart of a vehicle autonomous driving control method according to an embodiment of this application is shown; Figure 2 A block diagram of a vehicle autonomous driving control device according to an embodiment of this application is shown; Figure 3 A schematic diagram of the vehicle structure in an embodiment of this application is shown. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0024] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. It should also be noted that, for the sake of simplicity, certain components in the drawings that do not affect the interpretation of the technical solution of this application have been appropriately omitted.

[0025] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined. Therefore, the actual execution order may change depending on the actual situation.

[0026] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "multiple" means two or more.

[0027] This application proposes a vehicle autonomous driving control scheme, the purpose of which is to improve the safety of vehicle autonomous driving control.

[0028] The core concept of this application is to use the external microphones that are standard on L3 and above autonomous vehicles as a supplementary perception means. Effective sound information is extracted through audio preprocessing technology and then fused with the bird's eye view (BEV) features of vision and LiDAR. This enables early perception and recognition of vehicles in blind spots that cannot be covered by vision and radar. Then, based on the vehicle's position, the type of vehicle approaching, and the collision risk level, differentiated control strategies are output, which fundamentally solves the perception failure problem of traditional autonomous driving systems in occluded scenarios and significantly improves the safety margin of autonomous driving.

[0029] Next, this application will elaborate on the proposed vehicle autonomous driving control scheme. (Refer to...) Figure 1 The flowchart of a vehicle autonomous driving control method according to an embodiment of this application is shown. This method can be executed by a device with computing processing capabilities, such as... Figure 1 As shown, the method includes at least steps 110 to 140, which are described in detail below: In step 110, external audio data, visual sensor data, and lidar data collected during vehicle operation are acquired.

[0030] In this application, the external audio data can be collected by multiple external microphones installed on the front and rear bumpers, rearview mirrors, or roof of the vehicle. In a preferred embodiment, the vehicle is equipped with at least a first front microphone and a second rear microphone, which are symmetrically arranged along the longitudinal centerline of the vehicle, with a spacing of 2 to 4 meters to ensure the accuracy of sound source localization. The visual sensing data can be collected by visual sensors such as monocular or binocular cameras or surround-view cameras installed above the windshield of the vehicle, mainly used to acquire image information of the road environment, including lane lines, traffic signs, and visible obstacles. The lidar data can be collected by lidar sensors installed on the roof or around the vehicle body to acquire three-dimensional point cloud information of the surrounding environment, enabling precise measurement of the distance, size, and shape of obstacles.

[0031] This application simultaneously collects data from three different modalities, which can give full play to the advantages of each sensor. The visual sensor is good at recognizing the color and texture features of objects, the lidar is good at accurately measuring the spatial position and geometry of objects, and the audio sensor can detect targets in blind spots through sound signals when the visual and lidar are blocked. The three complement each other to achieve full-scene environmental perception.

[0032] Continue to refer to Figure 1 In step 120, the external audio data is preprocessed to obtain audio spatial features. The visual sensing data is then converted into visual bird's-eye view features, and the lidar data is converted into lidar bird's-eye view features.

[0033] In this application, a significant amount of vehicle noise is generated during vehicle operation, such as tire noise, wind noise, and engine noise. Furthermore, sound propagation is affected by reflections from obstacles such as mountains, buildings, and vegetation, resulting in reverberation. These interferences severely impact the accuracy of target sound recognition and localization. Therefore, a series of preprocessing operations are required on the external audio data to remove interference and extract audio spatial features that reflect the spatial location and type of the target sound.

[0034] In this application, the visual bird's-eye view feature is a three-dimensional feature map obtained by converting a two-dimensional perspective image captured by a camera into a top-down view from the air. This feature map can uniformly represent the spatial positional relationships of all objects around the vehicle, facilitating fusion with LiDAR data and audio data. The LiDAR bird's-eye view feature is a feature map obtained by projecting three-dimensional point cloud data collected by LiDAR onto the bird's-eye view plane. This feature map can accurately reflect the position, size, and motion state of obstacles.

[0035] In this application, by unifying visual and laser data into bird's-eye view features, the differences between different sensor coordinate systems can be eliminated, laying the foundation for subsequent multimodal fusion. At the same time, the bird's-eye view can intuitively display the global environment around the vehicle, making it more suitable for autonomous driving systems to perform path planning and decision control.

[0036] In such Figure 1 In step 120, the preprocessing of the external audio data to obtain audio spatial features can be performed according to steps 121 to 123 as follows: Step 121: Perform adaptive filtering on the external audio data to remove vehicle noise and obtain a preliminary denoised frequency.

[0037] Step 122: Perform weighted dreverberation processing on the initial noise-reduced frequency to remove the reverberation caused by sound reflection and obtain the direct sound frequency.

[0038] Step 123: Perform sound source localization processing on the direct sound audio to determine the incident angle of the sound and obtain the audio spatial characteristics.

[0039] In this application, the adaptive filtering is a signal processing technique that can automatically adjust filter parameters based on the statistical characteristics of the input signal, effectively removing non-stationary vehicle noise. Compared with traditional fixed-parameter filters, adaptive filters can track changes in vehicle noise in real time, such as changes in engine noise during vehicle acceleration and changes in tire noise on different road surfaces, thereby achieving better noise reduction results.

[0040] In this application, reverberation refers to the superimposed signal formed after sound propagates through a closed or semi-closed space and undergoes multiple reflections, which leads to unclear sound signals and seriously affects the accuracy of sound source localization. In this application, the WPE dereverberation algorithm can be used to remove early reflections and late reverberation, preserving a clear direct sound signal.

[0041] In this application, sound source localization refers to determining the spatial location of a sound source by using the time difference, intensity difference, or phase difference of the same sound signal received by multiple microphones. In this application, the time difference between two microphones (front and rear) can be used to calculate the incident angle of the sound, enabling precise localization of sound sources within a 180-degree range in front of the vehicle, with a localization accuracy within ±5 degrees.

[0042] In this application, through the above three preprocessing operations, clear and accurate audio spatial features can be extracted from noisy external audio data, providing reliable input for subsequent multimodal fusion and blind spot vehicle recognition, thereby solving the problems of poor anti-interference ability and low positioning accuracy of traditional audio perception technology in the vehicle environment.

[0043] In step 121 above, the adaptive filtering of the external audio data to remove vehicle noise and obtain a preliminary denoised frequency can be performed according to steps 1211 to 1214 as follows: Step 1211: Collect first audio data containing the target sound and vehicle noise through the first microphone at the front of the vehicle.

[0044] Step 1212: Collect second audio data, mainly containing vehicle noise, through a second microphone at the rear of the vehicle.

[0045] Step 1213: Generate a vehicle noise estimate based on the second audio data.

[0046] Step 1214: Subtract the vehicle noise estimate from the first audio data to obtain the preliminary denoised frequency.

[0047] In this application, the first microphone can be installed at the front of the vehicle, such as in the middle of the front bumper, facing directly forward. It can preferentially collect target sounds in front of the vehicle, such as horns and engine sounds from oncoming vehicles, while also collecting noise signals such as tire noise and wind noise from the vehicle itself. The first audio data can be represented as:

[0048] in, For the target sound signal, This is a vehicle noise signal. These are the time sampling points.

[0049] In this application, the second microphone can be installed at the rear of the vehicle, such as in the middle of the rear bumper. Because it is far from the target sound source in front of the vehicle and is obstructed by the vehicle body, the collected signal mainly contains its own vehicle noise, while the target sound signal is very weak. Therefore, the second audio data can be approximated as containing only its own vehicle noise. The second audio data can be represented as:

[0050] In this application, the Normalized Least Mean Square (NLMS) adaptive filtering algorithm can be used to generate vehicle noise estimates. The NLMS algorithm is a classic adaptive filtering algorithm with advantages such as low computational cost, fast convergence speed, and good stability, making it very suitable for real-time signal processing scenarios in vehicles. The vehicle noise estimate can be calculated using the following formula:

[0051] in, Let be the weight vector of the adaptive filter. The input vector for the second audio data contains sample values ​​from the current time and several previous time points.

[0052] In this application, the initial denoised frequency is the estimated value of the target sound signal obtained after adaptive filtering, which can be calculated using the following formula:

[0053] The weight vector of the NLMS algorithm is updated in real time based on the error signal, and the update formula is:

[0054] in, The step size parameter controls the convergence speed and steady-state error of the algorithm. In automotive applications, The value can range from 0.01 to 0.1; It is a very small positive number used to prevent the denominator from being zero, thus ensuring the numerical stability of the algorithm.

[0055] In this application, by adopting a dual-microphone NLMS adaptive filtering scheme, various noises generated by the vehicle can be effectively removed, leaving only the target sound signal from outside the vehicle. The processing delay is only at the millisecond level, which can fully meet the real-time requirements of the autonomous driving system, thus providing clean audio input for subsequent sound source localization and target recognition.

[0056] In step 122 above, the weighted denoising process performed on the initially denoised frequency to remove reverberation caused by sound reflections and obtain the direct sound frequency can be executed according to steps 1221 to 1223 as follows: Step 1221: Obtain the historical continuous multi-frame signal of the preliminary denoised frequency.

[0057] Step 1222: Based on the weighted prediction of the reverberation component of the current frame using the historical consecutive frames of signal, predict the reverberation component of the current frame.

[0058] Step 1223: Subtract the reverberation component from the current frame signal of the preliminary denoised frequency to obtain the direct audio frequency.

[0059] In a preferred embodiment of this application, the past 10 consecutive frames of the initial denoised frequency signal can be obtained. The duration of each frame signal can be set to 10 to 20 milliseconds, and there can be 50% overlap between frames to ensure the continuity of the signal.

[0060] In this application, it should be explained that the core idea of ​​the WPE dereverberation algorithm is to use past signal frames to predict the reverberation component of the current frame. Since the reverberation signal is formed after multiple reflections of the direct sound signal, the reverberation component of the current frame has a linear relationship with the signals of multiple past frames. By minimizing the prediction error, the optimal prediction weight can be obtained, thereby accurately predicting the reverberation component of the current frame.

[0061] After WPE dreverberation processing, the resulting direct audio audio only contains the sound signal that travels directly from the sound source to the microphone, removing all reverberation signals generated by reflections, thus significantly improving the clarity and intelligibility of the sound.

[0062] In this application, weighted dereverberation processing using WPE can effectively eliminate the reflection effects of obstacles such as mountains, buildings, and green belts on sound, and prevent the algorithm from misidentifying the reverberation signal as multiple sound sources, thereby improving the accuracy of sound source localization and the reliability of target recognition. Especially in scenarios with severe reverberation such as mountain bends and tunnels, it can significantly improve the performance of the system.

[0063] In step 123 above, the sound source localization processing of the direct sound audio to determine the incident angle of the sound can be performed according to the following steps 1231 to 1233: Step 1231: Calculate the time difference between the arrival of the direct audio signal at the first microphone at the front of the vehicle and the second microphone at the rear.

[0064] Step 1232: Calculate the incident angle of the sound based on the time difference and the distance between the two microphones.

[0065] Step 1233: Perform sliding window smoothing on the calculated incident angle and output a stable incident angle as part of the audio spatial features.

[0066] In this application, the time difference between the arrival of the same sound signal at two microphones can be calculated using a cross-correlation function. The cross-correlation function measures the similarity between two signals; it reaches its maximum value when the two signals are perfectly aligned, and the corresponding time offset is the time difference between the two microphones. The formula for calculating the cross-correlation function is:

[0067] in, The direct audio signal collected by the first microphone. The direct audio signal collected by the second microphone. This represents the time offset. Cross-correlation function. When the maximum value is obtained, the corresponding The value is the time difference between the two microphones. .

[0068] Assuming sound propagates as a plane wave, the distance between the two microphones is... The speed of sound in air is (Typically taken as 340 m / s), the angle of incidence of the sound is... (Defined as the angle between the direction of sound propagation and the longitudinal centerline of the vehicle, ranging from -90 degrees to 90 degrees), then the time difference With the angle of incidence The relationship between them is:

[0069] The formula for calculating the incident angle can be obtained through deformation:

[0070] Due to the influence of noise and signal fluctuations, the incident angle calculated in a single step may contain significant errors. Therefore, a sliding window smoothing method can be used to average or weighted average the incident angles calculated over multiple consecutive time points to obtain a more stable and accurate angle estimate. The size of the sliding window can be set to 5 to 10 frames, which can be adjusted according to the actual application scenario.

[0071] In this application, by using a sound source localization method based on the time difference of dual microphones, the incident angle of the target sound can be accurately calculated, thereby determining the approximate direction of the vehicle approaching in the blind spot. This provides key spatial information for subsequent multimodal fusion and target recognition, enabling the autonomous driving system to perceive the presence and position of the vehicle in the blind spot in advance even when vision and radar are completely ineffective.

[0072] Continue to refer to Figure 1 In step 130, the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features are fused to obtain multimodal fusion features.

[0073] In this application, multimodal fusion refers to integrating information from different modalities collected by different sensors to obtain a more comprehensive, accurate, and reliable environmental perception result than single-modal information. In this application, a Transformer-based multimodal fusion architecture can be adopted, which can effectively mine the complementarity and correlation between different modal features, achieving accurate perception of complex environments.

[0074] This application integrates audio spatial features with visual and laser bird's-eye view features, which can fully utilize the advantages of audio sensors in blind spot perception and make up for the shortcomings of visual and radar sensors in occluded scenarios, thereby achieving true blind spot-free perception and significantly improving the perception capability and safety of autonomous driving systems in complex environments.

[0075] In such Figure 1 In step 130, the fusion of the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features to obtain multimodal fusion features can be performed according to the following steps 131 to 132: Step 131: Convert the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features into feature tokens of a unified format, and add spatiotemporal location encoding to each feature token.

[0076] Step 132: The feature tokens are weighted and fused using an attention mechanism to obtain the multimodal fused features.

[0077] In this application, in order to achieve the fusion of different modal features, it is necessary to map the visual bird's-eye view features, laser bird's-eye view features and audio spatial features into the same feature space and convert them into feature token sequences with the same dimension.

[0078] The spatiotemporal location encoding is used to mark the spatial location and temporal information corresponding to each feature token, enabling the Transformer model to understand the spatial and temporal relationships between different features. The spatial location encoding can be generated based on the coordinates of the feature in the bird's-eye view, while the temporal location encoding can be generated based on the timestamp of feature acquisition.

[0079] The attention mechanism is the core of the Transformer model, enabling it to automatically learn the importance weights among different features and focus on features more valuable to the current task. During multimodal fusion, the attention mechanism automatically adjusts the weights of different modal features based on the confidence level of each modality's data. When the confidence level of visual and radar data is high, it relies more on visual and laser features; when the confidence level of visual and radar data is low due to occlusion, it relies more on audio features.

[0080] In this application, by converting features from different modalities into unified feature tokens and adding spatiotemporal location encoding, and then using the attention mechanism of Transformer for weighted fusion, deep fusion of features from different modalities can be achieved, giving full play to the advantages of each modality, thereby obtaining more comprehensive, accurate and robust multimodal fusion features, providing a solid foundation for subsequent blind spot vehicle recognition.

[0081] In step 132 above, the weighted fusion of the feature tokens using an attention mechanism can be performed according to steps 1321 to 1322 as follows: Step 1321: Assign different weights to the corresponding feature tokens according to the confidence level of each modality data, so that the feature tokens corresponding to the laser bird's-eye view features focus on the texture information of the feature tokens corresponding to the visual bird's-eye view features, the feature tokens corresponding to the visual bird's-eye view features focus on color and texture information, and the feature tokens corresponding to the audio spatial features focus on the sound information in the blind zone direction.

[0082] Step 1322: Based on the assigned weights and the attention results of each feature token, the multimodal fusion feature is obtained.

[0083] In this application, the confidence level of each modal data can be evaluated based on the sensor's operating state, environmental conditions, and the quality of the features. For example, in a clear, unobstructed environment, the confidence level of visual and laser data is relatively high; in rainy, foggy, or obstructed environments, the confidence level of visual and laser data decreases, while the confidence level of audio data is relatively high.

[0084] By guiding feature tokens from different modalities to focus on the information they excel at, the quality of fused features can be further improved. LiDAR excels at acquiring the three-dimensional geometric information of objects, so it can be guided to focus on the texture information in visual features to better identify the object category; visual sensors excel at acquiring color and texture information, so they can be guided to focus on their own color and texture features to distinguish different objects; audio sensors excel at perceiving sound information in blind areas, so they can be guided to focus on sound features in the direction of the blind area to detect targets in the blind area in advance.

[0085] In this application, all feature tokens are weighted and summed according to their assigned weights to obtain the final multimodal fusion feature. The multimodal fusion feature contains information from three modalities: visual, laser, and audio, and can comprehensively reflect the environmental conditions around the vehicle, including visible areas and invisible blind spots.

[0086] This application utilizes an attention-based multimodal fusion method to achieve adaptive fusion of features from different modalities. This allows the system to automatically adjust the contribution of each modality according to different environmental conditions, thereby maintaining excellent perception performance in various complex scenarios. In particular, in blind spot scenarios where vision and radar are obstructed, the system can accurately identify oncoming vehicles in the blind spot through audio features, effectively avoiding collisions.

[0087] Continue to refer to Figure 1 In step 140, vehicles approaching from blind spots are identified based on the multimodal fusion features, and the vehicle is controlled to move according to the identification results.

[0088] In this application, blind spot vehicle recognition is performed based on multimodal fusion features. This method comprehensively utilizes information from three modalities—visual, laser, and audio—to achieve accurate detection, classification, and tracking of vehicles approaching from blind spots. Based on the recognition results, corresponding control strategies are output, enabling differentiated control measures to be implemented according to different risk levels. This maximizes traffic efficiency and passenger comfort while ensuring safety.

[0089] This application combines multimodal perception with intelligent control to achieve a complete closed loop from perception to decision-making to control, enabling the autonomous driving system to proactively respond to blind spot oncoming traffic scenarios, transforming from the traditional "reacting only after seeing" to "predicting and proactively avoiding in advance," thereby significantly improving the safety and reliability of the autonomous driving system.

[0090] In such Figure 1 In step 140 shown, the process of identifying vehicles approaching from blind spots based on the multimodal fusion features can be performed according to steps 141 to 143 as follows: Step 141: Perform three-dimensional target detection based on the multimodal fusion features.

[0091] Step 142: When a vehicle is detected in a blind spot not covered by visual sensing data and lidar data, the recognition of the vehicle in the blind spot is enhanced by combining audio spatial features.

[0092] Step 143: Track the trajectory of vehicles approaching from the blind spot and predict their direction and speed.

[0093] In this application, the 3D target detection refers to detecting all target objects in the surrounding environment from multimodal fusion features and determining their 3D position, size, orientation, and motion state. In this application, a Transformer-based 3D target detection algorithm can be used to directly detect targets on the multimodal fusion features, enabling simultaneous processing of targets in both visible and blind areas.

[0094] When a target is located in the blind zone of both vision and radar, and there is no information about the target in the visual and laser signatures, then the audio spatial signature becomes the sole basis for identifying the target. Based on information such as the incident angle, sound frequency, and intensity in the audio spatial signature, the existence, type, and approximate location of the target can be further confirmed, and a virtual target can be generated at the corresponding blind zone location in the bird's-eye view.

[0095] By combining multimodal fusion features from multiple consecutive frames, vehicles approaching from blind spots can be tracked, and their motion trajectory models can be established. Based on historical trajectory information, the future driving direction and speed of vehicles approaching from blind spots can be predicted, thereby calculating the time to collision (TTC) between the vehicle and the vehicle approaching from the blind spot, providing a basis for subsequent decision-making and control.

[0096] In this application, through the above-mentioned three-dimensional target detection, audio enhancement recognition and trajectory tracking steps, accurate, stable and continuous perception of vehicles approaching from blind spots can be achieved. Even if the target is in the blind spot for a long time, it will not lose tracking, thereby providing sufficient reaction time for the autonomous driving system and effectively avoiding collision accidents caused by vehicles suddenly appearing in the blind spot.

[0097] In such Figure 1 In step 140 shown, controlling the vehicle's movement based on the recognition result can be performed according to steps 144 to 147 as follows: Step 141: Determine the vehicle type of the vehicle approaching from the blind spot and the collision time between the vehicle and the vehicle approaching from the blind spot.

[0098] Step 142: When the vehicle is in the blind spot before the curve, which is more than a first preset distance from the center of the curve, a first vehicle control strategy is determined based on the vehicle type and the collision time.

[0099] Step 143: When the vehicle is in a blind spot curve at a distance less than or equal to the center of the curve from the first preset distance, a second vehicle control strategy is determined based on the vehicle type and the collision time.

[0100] Step 144: Control the vehicle's speed and lateral position according to the first vehicle control strategy or the second vehicle control strategy.

[0101] In this application, vehicle types can include small vehicles and large vehicles. Large vehicles include trucks, buses, etc. Due to their large size, long braking distance, and high inertia, the consequences of a collision would be more severe, thus requiring a more conservative control strategy. Collision time refers to the time required for a vehicle to collide with a vehicle in the blind spot at the current speed and direction of travel of both parties, and is a core indicator for measuring the collision risk level.

[0102] In some embodiments, the first preset distance can be set to 50 meters, that is, when the distance between the vehicle and the center of the curve is greater than 50 meters, the vehicle is considered to be in the position before the curve, at which time the vehicle has relatively sufficient time and space to take measures such as deceleration and avoidance.

[0103] When a vehicle is less than or equal to 50 meters from the center of a curve, it is considered to have entered the curve. At this point, the vehicle's driving space is limited, the difficulty of operation increases, and the risk of collision is higher. Therefore, more stringent control strategies are required.

[0104] Vehicle speed control can be achieved by controlling the engine's output power and the braking force of the braking system, while lateral position control can be achieved by controlling the steering angle of the steering system. During control, it is essential to ensure smooth vehicle operation and avoid sudden acceleration, deceleration, and sharp turns to enhance ride comfort.

[0105] In this application, by determining differentiated control strategies based on vehicle location, oncoming vehicle type, and collision time, refined control of blind spot oncoming traffic scenarios can be achieved. Under the premise of ensuring safety, the impact on normal vehicle operation can be minimized as much as possible, balancing safety and traffic efficiency, while improving the passenger riding experience.

[0106] In step 141 above, determining the vehicle type of the vehicle approaching from the blind spot can be performed as follows: step 1411: Step 1411: Based on the sound frequency and intensity features in the audio spatial features, identify the vehicle type of the vehicle approaching from the blind spot as a small vehicle or a large vehicle.

[0107] In this application, different types of vehicles emit sounds with different frequency and intensity characteristics. Small vehicles have higher engine frequencies and relatively lower intensity, while large vehicles have lower engine frequencies and relatively higher intensity, especially large trucks with diesel engines, whose sounds exhibit significant low-frequency characteristics. By extracting features such as Mel-Frequency Cepstral Coefficients (MFCCs) from the audio signals and inputting them into a pre-trained classification model, the vehicle type approaching from the blind spot can be accurately identified.

[0108] In this application, vehicle type identification is based on audio features, without relying on visual and radar data. Therefore, even when the vehicle is completely obscured, it can accurately distinguish between small and large vehicles, thus providing a basis for subsequent differentiated control strategies. In particular, adopting a more conservative control strategy for large vehicles can effectively reduce the probability of major traffic accidents.

[0109] In some embodiments of step 142 above, determining the first vehicle control strategy based on the vehicle type and the collision time may include the following: In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time exceeds a first time threshold, the vehicle is controlled to smoothly decelerate to a first speed threshold, and the vehicle is controlled to laterally deviate a first distance within the lane. The first time threshold can be set to 6 seconds, the first speed threshold can be set to 40 kilometers per hour, and the first distance can be set to 0.2 meters. At this time, the collision risk is low, and the vehicle only needs to decelerate appropriately and move slightly to the right, leaving sufficient safety space for possible oncoming vehicles.

[0110] In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time is less than or equal to a first time threshold and greater than or equal to a second time threshold, the vehicle is controlled to smoothly decelerate to the second speed threshold, and the vehicle is controlled to laterally deviate a second distance within the lane. The second speed threshold is less than the first speed threshold, and the second distance is greater than the first distance. The second time threshold can be set to 3 seconds, the second speed threshold can be set to 25 kilometers per hour, and the second distance can be set to 0.4 meters. At this point, the collision risk is moderate, and the vehicle needs to further reduce its speed and increase the lateral deviation to improve the safety margin.

[0111] In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time is less than a second time threshold, the vehicle is controlled to smoothly decelerate to a third speed threshold, and the vehicle is controlled to laterally deviate a third distance within the lane. The third speed threshold is less than the second speed threshold, and the third distance is greater than the second distance. The third speed threshold can be set to 15 kilometers per hour, and the third distance can be set to 0.6 meters. At this time, the collision risk is high, and the vehicle needs to significantly reduce its speed and drive as far to the right as possible to avoid a collision with an oncoming vehicle.

[0112] In this application, when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than a third time threshold, the vehicle is controlled to brake to a stop until the vehicle in the blind spot has completely moved away before starting again. The third time threshold is less than the second time threshold. The third time threshold can be set to 1 second. At this time, the risk of collision is extremely high, and the vehicle needs to brake to a stop immediately, waiting for the vehicle in the blind spot to completely move away before continuing to drive, to ensure absolute safety.

[0113] In this application, when a large vehicle approaches from the blind spot and the collision time exceeds a first time threshold, the vehicle is controlled to smoothly decelerate to a fourth speed threshold, and then laterally deviates a fourth distance within the lane. The fourth speed threshold is less than the first speed threshold, and the fourth distance is greater than the first distance. The fourth speed threshold can be set to 30 kilometers per hour, and the fourth distance can be set to 0.3 meters. Because large vehicles pose a higher risk, even in low-risk situations, a more conservative control strategy is required compared to that for smaller vehicles.

[0114] In this application, when the oncoming vehicle in the blind spot is a large vehicle and the collision time is less than or equal to a first time threshold and greater than or equal to a second time threshold, the vehicle is controlled to smoothly decelerate to a fifth speed threshold, and the vehicle is controlled to laterally deviate a fifth distance within the lane. The fifth speed threshold is less than the second speed threshold and less than the fourth speed threshold, and the fifth distance is greater than the second distance and greater than the fourth distance. The fifth speed threshold can be set to 20 kilometers per hour, and the fifth distance can be set to 0.5 meters. At this point, the collision risk is moderate, and it is necessary to further reduce the vehicle speed and increase the lateral deviation.

[0115] In this application, when a large vehicle approaches from the blind spot and the collision time is less than a second time threshold, the vehicle is controlled to smoothly decelerate to a sixth speed threshold, and the vehicle is controlled to laterally deviate a third distance within the lane. The sixth speed threshold is less than the third speed threshold and less than the fifth speed threshold, and the third distance is greater than the fifth distance. The sixth speed threshold can be set to 10 kilometers per hour. At this point, the collision risk is high, requiring the vehicle speed to be reduced to an even lower level and the vehicle to travel as far to the right as possible.

[0116] In this application, when the oncoming vehicle in the blind spot is a large vehicle and the collision time is less than a third time threshold, the vehicle is controlled to come to a complete stop until the oncoming vehicle in the blind spot has moved away before starting again. This indicates that the risk of collision is extremely high at this point, and the vehicle must stop immediately to yield and ensure that a collision with the large vehicle is not possible.

[0117] This application develops refined speed and lateral position control strategies for different types of oncoming vehicles and different collision risk levels in the scenario before a curve. This can maximize traffic efficiency while ensuring safety, and at the same time avoid sudden braking, sharp turns and other violent operations, improve ride comfort, and prevent passengers from feeling motion sickness or panic.

[0118] In some embodiments of step 143 above, determining the second vehicle control strategy based on the vehicle type and the collision time may include the following: In this application, when the vehicle approaching from the blind spot is a small vehicle and the collision time exceeds a first time threshold, the vehicle is controlled to smoothly decelerate to a first speed threshold, and the vehicle is controlled to laterally deviate a fourth distance within the lane. At this time, the vehicle is in a curve, and the driving space is limited. Therefore, even in a low-risk situation, a larger lateral deviation is required than in the scenario before the curve to reserve more safety space.

[0119] In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time is less than or equal to a first time threshold and greater than or equal to a second time threshold, the vehicle is controlled to smoothly decelerate to a second speed threshold, and the vehicle is controlled to laterally deviate a fifth distance within the lane. The second speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. At this point, the collision risk is moderate, and it is necessary to further reduce the vehicle speed and increase the lateral deviation.

[0120] In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time is less than a second time threshold, the vehicle is controlled to smoothly decelerate to a third speed threshold, and the vehicle is controlled to laterally deviate a third distance within the lane. The third speed threshold is less than the second speed threshold, and the third distance is greater than the fifth distance. At this time, the collision risk is high, requiring a significant reduction in vehicle speed and driving as far to the right as possible.

[0121] In this application, when the oncoming vehicle in the blind spot is a small vehicle and the collision time is less than a third time threshold, the vehicle is controlled to brake to a stop until the oncoming vehicle in the blind spot has moved away before starting again. The third time threshold is less than the second time threshold. At this time, the risk of collision is extremely high, and the vehicle must stop immediately to yield.

[0122] In this application, when a large vehicle approaches from the blind spot and the collision time exceeds a first time threshold, the vehicle is controlled to smoothly decelerate to a seventh speed threshold, and the vehicle is controlled to laterally deviate a fifth distance within the lane. The seventh speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. The seventh speed threshold can be set to 20 kilometers per hour. Because large vehicles are wide, they occupy more space in curves, thus requiring a more conservative control strategy.

[0123] In this application, when a large vehicle approaches from the blind spot and the collision time is less than or equal to a first time threshold, the vehicle is controlled to brake to a stop, and then laterally deviated by a sixth distance within the lane, which is greater than the third distance. The sixth distance can be set to 0.8 meters. At this point, the collision risk is high, and the vehicle needs to stop immediately and move as far to the right as possible to allow sufficient passage for the large vehicle.

[0124] In this application, when the oncoming vehicle in the blind spot is a large vehicle and the collision time is less than a fourth time threshold, the vehicle is controlled to brake to a stop and then reverse a preset distance until the oncoming vehicle in the blind spot has left before starting again. The fourth time threshold is less than the second time threshold and greater than the third time threshold. The fourth time threshold can be set to 2 seconds, and the preset distance can be set to 1 to 2 meters. In particularly urgent situations, the vehicle can reverse appropriately after braking to further increase the safe distance between itself and the large vehicle, avoiding scraping or collision accidents.

[0125] This application develops more stringent and conservative control strategies for different types of oncoming vehicles and different collision risk levels in curve scenarios. It can effectively address issues such as limited driving space, poor visibility, and high operational difficulty in curves. In particular, it adds control measures for reversing to yield to large vehicles, which can minimize collision risks and ensure driving safety even in the most dangerous situations.

[0126] In this application, the following step 151 may also be performed: Step 151: When a vehicle is detected in the blind spot, display a blind spot vehicle warning message through the vehicle's human-machine interface; and / or, play a blind spot vehicle warning voice message through the vehicle's voice broadcast system.

[0127] In this application, the human-machine interface can be a vehicle's central control display screen, instrument panel, or head-up display (HUD) system. The prompt information can include the direction, type, and risk level of vehicles approaching from the blind spot. The voice broadcast system can remind the driver to pay attention to vehicles approaching from the blind spot via voice when the risk level is high, such as broadcasting "Please note that there is a vehicle in the left front blind spot."

[0128] By providing drivers with blind spot warnings through a human-machine interface and voice broadcast system, drivers can be aware of the surrounding environment and take over control of the vehicle when necessary, thereby improving driving safety and enhancing drivers' trust in the autonomous driving system.

[0129] In this application, steps 152 to 154 may also be performed: Step 152: During the process of controlling the vehicle's movement, update the external audio data, visual sensor data, and lidar data in real time.

[0130] Step 153: Update the multimodal fusion features based on the updated external audio data, visual sensing data, and LiDAR data.

[0131] Step 154: Re-identify vehicles approaching from the blind spot based on the updated multimodal fusion features, and dynamically adjust the vehicle control strategy according to the re-identification results.

[0132] In this application, the data update frequency can be set to 10 Hz to 30 Hz to ensure that the system can perceive changes in the vehicle's surrounding environment in real time. Whenever new data is input, the preprocessing, feature transformation, and feature fusion steps can be re-executed to generate the latest multimodal fusion features.

[0133] In this application, the position, speed, and type information of oncoming vehicles in the blind spot can be updated in real time based on the latest identification results, and the collision time and risk level can be recalculated. This allows for dynamic adjustment of the vehicle's speed and lateral position to adapt to constantly changing environmental conditions. By updating data in real time and dynamically adjusting the control strategy, closed-loop control of blind spot oncoming vehicle scenarios can be achieved, ensuring that the system can react promptly to changes in the environment and always maintain the optimal control state, thereby further improving the safety and reliability of the autonomous driving system.

[0134] Based on the vehicle autonomous driving control method proposed in this application, by utilizing the vehicle's standard external microphones as a supplementary perception means, and combining advanced audio preprocessing technology with Transformer-based multimodal fusion technology, it can achieve early perception and accurate identification of vehicles approaching from visual and radar blind spots, and output differentiated control strategies based on vehicle position, vehicle type, and collision risk level. This method does not require additional hardware costs, effectively solves the perception failure problem of traditional autonomous driving systems in occluded scenarios, significantly improves the safety margin of L3 and above autonomous driving systems, and simultaneously considers traffic efficiency and passenger comfort, enabling the autonomous driving system to operate safely and reliably in complex road scenarios such as mountainous and rural areas.

[0135] The following describes an apparatus embodiment of this application, which can be used to execute the vehicle autonomous driving control method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the vehicle autonomous driving control method described above.

[0136] See Figure 2 The diagram shows a block diagram of a vehicle autonomous driving control device according to an embodiment of this application.

[0137] like Figure 2 As shown, the vehicle autonomous driving control device 200 according to an embodiment of this application includes: an acquisition unit 201, a processing unit 202, a fusion unit 203, and a control unit 204.

[0138] The system includes an acquisition unit 201 for acquiring external audio data, visual sensor data, and lidar data collected during vehicle operation; a processing unit 202 for preprocessing the external audio data to obtain audio spatial features; converting the visual sensor data into visual bird's-eye view features; and converting the lidar data into lidar bird's-eye view features; a fusion unit 203 for fusing the audio spatial features, the visual bird's-eye view features, and the lidar bird's-eye view features to obtain multimodal fusion features; and a control unit 204 for identifying vehicles approaching from blind spots based on the multimodal fusion features and controlling the vehicle's movement according to the identification results.

[0139] In some embodiments of this application, based on the foregoing scheme, the processing unit 202 is configured to: perform adaptive filtering on the external audio data to remove vehicle noise and obtain a preliminary denoised frequency; perform weighted denoising on the preliminary denoised frequency to remove reverberation caused by sound reflection and obtain a direct sound audio frequency; and perform sound source localization on the direct sound audio frequency to determine the incident angle of the sound and obtain the audio spatial features.

[0140] In some embodiments of this application, based on the foregoing scheme, the fusion unit 203 is configured to: convert the audio spatial features, the visual bird's-eye view features and the laser bird's-eye view features into feature tokens of a unified format, and add spatiotemporal location encoding to each feature token; and perform weighted fusion of the feature tokens through an attention mechanism to obtain the multimodal fusion features.

[0141] In some embodiments of this application, based on the foregoing scheme, the fusion unit 203 is configured to: assign different weights to the corresponding feature tokens according to the confidence level of each modal data, so that the feature tokens corresponding to the laser bird's-eye view features pay attention to the texture information of the feature tokens corresponding to the visual bird's-eye view features, the feature tokens corresponding to the visual bird's-eye view features pay attention to color and texture information, and the feature tokens corresponding to the audio spatial features pay attention to the sound information in the blind zone direction; and fuse the multimodal fusion features based on the assigned weights and the attention results of each feature token.

[0142] In some embodiments of this application, based on the aforementioned scheme, the control unit 204 is configured to: perform three-dimensional target detection based on the multimodal fusion features; when a vehicle is detected in a blind spot not covered by visual sensing data and lidar data, enhance the identification of the vehicle in the blind spot by combining audio spatial features; track the movement trajectory of the vehicle in the blind spot, and predict the driving direction and speed of the vehicle in the blind spot.

[0143] In some embodiments of this application, based on the foregoing scheme, the control unit 204 is configured to: determine the vehicle type of the vehicle approaching from the blind spot and the collision time between the vehicle and the vehicle approaching from the blind spot; when the vehicle is in the blind spot before the curve at a distance greater than a first preset distance from the center of the curve, determine a first vehicle control strategy based on the vehicle type and the collision time; when the vehicle is in the blind spot in the curve at a distance less than or equal to the center of the curve at a distance less than or equal to the first preset distance, determine a second vehicle control strategy based on the vehicle type and the collision time; and control the vehicle's speed and lateral position according to the first vehicle control strategy or the second vehicle control strategy.

[0144] In some embodiments of this application, based on the foregoing scheme, the control unit 204 is configured to: when the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than a first time threshold, control the vehicle to smoothly decelerate to a first speed threshold and control the vehicle to laterally deviate a first distance within the lane; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold and greater than or equal to a second time threshold, control the vehicle to smoothly decelerate to a second speed threshold and control the vehicle to laterally deviate a second distance within the lane, the second speed threshold being less than the first speed threshold and the second distance being greater than the first distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, control the vehicle to smoothly decelerate to a third speed threshold and control the vehicle to laterally deviate a third distance within the lane, the third speed threshold being less than the second speed threshold and the third distance being greater than the second distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the third time threshold, control the vehicle to brake to a stop until the vehicle in the blind spot has left before starting again, the third time threshold being less than the second time threshold; when the vehicle approaching from the blind spot is a large vehicle... When a vehicle in the blind spot is a large vehicle and the collision time is greater than a first time threshold, the vehicle is controlled to smoothly decelerate to a fourth speed threshold and then laterally deviate from the lane by a fourth distance. The fourth speed threshold is less than the first speed threshold, and the fourth distance is greater than the first distance. When a large vehicle is approaching from the blind spot and the collision time is less than or equal to the first time threshold and greater than or equal to the second time threshold, the vehicle is controlled to smoothly decelerate to a fifth speed threshold and then laterally deviate from the lane by a fifth distance. The fifth speed threshold is less than the second speed threshold and less than the fourth speed threshold, and the fifth distance is greater than the second distance and greater than the fourth distance. When a large vehicle is approaching from the blind spot and the collision time is less than the second time threshold, the vehicle is controlled to smoothly decelerate to a sixth speed threshold and then laterally deviate from the lane by a third distance. The sixth speed threshold is less than the third speed threshold and less than the fifth speed threshold, and the third distance is greater than the fifth distance. When a large vehicle is approaching from the blind spot and the collision time is less than the third time threshold, the vehicle is controlled to stop and then restart after the vehicle has left the blind spot.

[0145] In some embodiments of this application, based on the foregoing scheme, the control unit 204 is configured to: when the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than a first time threshold, control the vehicle to smoothly decelerate to a first speed threshold, and control the vehicle to laterally deviate a fourth distance within the lane; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold, and greater than or equal to a second time threshold, control the vehicle to smoothly decelerate to a second speed threshold, and control the vehicle to laterally deviate a fifth distance within the lane, wherein the second speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance; when the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, control the vehicle to smoothly decelerate to a third speed threshold, and control the vehicle to laterally deviate a third distance within the lane, wherein the third speed threshold is less than the second speed threshold, and the third distance is greater than the fifth distance; when the vehicle approaching from the blind spot is a small vehicle... When a vehicle is a small vehicle and the collision time is less than a third time threshold, the vehicle is controlled to stop and continue until the vehicle in the blind spot leaves the lane. The third time threshold is less than the second time threshold. When the vehicle in the blind spot is a large vehicle and the collision time is greater than a first time threshold, the vehicle is controlled to smoothly decelerate to a seventh speed threshold and then laterally deviate a fifth distance within the lane. The seventh speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. When the vehicle in the blind spot is a large vehicle and the collision time is less than or equal to the first time threshold, the vehicle is controlled to stop and then laterally deviate a sixth distance within the lane. The sixth distance is greater than the third distance. When the vehicle in the blind spot is a large vehicle and the collision time is less than a fourth time threshold, the vehicle is controlled to stop and then reverse a preset distance until the vehicle in the blind spot leaves the lane. The fourth time threshold is less than the second time threshold and greater than the third time threshold.

[0146] In some embodiments of this application, based on the foregoing scheme, the device further includes: an updating unit, configured to update external audio data, visual sensing data, and lidar data in real time during the process of controlling the vehicle to drive; update multimodal fusion features based on the updated external audio data, visual sensing data, and lidar data; re-identify vehicles approaching from blind spots based on the updated multimodal fusion features, and dynamically adjust the vehicle control strategy according to the re-identification result.

[0147] Based on the same inventive concept, embodiments of this application provide a computer program product, the computer program product including computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor, so as to cause a computer device having the processor to perform operations performed by the vehicle autonomous driving control method as described above.

[0148] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing at least one computer program instruction, which is loaded and executed by a processor to implement the operations performed by the vehicle autonomous driving control method described above.

[0149] Based on the same inventive concept, this application also provides a vehicle, see reference. Figure 3 The diagram shows a structural schematic of a vehicle in an embodiment of this application. The vehicle includes one or more memories 304, one or more processors 302, and at least one computer program (computer program instruction) stored in the memory 304 and executable on the processor 302. When the processor 302 executes the computer program, it implements the vehicle autonomous driving control method as described above.

[0150] Among them, Figure 3 In this document, a bus architecture (represented by bus 300) is used. Bus 300 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 302 and memory represented by memory 304. Bus 300 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 can be used to store data used by processor 302 during operation.

[0151] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. When implemented in software executed by a processor, the functions can be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.

[0152] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0153] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0154] When the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer program instructions, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0155] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for controlling autonomous driving of a vehicle, characterized in that, The method includes: Acquire external audio data, visual sensor data, and LiDAR data collected during vehicle operation; The external audio data is preprocessed to obtain audio spatial features; and the visual sensing data is converted into visual bird's-eye view features, and the lidar data is converted into lidar bird's-eye view features. The audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features are fused to obtain multimodal fusion features; The system identifies vehicles approaching from blind spots based on the multimodal fusion features and controls the vehicle's movement according to the identification results.

2. The method according to claim 1, characterized in that, The preprocessing of the external audio data to obtain audio spatial features includes: The external audio data is subjected to adaptive filtering to remove vehicle noise, resulting in a preliminary denoised frequency. The initial noise-reduced frequency is subjected to weighted dreverberation processing to remove the reverberation caused by sound reflection, thereby obtaining the direct sound frequency. The direct sound audio is subjected to sound source localization processing to determine the incident angle of the sound and obtain the audio spatial characteristics.

3. The method according to claim 1, characterized in that, The process of fusing the audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features to obtain multimodal fusion features includes: The audio spatial features, the visual bird's-eye view features, and the laser bird's-eye view features are converted into feature tokens of a unified format, and a spatiotemporal location code is added to each feature token. The multimodal fusion feature is obtained by weighting and fusing the feature tokens through an attention mechanism.

4. The method according to claim 3, characterized in that, The weighted fusion of the feature tokens through an attention mechanism includes: Different weights are assigned to the corresponding feature tokens based on the confidence level of each modality data, so that the feature tokens corresponding to the laser bird's-eye view feature focus on the texture information of the feature tokens corresponding to the visual bird's-eye view feature, the feature tokens corresponding to the visual bird's-eye view feature focus on color and texture information, and the feature tokens corresponding to the audio spatial feature focus on the sound information in the blind zone direction. The multimodal fusion feature is obtained by fusing the assigned weights and the attention results of each feature token.

5. The method according to claim 1, characterized in that, The method of identifying vehicles approaching from blind spots based on the multimodal fusion features includes: Three-dimensional target detection is performed based on the aforementioned multimodal fusion features; When a vehicle is detected in a blind spot not covered by visual sensing data and lidar data, the identification of the vehicle in the blind spot is enhanced by combining audio spatial features. Track the trajectory of vehicles approaching from the blind spot and predict their direction and speed.

6. The method according to claim 1, characterized in that, The step of controlling the vehicle's movement based on the recognition result includes: Determine the type of vehicle approaching from the blind spot and the collision time between the vehicle and the vehicle approaching from the blind spot; When the vehicle is in the blind spot before the curve, which is more than a first preset distance from the center of the curve, a first vehicle control strategy is determined based on the vehicle type and the collision time. When the vehicle is in a blind spot curve at a distance less than or equal to the center of the curve, a second vehicle control strategy is determined based on the vehicle type and the collision time. The vehicle's speed and lateral position are controlled according to the first vehicle control strategy or the second vehicle control strategy.

7. The method according to claim 6, characterized in that, The step of determining the first vehicle control strategy based on the vehicle type and the collision time includes: When the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than the first time threshold, the vehicle is controlled to smoothly decelerate to the first speed threshold, and the vehicle is controlled to laterally deviate a first distance within the lane. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold and greater than or equal to the second time threshold, the vehicle is controlled to smoothly decelerate to the second speed threshold, and the vehicle is controlled to laterally deviate a second distance within the lane. The second speed threshold is less than the first speed threshold, and the second distance is greater than the first distance. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, the vehicle is controlled to smoothly decelerate to the third speed threshold, and the vehicle is controlled to laterally deviate a third distance within the lane. The third speed threshold is less than the second speed threshold, and the third distance is greater than the second distance. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the third time threshold, the vehicle is controlled to stop and continue until the vehicle in the blind spot leaves before starting again. The third time threshold is less than the second time threshold. When a large vehicle approaches from the blind spot and the collision time is greater than the first time threshold, the vehicle is controlled to smoothly decelerate to the fourth speed threshold, and the vehicle is controlled to laterally deviate by a fourth distance within the lane. The fourth speed threshold is less than the first speed threshold, and the fourth distance is greater than the first distance. When a large vehicle approaches from the blind spot and the collision time is less than or equal to the first time threshold and greater than or equal to the second time threshold, the vehicle is controlled to smoothly decelerate to the fifth speed threshold, and the vehicle is controlled to laterally deviate by the fifth distance within the lane. The fifth speed threshold is less than the second speed threshold and less than the fourth speed threshold, and the fifth distance is greater than the second distance and greater than the fourth distance. When the vehicle approaching from the blind spot is a large vehicle and the collision time is less than the second time threshold, the vehicle is controlled to smoothly decelerate to the sixth speed threshold, and the vehicle is controlled to laterally deviate by a third distance within the lane. The sixth speed threshold is less than the third speed threshold and less than the fifth speed threshold, and the third distance is greater than the fifth distance. When the vehicle approaching from the blind spot is a large vehicle and the collision time is less than the third time threshold, the vehicle is controlled to stop and continue until the vehicle in the blind spot leaves before starting again.

8. The method according to claim 6, characterized in that, The step of determining the second vehicle control strategy based on the vehicle type and the collision time includes: When the vehicle approaching from the blind spot is a small vehicle and the collision time is greater than the first time threshold, the vehicle is controlled to smoothly decelerate to the first speed threshold, and the vehicle is controlled to laterally deviate a fourth distance within the lane. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than or equal to the first time threshold and greater than or equal to the second time threshold, the vehicle is controlled to smoothly decelerate to the second speed threshold, and the vehicle is controlled to laterally deviate a fifth distance within the lane. The second speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the second time threshold, the vehicle is controlled to smoothly decelerate to the third speed threshold, and the vehicle is controlled to laterally deviate a third distance within the lane. The third speed threshold is less than the second speed threshold, and the third distance is greater than the fifth distance. When the vehicle approaching from the blind spot is a small vehicle and the collision time is less than the third time threshold, the vehicle is controlled to stop and continue until the vehicle in the blind spot leaves before starting again. The third time threshold is less than the second time threshold. When a large vehicle approaches from the blind spot and the collision time is greater than the first time threshold, the vehicle is controlled to smoothly decelerate to the seventh speed threshold, and the vehicle is controlled to laterally deviate by a fifth distance within the lane. The seventh speed threshold is less than the first speed threshold, and the fifth distance is greater than the fourth distance. When a large vehicle approaches from the blind spot and the collision time is less than or equal to the first time threshold, the vehicle is controlled to stop and then laterally deviated by a sixth distance within the lane, the sixth distance being greater than the third distance. When the vehicle approaching from the blind spot is a large vehicle and the collision time is less than the fourth time threshold, the vehicle is controlled to brake and then reverse a preset distance until the vehicle in the blind spot leaves before starting again. The fourth time threshold is less than the second time threshold and greater than the third time threshold.

9. The method according to claim 1, characterized in that, The method further includes: During the process of controlling the vehicle's movement, external audio data, visual sensor data, and lidar data are updated in real time. Based on the updated external audio data, visual sensor data, and LiDAR data, the multimodal fusion features are updated; Vehicles approaching from blind spots are re-identified based on the updated multimodal fusion features, and the vehicle control strategy is dynamically adjusted according to the re-identification results.

10. A vehicle, characterized in that, The vehicle includes one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to implement the method as described in any one of claims 1 to 9.