Method and system for cooperative interaction of multi-mode blind area outside vehicle

By combining multimodal perception and interaction, high-definition cameras and millimeter-wave radar are used to collect data, and edge computing units are used to identify the status of traffic participants. Visual and voice prompts are output through laser projection and directional sound wave arrays, which solves the problem of insufficient intention transmission between vehicles, pedestrians and non-motorized vehicles in blind spot environments and improves traffic safety.

CN122050161APending Publication Date: 2026-05-15YIXIAN INTELLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YIXIAN INTELLIGENCE
Filing Date
2026-01-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the means of communication between vehicles and pedestrians and non-motorized vehicles are limited, especially in blind spot environments where the vehicle's driving intentions cannot be effectively conveyed, leading to frequent traffic accidents. Novice drivers lack the ability to judge and respond in complex environments, and autonomous vehicles cannot proactively convey decision-making intentions to vulnerable road users.

Method used

By employing a collaborative approach of multimodal perception and interaction, data is collected through high-definition cameras and millimeter-wave radar, and fused and analyzed using edge computing units to identify the status of traffic participants. Visual and voice prompts are then output to the target area via laser projection and directional acoustic arrays, enabling accurate perception and intent transmission in vehicle blind spots and surrounding areas.

Benefits of technology

It significantly reduces the risk of accidents caused by novice drivers' lack of experience or misjudgment of intent in blind spots and complex environments, improves traffic safety in driving school training and urban commuting scenarios, and provides an effective external interaction solution for autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050161A_ABST
    Figure CN122050161A_ABST
Patent Text Reader

Abstract

The invention discloses an out-of-vehicle multi-mode blind area cooperative interaction method and an out-of-vehicle multi-mode blind area cooperative interaction system. Starting blind area sensing when a preset starting condition is met; the camera and the millimeter-wave radar collect target position, speed, attitude and track information; the edge calculation unit performs multi-source fusion and behavior recognition, and outputs a target behavior intention and / or a risk result; the strategy decision-making module is used for matching an interaction strategy, controlling laser projection equipment to project a warning or guiding graph in a target area, and controlling a directional sound wave array to output a directional voice prompt; and continuously monitoring the target state and updating the strategy, and stopping output and waiting when a termination condition is met. The system comprises a sensing layer module, a multi-mode interaction terminal, a strategy decision module and a template adaptation module, and template parameters are used for generating projection and voice output parameters and can be linked with a vehicle light and / or steering system through a vehicle-mounted bus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automotive safety technology, specifically to a method and system for collaborative interaction in multimodal blind spots outside the vehicle. Background Technology

[0002] With the continuous increase in urban traffic flow, the frequency of interactions between vehicles, pedestrians, and non-motorized vehicles at intersections has increased significantly, exacerbating the risk of traffic accidents. This is especially true in driving school settings, where novice drivers lack experience, have weak judgment and emergency response capabilities in complex environments, and are frequently exposed to crowded conditions, students crossing roads, and non-motorized vehicles weaving through traffic, further amplifying safety hazards. Furthermore, large training vehicles generally have significant blind spots, and current communication between vehicles and pedestrians / non-motorized vehicles relies solely on traditional light signals and horn cues. This method is not only limited in information transmission but also fails to effectively address the perception gaps caused by blind spots, making it difficult to accurately convey vehicle intentions (such as turning radius and traffic priority). In blind spot environments, vulnerable road users are particularly prone to not being perceived or misjudging the vehicle's driving status, leading to collisions.

[0003] Existing V2X (Vehicle-to-Everything) technologies primarily focus on communication and interaction between vehicles, aiming to achieve functions such as speed coordination and obstacle avoidance warnings. However, they pay insufficient attention to the interaction needs of vulnerable road users such as pedestrians and non-motorized vehicles, especially in terms of active perception and interaction in blind spots outside the vehicle, lacking effective solutions. Particularly in driving school scenarios, high-frequency safety hazards have not been addressed, resulting in gaps in vehicle-pedestrian interaction. In urban commuting scenarios, the behavior of pedestrians and non-motorized vehicles is highly unpredictable (such as looking down at a phone, suddenly running, or illegally crossing the road), and these behaviors are more likely to cause sudden accidents in vehicle blind spots. Novice drivers (including driving school students and newly licensed drivers) often have weak anticipation and response capabilities when faced with such unpredictable behavior. Existing traffic interaction methods also cannot accurately identify dynamic behavior and make timely intention responses, thus increasing traffic risks.

[0004] Currently, autonomous driving technology is developing rapidly. Although autonomous driving systems have accurate environmental perception capabilities, they cannot convey their decision-making intentions (such as turning or slowing down to avoid obstacles) to vulnerable road users outside the vehicle in a traditional way. This may lead to conflicts among road users due to insufficient prediction of the behavior of autonomous vehicles, and there is still a gap in the interaction between vehicles and people.

[0005] Therefore, there is an urgent need for an external multimodal blind spot collaborative interaction method that can accurately perceive the status of vulnerable road users in both blind spots and non-blind spots, and proactively convey the vehicle's driving intentions. This method can not only be prioritized for application in driving school scenarios to address high-frequency safety hazards and assist novice learners and newly licensed drivers in dealing with complex traffic environments, but it can also be further adapted to autonomous vehicles, filling existing technological gaps and comprehensively reducing traffic accident rates in different scenarios. Summary of the Invention

[0006] To address the shortcomings of existing technologies, such as limited means of communication between vehicles and pedestrians / non-motorized vehicles, gaps in vehicle-to-human interaction, high accident rates at complex intersections and blind spots, and the limited judgment and response capabilities of novice drivers (including driving school students and newly licensed drivers) in complex traffic environments, this invention provides a method and system for collaborative interaction in multimodal blind spots outside the vehicle. Applied to driving school scenarios, it addresses high-frequency safety hazards caused by dense crowds and unfamiliarity with driving techniques among novice students, while also providing safety assistance to newly licensed drivers on independent roads. In the future, the method can be further adapted to autonomous vehicles, filling the gaps in vehicle-to-human interaction. Through the coordinated use of multimodal perception and interaction, this invention can accurately perceive the status of vulnerable road users in and around the vehicle's blind spots, proactively conveying the vehicle's driving intentions (such as turning radius, traffic area, and warning prompts). This solution helps vulnerable road users clearly understand vehicle dynamics and avoid risks in a timely manner, while also relieving novice drivers of the pressure of environmental judgment and assisting them in dealing with complex intersections and training scenarios. This achieves efficient matching of vehicle and driver intentions, significantly reducing the risk of accidents caused by insufficient driving experience, misjudgment of intentions, or lack of perception in blind spots and complex environments. It comprehensively improves traffic safety in various application scenarios such as driving school training and urban commuting. Simultaneously, it provides an effective solution for external interaction of autonomous vehicles.

[0007] To achieve the above objectives, the first aspect of the present invention provides a method for cooperative interaction in multimodal blind spots outside a vehicle, comprising the following steps: S1. Perception Activation: When preset activation conditions are met, the perception of the vehicle's blind spot and surrounding area is activated. S2. Multi-source data acquisition: High-definition cameras installed at key locations in the vehicle's blind spot are used to collect data on the appearance characteristics, location distribution, and behavioral posture of traffic participants in the blind spot and surrounding area. Millimeter-wave radar is used to collect data on the speed, distance, and movement trajectory of the traffic participants. S3. Data Fusion and Intent Prediction: The edge computing unit is used to fuse and analyze the data collected by the high-definition camera and millimeter-wave radar. Based on the preset behavior recognition algorithm, the movement state of traffic participants is identified. Combined with environmental information such as intersection traffic signals and lane distribution, the behavioral intent and / or risk judgment result of traffic participants is output. S4. Multimodal interaction strategy execution: Match a multimodal interaction strategy according to the behavioral intent and / or risk assessment result, and control the multimodal interaction terminal set outside the vehicle to output visual projection information and directional voice prompt information to the target area; wherein, the multimodal interaction terminal includes at least one laser projection device and at least one directional sound wave array; S5. Interaction Update and Termination: During vehicle operation, the status of traffic participants in the target area is continuously monitored, and the multimodal interaction strategy is updated based on the monitoring results. When the preset termination condition is met, the output of the visual projection information and directional voice prompt information is stopped and the system enters standby mode.

[0008] Furthermore, S4 includes: when it is predicted that a pedestrian intends to cross the road or enter the blind spot of a vehicle, controlling the laser projection device to project a stop line and a warning frame onto the road surface in front of the pedestrian, and controlling the directional sound wave array to output directional voice prompts to the pedestrian, prompting them to stop or avoid the vehicle.

[0009] Furthermore, S4 includes: when it is predicted that a non-motorized vehicle is within the blind spot of a vehicle in a turning scenario at an intersection, controlling multiple laser projection devices to collaboratively project turning radius markings and traffic guidance areas onto the road surface and / or surrounding facades in the area where the non-motorized vehicle is located; wherein, the turning radius markings are presented in the form of dynamic arcs and are updated according to the turning trajectory corresponding to the current turning angle of the vehicle, and are prompted by the vehicle's turn signals in conjunction with the vehicle bus.

[0010] Furthermore, S4 also includes: when it is predicted that the target object is in normal passage and there is no risk of intention conflict, controlling the laser projection device and the directional acoustic array to be in a low-power standby state; when the intention conflict risk is detected, initiating the execution of the multimodal interaction strategy.

[0011] Furthermore, it also includes a custom template configuration step: providing a visual configuration interface to allow users to configure interactive template parameters based on different city traffic rules, regional cultural habits, and vehicle blind spot ranges; the interactive template parameters include the graphic style of the laser projection, the projection position, the projection size, and the prompting script for directional voice prompts; and in step S4, the corresponding visual projection information and directional voice prompts are generated and output according to the interactive template parameters.

[0012] A second aspect of this invention provides an external multimodal blind spot cooperative interaction system, comprising a multimodal interaction terminal, a perception layer module, an interaction strategy decision module, and a custom template adaptation module, wherein each module interacts with data via an in-vehicle bus; wherein: The multimodal interactive terminal is located outside the vehicle and includes at least one laser projection device and at least one directional sound wave array, used to output visual projection information and directional voice prompt information to the vehicle's blind spot and surrounding area; The perception layer module includes a high-definition camera, millimeter-wave radar, and an edge computing unit, which are used to collect the position, speed, and attitude data of traffic participants in the vehicle's blind spot and surrounding area, and output the traffic participants' behavioral intentions and / or risk assessment results. The interaction strategy decision module is communicatively connected to the perception layer module and is used to match a multimodal interaction strategy according to the behavioral intention and / or risk judgment result, and output control commands to the multimodal interaction terminal to control the multimodal interaction terminal to execute information output; The custom template adaptation module is used to provide a visual configuration interface and store interaction template parameters, so that the interaction strategy decision module can generate output parameters corresponding to the visual projection information and the directional voice prompt information according to the interaction template parameters.

[0013] Furthermore, the edge computing unit is equipped with a deep learning-based behavior recognition model, which is used to fuse and analyze the data collected by the high-definition camera and the millimeter-wave radar to output the behavioral intentions and / or risk assessment results of the traffic participants.

[0014] Furthermore, the multimodal interactive terminal includes multiple laser projection devices, and the multiple laser projection devices adopt a multi-angle collaborative projection structure to cover the corresponding areas of the blind spots on the front, sides and / or rear of the vehicle; the directional sound wave array adopts a directional sound wave propagation structure to propagate the voice prompt information directionally to the target area.

[0015] Furthermore, the interactive strategy decision module is configured as follows: when it is predicted that a pedestrian intends to cross the road or enter the vehicle's blind spot, it controls the laser projection device to project a stop line and a warning box, and controls the directional sound wave array to output directional voice prompts to prompt the pedestrian to stop or give way to the vehicle; when it is predicted that a non-motorized vehicle is within the vehicle's blind spot when turning at an intersection, it controls multiple laser projection devices to collaboratively project turning radius markers and traffic guidance areas, and links the vehicle's turn signals through the vehicle bus to provide a prompt.

[0016] Furthermore, the vehicle bus is a CAN bus, and the system interacts with the vehicle's steering system and / or lighting system via the CAN bus to match the multimodal interaction strategy with the vehicle's driving state. The present invention, by adopting the above technical solution, has at least the following beneficial effects: This invention addresses scenarios where pedestrians and non-motorized vehicles frequently intersect, such as driving school grounds, intersections, and blind spots. By perceiving and predicting the intentions of targets in blind spots and surrounding areas, and proactively initiating external interactive prompts when a risky intention is detected, it reduces the risk of collisions caused by factors such as obstructed vision and lack of experience during training or travel, providing safety assistance for driving training and daily travel.

[0017] This invention combines visual projection with directional voice prompts. Laser projection is used to present visual information such as stop lines, warning signs, turning radius signs, and traffic guidance areas in the target area, while directional sound waves are used to output voice prompts to the target area. The two work together to improve the understanding of vehicle driving intentions by vulnerable road users and reduce the impact of traditional horn honking or loud external announcements on the surrounding environment.

[0018] This invention uses multi-source data acquisition from high-definition cameras and millimeter-wave radar, and performs real-time fusion analysis by an edge computing unit. It can identify information such as the location, speed, movement trajectory and behavioral posture of traffic participants, and output their behavioral intentions such as crossing, turning, and entering blind spots, as well as / or risk assessment results, thereby improving the pertinence and timeliness of interactive strategy matching.

[0019] This invention continuously monitors the status of traffic participants within the target area during vehicle operation and updates the interaction strategy in real time based on changes in the target status. It can adjust the content or output method of projection and voice prompts in a timely manner when the target behavior changes, thereby improving the continuity and effectiveness of interactive prompts.

[0020] This invention provides a visual configuration interface that allows users to configure template parameters such as the style, position, size, and voice prompts of projected graphics based on urban traffic rules, regional cultural habits, and vehicle blind spot range. This enables the system to flexibly adapt to the application needs of different regions and vehicle models, reducing interaction incompatibility issues caused by regional and vehicle model differences.

[0021] This invention can be integrated and upgraded with existing vehicle perception hardware and in-vehicle systems to achieve external collaborative interaction without changing the basic vehicle architecture. At the same time, this solution can be used in scenarios such as driver training and urban commuting, and can be extended to vehicles with autonomous driving functions, providing an achievable technical means for the externalization of vehicle driving intentions to vulnerable traffic participants. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this embodiment. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of the method for collaborative interaction of multimodal blind spots outside the vehicle according to the present invention. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this embodiment. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this embodiment as detailed in the appended claims. Example 1

[0025] like Figure 1 As shown, this embodiment provides a method for cooperative interaction in multimodal blind spots outside a vehicle, including the following steps: S1. Perception Activation: When preset activation conditions are met, the perception of the vehicle's blind spot and surrounding area is activated. The preset activation conditions specifically include the vehicle identifying the intersection, school zone, driving school training area, or other complex traffic areas through onboard GPS positioning and electronic map matching, or the detection of pedestrians, non-motorized vehicles, or other traffic participants within 50 meters of the vehicle through the perception device. By clearly defining the triggering scenario and detection range, the perception mode can be accurately triggered, avoiding energy consumption caused by invalid activation, while covering high-frequency risk scenarios and improving the pertinence of safety protection.

[0026] S2. Multi-source data acquisition: High-definition cameras positioned at key locations in the vehicle's blind spots collect data on the appearance, location distribution, and behavioral posture of traffic participants in and around the blind spots. Millimeter-wave radar collects data on the speed, distance, and trajectory of these traffic participants. The high-definition cameras are primarily deployed at key blind spot nodes such as the A-pillar, below the left and right rearview mirrors, on both sides of the front bumper, and the tailgate. They possess a wide dynamic range, clearly capturing the posture of traffic participants in strong light and backlight conditions. The millimeter-wave radar is deployed in a staggered manner with the cameras, focusing on covering blind spots such as turning and close-range blind spots where cameras are easily obstructed. It can penetrate adverse weather conditions such as rain, fog, and dust, as well as minor obstacles. The complementary data collection by both systems compensates for the limitations of a single device and provides accurate data sources for subsequent intention prediction through multi-dimensional data acquisition, improving the comprehensiveness and reliability of data collection.

[0027] S3. Data Fusion and Intent Prediction: The edge computing unit fuses and analyzes the data collected by the high-definition camera and millimeter-wave radar. Based on a preset behavior recognition algorithm, it identifies the movement state of traffic participants and, combined with environmental information such as intersection traffic signals and lane distribution, outputs the behavioral intent and / or risk assessment results of traffic participants. The edge computing unit is integrated into the vehicle's central control unit, adopting a localized data processing mode to avoid the delay caused by data transmission to the cloud, and the prediction delay can be controlled within an extremely low range. The preset behavior recognition algorithm is optimized based on the YOLOv8 deep learning model and can accurately identify high-risk behavioral postures such as looking down at a mobile phone, running, standing still and looking around, and suddenly changing direction. Combined with environmental information for comprehensive prediction, it can achieve rapid and accurate determination of the behavioral intent of traffic participants, providing sufficient reaction time for subsequent interaction strategy execution and reducing the safety risks caused by misjudgment.

[0028] S4. Multimodal interaction strategy execution: Match a multimodal interaction strategy according to the behavioral intent and / or risk assessment result, and control the multimodal interaction terminal set outside the vehicle to output visual projection information and directional voice prompt information to the target area; wherein, the multimodal interaction terminal includes at least one laser projection device and at least one directional sound wave array; As one implementation method, S4 in this embodiment includes: when it is predicted that a pedestrian intends to cross the road or enter the blind spot of a vehicle, controlling the laser projection device to project a stop line and a warning frame onto the road surface in front of the pedestrian, and controlling the directional sound wave array to output directional voice prompts to the pedestrian, prompting them to stop or give way to vehicles; the graphics projected by the laser projection device have high definition and can still be clearly identified even in strong midday light; the directional sound wave array adopts a ±15° directional propagation angle design, only transmitting voice to the target pedestrian, without causing noise interference to other traffic participants in the vicinity; the visual and auditory synergy of prompts can ensure that pedestrians quickly receive warning information, especially suitable for pedestrians with poor concentration, improving the effectiveness of the prompts.

[0029] As one implementation method, S4 in this embodiment includes: when it is predicted that a non-motorized vehicle is within the blind spot of a vehicle turning at an intersection, controlling multiple laser projection devices to collaboratively project turning radius markings and traffic guidance areas onto the road surface and / or surrounding facades in the area where the non-motorized vehicle is located; wherein, the turning radius markings are presented in the form of dynamic arcs and are updated in real time according to the turning trajectory corresponding to the current turning angle of the vehicle, and are linked to the vehicle's turn signals via the vehicle bus for prompting; the multiple laser projection devices adopt a multi-angle collaborative projection structure, which can achieve no blind spots in the coverage of the turning radius markings and traffic guidance areas, and the dynamic arcs are synchronized with the vehicle's turning actions to accurately transmit the actual driving trajectory of the vehicle, which can not only make the safe passage range clear for non-motorized vehicles, but also enhance the prompting effect through the linkage of turn signals, effectively avoiding the "inner wheel difference" risk when large vehicles turn.

[0030] As one implementation method, S4 in this embodiment further includes: when it is predicted that the target object is passing normally and there is no risk of intention conflict, controlling the laser projection device and the directional acoustic array to be in a low-power standby state; when the intention conflict risk is detected, starting the execution of the multimodal interaction strategy; the low-power standby state can significantly reduce device energy consumption and extend device life, while the device maintains a fast wake-up capability to ensure that interaction can be started immediately when the risk occurs, achieving a balance between energy saving and safety.

[0031] As one implementation method, this embodiment also includes a custom template configuration step: providing a visual configuration interface to allow users to configure interactive template parameters based on different city traffic rules, regional cultural habits, and vehicle blind spot ranges; the interactive template parameters include the graphic style, projection position, projection size, and prompting scripts for directional voice prompts; and in step S4, generating and outputting corresponding visual projection information and directional voice prompts based on the interactive template parameters; the visual configuration interface supports touch operation, presets multiple templates adapted to different city traffic rules, which users can switch and fine-tune with one click, and also supports multilingual voice prompt customization, adapting to diverse scenarios such as driving schools and urban commuting, improving the system's versatility and flexibility, and facilitating large-scale promotion and application.

[0032] S5. Interaction Update and Termination: During vehicle operation, the status of traffic participants within the target area is continuously monitored, and the multimodal interaction strategy is updated based on the monitoring results. When a preset termination condition is met, the output of the visual projection information and directional voice prompt information is stopped, and the system enters a standby state. The preset termination conditions include scenarios such as traffic participants completely leaving the vehicle's blind spot and risk area, the vehicle completing a turn or passing through an intersection, and the vehicle leaving a complex traffic area. Real-time strategy updates and precise termination of interaction can avoid invalid prompts from continuously interfering with traffic order, while quickly switching to a standby state to further optimize energy consumption.

[0033] Example 2 This embodiment provides a multimodal blind spot cooperative interaction system for vehicles, including a multimodal interaction terminal, a perception layer module, an interaction strategy decision module, and a custom template adaptation module, and the modules interact with each other via an in-vehicle bus; wherein: The multimodal interactive terminal is located outside the vehicle and includes at least one laser projection device and at least one directional acoustic array. It is used to output visual projection information and directional voice prompts to the vehicle's blind spots and surrounding areas. The multimodal interactive terminal is preferably configured with more than four sets of laser projection devices, which are deployed at the four corners of the vehicle's top and the sides of the vehicle body, covering the entire blind spot of the front, sides, and rear of the vehicle. The directional acoustic array is deployed one-to-one with the laser projection devices. Each array consists of at least eight ultrasonic transducers with an effective propagation distance of 1-8 meters. It can achieve accurate transmission of visual and auditory information in the same area, ensuring that traffic participants in all directions of the blind spot are covered and improving the comprehensiveness of the interaction.

[0034] The perception layer module includes a high-definition camera, a millimeter-wave radar, and an edge computing unit. It is used to collect the position, speed, and posture data of traffic participants in the vehicle's blind spot and surrounding area, and output the behavioral intentions and / or risk assessment results of traffic participants. The high-definition camera uses a high-definition imaging module to capture subtle changes in the posture of traffic participants. The millimeter-wave radar has multi-target detection capabilities and can simultaneously track the movement trajectories of multiple traffic participants. The edge computing unit has a built-in data fusion algorithm that can perform real-time correlation and calibration of data collected by the two devices. The three work together to achieve accurate perception of the state of traffic participants in complex environments, providing solid data support for intention prediction.

[0035] The interaction strategy decision module is communicatively connected to the perception layer module. It is used to match multimodal interaction strategies based on the behavioral intent and / or risk assessment results, and output control commands to the multimodal interaction terminal to control the multimodal interaction terminal to execute information output. The interaction strategy decision module has a built-in preset strategy library that stores interaction schemes for different traffic participants and different risk scenarios. It adopts a fast matching algorithm and can complete strategy matching and command issuance within milliseconds, ensuring the instantaneous response of interactive actions and avoiding safety hazards caused by command delays.

[0036] The custom template adaptation module provides a visual configuration interface and stores interaction template parameters, enabling the interaction strategy decision module to generate output parameters corresponding to the visual projection information and directional voice prompts based on the interaction template parameters. This module supports the categorized storage and automatic matching of templates, and can automatically call the adapted interaction templates based on vehicle positioning information and vehicle model information, eliminating the need for repetitive manual operations. It also supports the import and export of templates, facilitating batch configuration management for driving schools, transportation companies, etc., and improving the system's usability and operational efficiency.

[0037] As one implementation method, the edge computing unit in this embodiment is equipped with a deep learning-based behavior recognition model, which is used to fuse and analyze the data collected by the high-definition camera and the millimeter-wave radar to output the behavioral intentions and / or risk assessment results of the traffic participants. The behavior recognition model has been trained with a large amount of traffic scene data, can accurately distinguish between normal passage and high-risk behavior, and has self-optimization capabilities. It can continuously optimize the recognition accuracy according to the data of actual application scenarios, adapt to the behavioral habits of different regions and different groups of people, and improve the universality and accuracy of intention prediction.

[0038] As one implementation method, the multimodal interactive terminal in this embodiment includes multiple laser projection devices, and the multiple laser projection devices adopt a multi-angle collaborative projection structure to cover the corresponding areas of the blind spots on the front, sides and / or rear of the vehicle; the directional sound wave array adopts a sound wave directional propagation structure to propagate the voice prompt information to the target area in a directional manner; the multi-angle collaborative projection structure can eliminate projection blind spots and ensure the complete presentation of graphic information under complex terrain, and the sound wave directional propagation structure can effectively isolate the surrounding environmental noise interference, allowing the target traffic participants to clearly receive the voice prompts. The combination of the two achieves the interactive effect of "accurate prompts without disturbing others".

[0039] As one implementation method, the interactive strategy decision module in this embodiment is configured as follows: when it is predicted that a pedestrian intends to cross the road or enter the vehicle's blind spot, the module controls the laser projection device to project a stop line and a warning box, and controls the directional sound wave array to output directional voice prompts to prompt the pedestrian to stop or give way to the vehicle; when it is predicted that a non-motorized vehicle is within the vehicle's blind spot when turning at an intersection, the module controls multiple laser projection devices to collaboratively project turning radius markings and traffic guidance areas, and links the vehicle's turn signals through the vehicle bus to provide a prompt; the interactive strategy decision module can dynamically adjust the prompt intensity according to the distance and speed of the traffic participant, such as increasing the voice volume and enlarging the projected graphics when the distance to the vehicle is close, to achieve differentiated prompts and further improve the pertinence and effectiveness of the prompts.

[0040] In one implementation method, the vehicle bus in this embodiment is a CAN bus, and the system is linked with the vehicle's steering system and / or lighting system through the CAN bus to match the multimodal interaction strategy with the vehicle's driving state. The CAN bus has the advantages of fast data transmission speed and strong anti-interference capability, and can synchronize data such as vehicle steering angle and lighting status in real time to ensure that the interaction strategy is accurately synchronized with the actual driving state of the vehicle, avoiding misjudgments caused by inconsistencies between interaction information and vehicle actions. At the same time, it is compatible with the existing vehicle bus system, which facilitates system retrofit and lowers the application threshold.

[0041] Example 3 This embodiment further details the specific structure, hardware configuration, workflow, and application effects of the multimodal blind spot collaborative interaction system outside the vehicle according to the present invention, so as to fully demonstrate the feasibility and practicality of the technical solution. The core of the system in this embodiment consists of a multimodal interaction terminal, a perception layer module, an interaction strategy decision module, and a custom template adaptation module. These modules cooperate through the vehicle bus to achieve precise interaction between the vehicle and vulnerable road users outside the vehicle. The specific structure and implementation method are as follows: I. Multimodal Interactive Terminal The multimodal interactive terminal is deployed on the top and sides of the vehicle, focusing on covering the blind spots in front, to the sides, and behind the vehicle. It consists of laser projection equipment and a directional sound wave array. Through the coordinated operation of multiple laser projection devices, it achieves full coverage and accurate transmission of visual and auditory information in the blind spots and surrounding areas. The laser projection equipment is equipped with multiple high-precision components and adopts a multi-angle collaborative projection structure, possessing wide-range and high-definition projection performance. It can accurately project graphic information related to the vehicle's driving intentions (including turning radii, traffic area markings, warning boxes, stop lines, etc.) onto the road surface around the vehicle. Even in complex lighting conditions such as midday sun, it maintains good visual recognition, ensuring that the target object clearly captures the guidance information. The directional sound wave array uses directional sound wave propagation technology to accurately focus voice prompts on the target area (such as the location of specific pedestrians or non-motorized vehicles), effectively avoiding noise pollution caused by voice diffusion and achieving a precise interactive effect of "only prompting the target, without interfering with the surroundings."

[0042] II. Perception Layer Module The perception layer module integrates high-definition cameras, millimeter-wave radar, and edge computing units. Its core function is to accurately perceive and predict the behavioral intentions of vulnerable road users within the vehicle's blind spot. At the same time, it extends to cover the surrounding areas outside the blind spot, ensuring no blind spots and providing reliable data support for the execution of subsequent interaction strategies.

[0043] (a) Data collection High-definition cameras are deployed at key locations around the vehicle, including the front, rear, left, right, and blind spots, to collect road image information. They comprehensively capture the appearance characteristics, location distribution, and behavioral postures (specifically, high-risk postures such as looking down at a mobile phone, running, and standing still to look around) of pedestrians and non-motorized vehicles in blind spots and surrounding areas. Millimeter-wave radar is deployed in a staggered manner with the high-definition cameras to simultaneously collect quantitative data such as the speed, distance, and movement trajectory of the target objects. This effectively compensates for the perception shortcomings of high-definition cameras in adverse weather conditions such as rain, snow, and fog, as well as in scenarios where blind spots are obstructed, forming a multi-source data complementary acquisition system of "image + radar".

[0044] (ii) Edge computing processing The edge computing unit performs real-time fusion analysis of dual-source data collected by high-definition cameras and millimeter-wave radar. It identifies the movement status of pedestrians and non-motorized vehicles through preset behavior recognition algorithms. At the same time, it combines environmental information such as traffic signals at intersections and lane distribution to comprehensively predict whether the target object has the intention to cross the road, turn, or enter blind spots, providing accurate basis for interaction strategy matching.

[0045] III. Interactive Strategy Decision Module Based on the target object's state and intent prediction results output by the perception layer module, the interaction strategy decision module automatically matches the corresponding multimodal interaction strategy and sends control commands to the multimodal interaction terminal to drive it to execute specific information transmission actions, specifically covering different risk scenarios: 1. For pedestrians predicted to be crossing the road: control the laser projection equipment to project a stop line and a red warning box onto the road in front of the pedestrian, and at the same time control the directional sound wave array to deliver a directional voice prompt to the pedestrian, "Please wait, and cross after the vehicle has passed," forming a dual warning effect of visual and auditory perception to ensure that the pedestrian receives the hazard avoidance information in a timely manner.

[0046] 2. For non-motorized vehicles within the blind spot of vehicles when turning at intersections: Control multiple sets of laser projection devices to work together to project turning radius markings and green passage zones onto the road surface or surrounding facades where the non-motorized vehicles are located. The turning radius markings are presented in the form of dynamic arcs, accurately matching the actual turning trajectory corresponding to the vehicle's current turning angle, intuitively informing the non-motorized vehicle driver of the inner wheel difference range and the area swept by the vehicle body when turning. The green passage zone is marked with clear green high-brightness blocks, clearly defining the area where non-motorized vehicles can safely pass, avoiding the vehicle's turning blind spot and vehicle trajectory. At the same time, it is linked to the vehicle's turn signals to flash as a reminder. Through visual trajectory and area guidance, non-motorized vehicle drivers can clearly know the avoidance direction and safety boundary, accurately avoiding the risk of collisions in the blind spot.

[0047] 3. For pedestrians and non-motorized vehicles passing through normally: Control the multimodal interactive terminal to maintain a low-power standby state, only continuously receiving monitoring data from the perception layer module, and immediately start interaction when an intention conflict risk is detected, so as to achieve a dynamic balance between energy saving and safety.

[0048] IV. Custom Template Adaptation Module This system allows users to configure interactive templates through a custom template adaptation module. Adaptation scope includes, but is not limited to, projected graphic styles, projection positions, projection sizes, and voice prompts. Users can adjust the interactive templates to suit different city traffic rules (such as traffic priority at specific intersections and regulations for non-motorized vehicles) and regional cultural habits (such as multilingual voice prompts and graphic signs that conform to local understanding), enhancing the system's regional adaptability. Furthermore, users can customize the laser projection coverage area and projection angle based on the blind spot differences of different vehicle models, further expanding the system's adaptability scenarios.

[0049] V. Specific Hardware Configuration and Workflow The system in this embodiment is primarily used in driving school training vehicles, and is also compatible with the private vehicles of newly licensed learners. In the future, it can be extended to autonomous vehicles. The hardware parameters and specific workflows of each module of the system are as follows: (a) Hardware configuration 1. Multimodal interactive terminal: It is equipped with 4 sets of 30W solid-state laser projectors as the core interactive components. It adopts multi-angle distributed deployment, with a single unit projection distance range of 1-10m and a projection accuracy of 1mm / px. Multiple sets can work together to achieve full coverage of blind spots and surrounding areas, accurately projecting onto the road surface and facades around the blind spots; the directional sound wave array consists of 8 ultrasonic transducers with a directional propagation angle of ±15° and an effective propagation distance of 1-8m, which can accurately transmit voice prompts to target objects in blind spots.

[0050] 2. Perception Layer Module: The high-definition camera uses an 8-megapixel CMOS camera with a frame rate of 30fps and a wide dynamic range of 120dB. It is mainly deployed in key blind spot locations such as the A-pillar, rearview mirror, and rear of the vehicle to achieve full coverage of the blind spot. The millimeter-wave radar uses a 77GHz frequency band radar sensor with a detection range of 0.1-150m and a detection angle of ±60°. It can penetrate some obstructions to achieve target detection in the blind spot. The edge computing unit uses an NVIDIA Jetson Xavier NX processor, which has efficient real-time data fusion and algorithm processing capabilities. It is equipped with a deep learning-based behavior recognition model (optimized by YOLOv8 algorithm) to realize real-time recognition of the position, speed, and posture of pedestrians and non-motorized vehicles in the blind spot and surrounding areas, with an intent prediction latency of ≤100ms.

[0051] 3. Interaction Strategy Decision Module: It has multiple built-in preset interaction strategy templates, which have the ability to quickly match strategies and issue commands; the custom template adaptation module provides a visual configuration interface, and users can modify the interaction templates through the vehicle central control screen or the background management system.

[0052] (II) Work Process 1. Perception Activation: When the vehicle is located by the vehicle GPS and the map is matched and identified, and the vehicle is driving to the intersection area, or when the perception layer module detects that there are pedestrians or non-motorized vehicles within 50m around the vehicle, the system automatically activates the full-dimensional perception mode and each module enters the working ready state.

[0053] 2. Data Acquisition and Analysis: High-definition cameras and millimeter-wave radar simultaneously acquire target object data. The edge computing unit fuses the dual-source data to identify the target object type (pedestrian / non-motorized vehicle), location coordinates, movement speed, and behavioral posture. Combined with the traffic light status at the intersection, it completes the prediction of behavioral intentions (crossing the road, turning, normal passage).

[0054] 3. Strategy Matching and Execution: Based on the intent prediction results, the system matches corresponding interaction strategies, enhances interactive prompts for target objects within the blind spot, and controls the multimodal interactive terminal to execute information transmission actions. For example, when it is predicted that a pedestrian is about to enter or is already in the vehicle's blind spot, the interaction strategy decision module immediately issues an instruction to the corresponding area's laser projection device, causing it to project a red warning box and a white stop line onto the road surface 2 meters in front of the pedestrian. At the same time, it controls the directional sound wave array to transmit a voice prompt, "Please wait, proceed only after the vehicle has passed." When it is predicted that a non-motorized vehicle will enter the vehicle's turning blind spot at an intersection and follow the turn, it controls multiple sets of laser projection devices to collaboratively project a green turning radius trajectory and a passing area graphic, while simultaneously coordinating with the vehicle's turn signals to flash at a frequency of 5Hz to clearly inform the non-motorized vehicle to avoid the blind spot.

[0055] 4. Interaction Termination: When the target object leaves the risk area (such as a pedestrian stopping and waiting, or a non-motorized vehicle entering a safe passage area), or after the vehicle completes a turn and passes through an intersection, the interaction strategy decision module controls the multimodal interaction terminal to stop information transmission, restore low-power standby state, and wait for the next perception activation.

[0056] (III) Template Customization Operation Users can flexibly adjust the parameters of the interactive template through the visual interface of the custom template adaptation module: for specific city non-motorized vehicle turning regulations, the graphic ratio, projection position and size of the projected turning radius can be adjusted; for intersections around schools, a "Caution: Children" projected graphic and corresponding voice prompt can be added to meet personalized scenario needs.

[0057] VI. Application Scenarios and Effects This system is primarily integrated into driving school training vehicles to address safety hazards caused by dense crowds and unfamiliarity with driving techniques among novice learners. It can also be directly integrated into intelligent connected vehicles used by newly licensed learners, providing safety assistance for their independent road use. In the future, it can be smoothly extended to autonomous vehicles, filling the gap in vehicle-to-human interaction in autonomous driving scenarios. This system is particularly suitable for large training vehicles, trucks, buses, and other vehicles with large blind spots. Through a CAN bus, it links with the vehicle's powertrain, steering, and lighting systems to ensure precise matching of interaction strategies with the vehicle's actual driving status (steering angle, speed).

[0058] Actual testing has verified that, in driving school scenarios, the system achieves an accuracy rate of ≥98% in recognizing pedestrians and non-motorized vehicles within the training area and an interactive information reception rate of ≥95%, which can reduce the risk of collision accidents by more than 90% during training, effectively ensuring teaching safety and efficiency; in scenarios where new students are on the road, it can reduce the risk of accidents at intersections and in blind spots by more than 70%; and in autonomous driving scenarios, it can reduce the risk of vehicle-human interaction conflicts by more than 80%.

[0059] The system in this embodiment not only significantly improves traffic safety in different scenarios and provides effective protection for novice learners to smoothly transition to proficient driving, but also lays the foundation for future application in the field of autonomous driving, possessing significant safety benefits, social value, and long-term prospects. Furthermore, the system can be upgraded by integrating existing vehicle perception hardware (such as cameras and millimeter-wave radar already equipped in some models), effectively controlling additional cost increases: for driving schools, it can significantly reduce losses such as personal injury compensation, vehicle repairs, and teaching interruptions caused by accidents; for families of new learners, upfront safety investment can avoid the risk of high-frequency accidents in the early stages of driving, possessing significant long-term economic and social value, and can be extended to autonomous vehicles without repeated research and development modifications, demonstrating strong technology reusability.

[0060] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for collaborative interaction in multimodal blind spots outside a vehicle, characterized in that: Includes the following steps: S1. Perception Activation: When preset activation conditions are met, the perception of the vehicle's blind spot and surrounding area is activated. S2. Multi-source data acquisition: High-definition cameras installed at key locations in the vehicle's blind spot are used to collect data on the appearance characteristics, location distribution, and behavioral posture of traffic participants in the blind spot and surrounding area. Millimeter-wave radar is used to collect data on the speed, distance, and movement trajectory of the traffic participants. S3. Data Fusion and Intent Prediction: The edge computing unit is used to fuse and analyze the data collected by the high-definition camera and millimeter-wave radar. Based on the preset behavior recognition algorithm, the movement state of traffic participants is identified. Combined with environmental information such as intersection traffic signals and lane distribution, the behavioral intent and / or risk judgment result of traffic participants is output. S4. Multimodal interaction strategy execution: Match a multimodal interaction strategy according to the behavioral intent and / or risk assessment result, and control the multimodal interaction terminal set outside the vehicle to output visual projection information and directional voice prompt information to the target area; wherein, the multimodal interaction terminal includes at least one laser projection device and at least one directional sound wave array; S5. Interaction Update and Termination: During vehicle operation, the status of traffic participants in the target area is continuously monitored, and the multimodal interaction strategy is updated based on the monitoring results. When the preset termination condition is met, the output of the visual projection information and directional voice prompt information is stopped and the system enters standby mode.

2. The method according to claim 1, characterized in that: S4 includes: when it is predicted that a pedestrian intends to cross the road or enter the blind spot of a vehicle, controlling the laser projection device to project a stop line and a warning box on the road surface in front of the pedestrian, and controlling the directional sound wave array to output directional voice prompts to the pedestrian to stop or avoid the vehicle.

3. The method according to claim 1, characterized in that: S4 includes: when it is predicted that a non-motorized vehicle is within the blind spot of a vehicle in a turning scenario at an intersection, controlling multiple laser projection devices to collaboratively project turning radius markings and traffic guidance areas onto the road surface and / or surrounding facades in the area where the non-motorized vehicle is located; wherein, the turning radius markings are presented in the form of dynamic arcs and are updated according to the turning trajectory corresponding to the current turning angle of the vehicle, and are prompted by the vehicle's turn signals in conjunction with the vehicle bus.

4. The method according to claim 1, characterized in that: The S4 further includes: when it is predicted that the target object is in normal passage and there is no risk of intention conflict, controlling the laser projection device and the directional acoustic array to be in a low-power standby state; when the risk of intention conflict is detected, the multimodal interaction strategy is activated.

5. The method according to claim 1, characterized in that: It also includes a custom template configuration step: providing a visual configuration interface so that users can configure interactive template parameters based on different city traffic rules, regional cultural habits and vehicle blind spot ranges; the interactive template parameters include the graphic style of the laser projection, the projection position, the projection size and the prompting words of the directional voice prompt information; and in step S4, the corresponding visual projection information and directional voice prompt information are generated and output according to the interactive template parameters.

6. A multimodal blind spot cooperative interaction system for vehicles, characterized in that: It includes a multimodal interactive terminal, a perception layer module, an interaction strategy decision module, and a custom template adaptation module, and each module interacts with other modules via an in-vehicle bus; among which: The multimodal interactive terminal is located outside the vehicle and includes at least one laser projection device and at least one directional sound wave array, used to output visual projection information and directional voice prompt information to the vehicle's blind spot and surrounding area; The perception layer module includes a high-definition camera, millimeter-wave radar, and an edge computing unit, which are used to collect the position, speed, and attitude data of traffic participants in the vehicle's blind spot and surrounding area, and output the traffic participants' behavioral intentions and / or risk assessment results. The interaction strategy decision module is communicatively connected to the perception layer module and is used to match a multimodal interaction strategy according to the behavioral intention and / or risk judgment result, and output control commands to the multimodal interaction terminal to control the multimodal interaction terminal to execute information output; The custom template adaptation module is used to provide a visual configuration interface and store interaction template parameters, so that the interaction strategy decision module can generate output parameters corresponding to the visual projection information and the directional voice prompt information according to the interaction template parameters.

7. The system according to claim 6, characterized in that: The edge computing unit is equipped with a deep learning-based behavior recognition model, which is used to fuse and analyze the data collected by the high-definition camera and the millimeter-wave radar to output the behavioral intentions and / or risk assessment results of the traffic participants.

8. The system according to claim 6, characterized in that: The multimodal interactive terminal includes multiple laser projection devices, and the multiple laser projection devices adopt a multi-angle collaborative projection structure to cover the corresponding areas of the blind spots on the front, sides and / or rear of the vehicle. The directional acoustic array employs a directional acoustic propagation structure to propagate voice prompts to the target area.

9. The system according to claim 6, characterized in that: The interaction strategy decision module is configured as follows: When it is anticipated that a pedestrian intends to cross the road or enter the blind spot of a vehicle, the laser projection device is controlled to project a stop line and a warning box, and the directional sound wave array is controlled to output directional voice prompts to prompt the pedestrian to stop or give way to the vehicle. When it is predicted that a non-motorized vehicle will be within the blind spot of a vehicle turning at an intersection, multiple laser projection devices are controlled to collaboratively project turning radius markings and traffic guidance areas, and the vehicle's turn signals are linked through the vehicle bus to provide a prompt.

10. The system according to any one of claims 6 to 9, characterized in that: The vehicle bus is a CAN bus, and the system is linked with the vehicle's steering system and / or lighting system through the CAN bus to match the multimodal interaction strategy with the vehicle's driving state.