Off-board voice interaction method, system, device, storage medium, and program product
By collecting external voice signals and external perception information, and combining sound source localization and multi-dimensional criterion mechanisms, the reliability and naturalness issues of the external voice interaction system are solved, achieving efficient and safe voice response in the external environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-19
AI Technical Summary
Existing in-vehicle voice interaction systems struggle to reliably respond to voice requests in external environments, are susceptible to environmental noise interference, exhibit stiff responses, and lack verification of the authenticity of the sound source and the interaction intent.
By collecting external voice signals and external perception information, combining sound source localization and multi-dimensional criterion mechanisms to judge the request event, conduct credibility assessment, and use a large language model to generate a natural response.
It improves the reliability and safety of external voice interaction, generates natural and context-appropriate responses, and reduces the false trigger rate.
Smart Images

Figure CN121708894B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent cockpit and human-machine interaction technology, and in particular to methods, systems, devices, storage media and program products for external voice interaction. Background Technology
[0002] With the development of intelligent vehicle technology, in-vehicle voice interaction systems have been widely used in human-machine interaction scenarios. However, existing systems are mainly designed for occupants and are difficult to effectively respond to voice requests from people outside the vehicle, such as being asked questions by passersby in a parking lot, being called out when picking up or dropping off passengers, or being communicated with by other drivers in narrow roads.
[0003] The current solution has obvious shortcomings: on the one hand, directly amplifying the sound outside the car or relying on the driver to roll down the window and shout results in a poor interactive experience and affects driving safety; on the other hand, if the response is triggered solely based on the voice content, it is very easy to be falsely triggered by environmental noise, advertising voices or non-interactive human voices, and there is a lack of reliable verification of the authenticity of the sound source and the interactive intent.
[0004] Furthermore, even when a valid external request is identified, the system often uses a fixed broadcast template and cannot combine external perception information (such as speaker characteristics and environmental context) to generate natural and appropriate response content, resulting in stiff interaction and a lack of context awareness.
[0005] Therefore, existing technologies lack a solution for intelligent voice interaction that can balance trigger reliability, response safety, and natural expression when voice interaction is initiated by people outside the vehicle.
[0006] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0007] The main purpose of this application is to provide a method, system, device, storage medium, and program product for external voice interaction, aiming to solve the technical problem of how to balance trigger reliability, response safety, and natural expression in intelligent voice interaction when people outside the vehicle initiate voice interaction.
[0008] To achieve the above objectives, this application proposes an external vehicle voice interaction method, which includes:
[0009] Collect external voice signals and external sensor information;
[0010] Based on the external voice signal and the sound source localization result, determine whether a passive request event from an external person is triggered;
[0011] If a passive request event from a person outside the vehicle is triggered, a credibility assessment is performed based on the external perception information to determine whether the system is allowed to respond.
[0012] If a response is allowed, the corresponding speech text is generated using a pre-defined large language model;
[0013] Based on the spoken text, a corresponding voice output is generated and broadcast through the in-vehicle screen / speaker.
[0014] In one embodiment, the step of determining whether a passive request event by a person outside the vehicle is triggered based on the external voice signal and the sound source localization result includes:
[0015] The external voice signal is subjected to voice energy detection and keyword detection, and combined with the sound source localization result, it is determined whether a passive request event by an external person is triggered.
[0016] In one embodiment, the step of performing voice energy detection and keyword detection on the external voice signal, and combining the sound source localization result to determine whether a passive request event by an external person is triggered includes:
[0017] Obtain the voice energy value corresponding to the external voice signal, and determine whether the voice energy value reaches a preset energy threshold;
[0018] If the voice energy value reaches a preset energy threshold, then semantic recognition is performed on the external voice signal to obtain the voice text;
[0019] The voice text is matched with a preset external trigger word library to obtain keyword matching results. At the same time, semantic understanding is performed on the voice text to obtain the confidence level of the interaction intent.
[0020] Based on the preset sound source localization strategy, the external speech signal is localized to obtain spatial location information of the sound source.
[0021] When the spatial location information of the sound source meets the preset area conditions, and the keyword matching result or the confidence level of the interaction intent meets the preset detection conditions, the trigger type is determined to be a passive request from a person outside the vehicle.
[0022] In one embodiment, the step of performing a credibility assessment based on the externally perceived information to determine whether to allow the system to respond includes:
[0023] Based on preset environmental sound energy weights, speech recognition confidence weights, speech recognition confidence, sound source localization confidence weights, and sound source localization confidence, an acoustic evidence score is calculated.
[0024] Based on the external perception information and sound source localization results, a visual confidence evidence score is constructed.
[0025] The joint confidence score is calculated by combining the acoustic evidence score, the visual confidence evidence score, and the current vehicle context confidence score, wherein the current vehicle context confidence score is calculated based on vehicle state and environmental scene information.
[0026] A reliability assessment is performed based on the joint confidence level and a preset confidence threshold, and based on the reliability assessment results, it is determined whether the system response is allowed.
[0027] In one embodiment, before the step of calculating the acoustic evidence score based on preset ambient sound energy weights, speech recognition confidence weights, speech recognition confidence, sound source localization confidence weights, and sound source localization confidence, the method further includes:
[0028] The sound source localization reliability is calculated by the degree of matching between the sound source direction and visual detection; or
[0029] Based on the positioning uncertainty parameters output by the Time Difference of Arrival (TDOA) algorithm, the reliability of the sound source location is quantitatively evaluated to obtain the sound source positioning reliability.
[0030] In one embodiment, the step of constructing a visual confidence evidence score based on the external perception information and the sound source localization result includes:
[0031] At least one pedestrian detection box is obtained from the external sensing information, and the sound source localization result is projected onto the image coordinate system to generate a sound source direction projection area;
[0032] Calculate the intersection-union ratio between the pedestrian detection box and the projection area in the direction of the sound source;
[0033] Based on the pedestrian detection box, the head orientation information and gesture features of the corresponding pedestrian are extracted to obtain the gesture or action matching score and the confidence level of the hand key points;
[0034] Calculate the spatial proximity score between the pedestrian and the target vehicle;
[0035] The visual confidence score is calculated based on the intersection-union ratio, the interactive behavior features, and the pedestrians and vehicles.
[0036] In one embodiment, the step of performing a reliability assessment based on the joint confidence level and a preset confidence threshold, and determining whether to allow the system response based on the reliability assessment result, includes:
[0037] When the joint confidence level is greater than or equal to a preset high confidence threshold, and the gesture or action matching score reaches a preset score threshold, the system is allowed to respond; and / or
[0038] When the joint confidence level is between a preset low confidence threshold and a preset high confidence threshold, and the confidence level of the hand keypoint is greater than a preset hand keypoint confidence threshold, the driver is prompted to provide manual confirmation via the human-machine interface to determine whether to respond; and / or
[0039] When the joint confidence level is less than the preset low confidence threshold, the system is not allowed to respond.
[0040] Furthermore, to achieve the above objectives, this application also proposes an external voice interaction system, which includes:
[0041] The acquisition module is used to acquire external voice signals and external perception information of the target vehicle;
[0042] The triggering module is used to determine whether to trigger a passive request event from people outside the vehicle based on the external voice signal and the sound source localization result.
[0043] The evaluation module is used to perform a credibility assessment based on the external voice signal and the external perception information if a passive request event from a person outside the vehicle is triggered, and to determine whether the system should be allowed to respond.
[0044] The text generation module is used to generate corresponding speech text using a preset large language model if a response is allowed.
[0045] The voice generation module is used to generate corresponding voice output based on the voice text and broadcast it through the in-vehicle screen / speaker.
[0046] In addition, to achieve the above objectives, this application also proposes an external vehicle voice interaction device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the external vehicle voice interaction method as described above.
[0047] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the vehicle-to-everything (V2X) voice interaction method described above.
[0048] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the vehicle-to-everything (V2X) voice interaction method described above.
[0049] This application proposes an external voice interaction method, system, device, storage medium, and program product. The method includes: acquiring external voice signals and external perception information; determining whether a passive request event from an external person is triggered based on the external voice signals and sound source localization results; if a passive request event from an external person is triggered, performing a credibility assessment based on the external perception information and external voice signals to determine whether a system response is allowed; if a response is allowed, generating corresponding speech text using a preset large language model; generating corresponding speech output based on the speech text, and broadcasting it through the in-vehicle screen / speaker. This solution effectively eliminates long-distance interference, false triggering due to background noise, and speech interference from non-target directions by integrating a multi-dimensional criterion mechanism that combines acoustic feature detection, sound source spatial location, and environmental context credibility assessment, thereby improving the reliability and anti-interference capability of triggering passive request events from outside the vehicle; simultaneously, by introducing a large language model to perform semantic understanding and context-aware natural language generation of the user's original request, the system's expression of requests from external persons becomes more natural and interactive. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart illustrating an embodiment of the vehicle-to-vehicle voice interaction method of this application.
[0053] Figure 2 This is a flowchart illustrating Embodiment 2 of the vehicle-exterior voice interaction method of this application;
[0054] Figure 3 This is a flowchart illustrating Embodiment 3 of the vehicle-exterior voice interaction method of this application;
[0055] Figure 4 This is a schematic diagram of the module structure of the vehicle external voice interaction system according to an embodiment of this application;
[0056] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the vehicle-to-vehicle voice interaction method in the embodiments of this application.
[0057] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0058] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0059] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0060] The main solution of this application embodiment is as follows: collect external voice signals and external perception information of the target vehicle; determine whether a passive request event by an external person is triggered based on the external voice signals and sound source localization results; if a passive request event by an external person is triggered, perform a credibility assessment based on the external voice signals and external perception information to determine whether the system is allowed to respond; if the response is allowed, generate corresponding voice text through a preset large language model; generate corresponding voice output based on the voice text and broadcast it through the in-vehicle screen / speaker.
[0061] In this embodiment, for ease of description, the following description will focus on the external voice interaction system.
[0062] In existing technologies, when responding to voice requests from people outside the vehicle, in-vehicle systems often lack a comprehensive assessment of the validity of the sound source, making them susceptible to interference from environmental noise, non-target directions, or brief shouts, which can lead to false triggering. Furthermore, even when a response is triggered, it often relies on pre-set fixed phrases for broadcasting, failing to understand the original request's semantics. This results in rigid and disjointed responses that are difficult to meet real-world interaction needs.
[0063] This application provides a solution that, through a multi-dimensional criterion mechanism integrating acoustic feature detection, spatial sound source localization, and environmental context credibility assessment, enables vehicles to reliably identify valid external voice requests in complex external environments. Simultaneously, by introducing a large language model to perform semantic understanding and context-aware natural language generation on the original request, it generates an intelligent response that is semantically accurate, naturally expressed, and context-appropriate, thereby improving the accuracy, safety, and user experience of external vehicle interaction.
[0064] It should be noted that the executing entity in this embodiment can be a computing terminal (such as a personal computer, tablet computer, or smartphone) or a vehicle-specific electronic device (such as a smart cockpit controller, vehicle domain controller, or vehicle infotainment system). The following description uses an external voice interaction system as an example to illustrate this embodiment and the subsequent embodiments.
[0065] Based on this, embodiments of this application provide a method for external voice interaction, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle-external voice interaction method of this application.
[0066] In this embodiment, the external voice interaction method includes steps S10 to S50:
[0067] Step S10: Collect external voice signals and external perception information of the target vehicle;
[0068] In this embodiment, the interaction trigger mainly relies on the voice input obtained by the in-vehicle and out-of-vehicle sound pickup units, and combines the perception results of the intelligent driving system as context input.
[0069] Specifically, a multi-directional microphone array is deployed externally on the vehicle, covering areas at least the front (e.g., inside the front grille), rear (e.g., above the license plate), and below the side mirrors, creating a 360-degree coverage area to ensure clear capture of voices from different directions. Simultaneously, this external microphone array integrates wind noise and traffic noise suppression modules, with a built-in dedicated signal processing unit (DSP) running a deep learning noise reduction algorithm optimized for the complex acoustic environment of vehicles. This algorithm can separate and suppress strong wind noise generated at high speeds, as well as complex traffic noises such as engine noise, tire noise, and horn honking in urban areas, thereby efficiently extracting clear and effective human voice signals.
[0070] In addition, this embodiment also utilizes sensors such as vehicle-mounted cameras, LiDAR, and external microphones to obtain external perception information, including but not limited to the characteristics of the communication object (such as approximate age, language or dialect type, emotional state) and the current environmental state (such as noise level, scene type: school, hospital, highway, etc.).
[0071] Step S20: Based on the external voice signal and the sound source localization result, determine whether a passive request event from an external person is triggered.
[0072] It should be noted that the passive request event from people outside the vehicle refers to an interactive event triggered when people outside the vehicle (such as pedestrians or passengers of other vehicles) make a voice call to the vehicle and it is captured by the external microphone array, such as a passerby calling out, "Excuse me, can I borrow a light?"
[0073] Understandably, given the complexity of the external environment, relying solely on voice energy or keyword matching can easily lead to misjudging non-target voices as valid requests, resulting in erroneous system responses or frequent disturbances to the driver. Therefore, step S20, through a multi-dimensional collaborative discrimination mechanism that integrates voice content features with sound source localization results, avoids false triggering caused by a single signal source, thereby achieving high-precision and robust recognition of passive requests from real people outside the vehicle.
[0074] In one feasible embodiment, step S20 may include step S21:
[0075] Step S21: Perform voice energy detection and keyword detection on the external voice signal, and combine the sound source localization results to determine whether a passive request event by an external person is triggered.
[0076] In this embodiment, since there are a lot of non-target voice interference in the external environment of the vehicle (such as distant conversations, broadcasts, vehicle horns or wind noise), if only a single criterion (such as keyword matching or sound intensity threshold) is used for judgment, it is very easy to cause the system to be falsely triggered, which not only wastes computing resources, but may also interfere with the driver or cause social misunderstanding due to invalid broadcasts.
[0077] Therefore, by introducing a triple filtering mechanism of energy screening, semantic verification, and spatial constraints, invalid signals are effectively eliminated, ensuring that only real, close-range external voice requests with clear interactive intentions are responded to.
[0078] Furthermore, step S21 may also include sub-steps a21~a25:
[0079] Step a21: Obtain the voice energy value corresponding to the external voice signal, and determine whether the voice energy value reaches a preset energy threshold.
[0080] In this embodiment, an ambient sound signal is continuously collected by a microphone array deployed outside the vehicle, and a sliding window is used to calculate the energy of short speech frames to obtain the speech energy value E(t) of the current speech segment:
[0081]
[0082] Where x(n) represents the amplitude of the nth sampling point; N is the number of sampling points corresponding to the frame length.
[0083] Then, E(t) is compared with a preset energy threshold. The threshold is compared. It is dynamically set or statically configured based on actual vehicle external environmental noise statistics, with a typical value of approximately 70 dB equivalent sound pressure level (significantly higher than the 55–60 dB background noise of urban roads). If E(t) is measured in several consecutive frames (e.g., more than 3 frames),... If a potentially valid human voice message is detected, the process proceeds to the next step.
[0084] Step a22: If the voice energy value reaches a preset energy threshold, then perform semantic recognition on the external voice signal to obtain the voice text;
[0085] In this embodiment, the system invokes a lightweight automatic speech recognition (ASR) module to perform fast endpoint detection and speech-to-text on the speech signals that have passed the initial energy screening, generating the original speech text. It also outputs the recognition confidence level simultaneously. This ASR module is optimized for vehicle-to-everything (V2X) interaction scenarios, prioritizing the recognition accuracy of keywords and short phrases rather than the full-text transcription accuracy, in order to balance real-time performance and resource efficiency.
[0086] Step a23: Match the voice text with a preset external trigger word library to obtain keyword matching results, and at the same time perform semantic understanding on the voice text to obtain the confidence level of the interaction intent;
[0087] The external trigger word library includes two categories: one category is general or emergency trigger words, such as "Hello, car owner", "Sir, stop for a moment", "There is a problem with your car", "Check your tires", etc., which are used to cover common help or warning scenarios; the other category is user-defined trigger words, which allow car owners to enter personalized wake-up phrases through the vehicle settings interface.
[0088] In this embodiment, the original speech text is first processed. Keyword matching is performed; if any trigger word is matched, it is marked as a preliminary valid request. Simultaneously, to further improve robustness, the speech understanding model is run in parallel to process the matched trigger words. The system categorizes interactions by intent, identifying whether they fall under valid interaction categories such as inquiries, requests, alerts, or guidance, and outputs the corresponding interaction intent confidence score. Only interactions deemed to possess valid semantic meaning are considered to have valid semantic meaning if the keyword match is successful or the semantic intent confidence score exceeds a preset threshold.
[0089] Step a24: Based on the preset sound source localization strategy, perform voice source localization on the external voice signal to obtain the sound source localization result;
[0090] In this embodiment, the direction angle of the sound source is calculated by utilizing the multi-channel audio input of an external microphone array combined with beamforming and time-delay estimation (TDOA) technology. , Top view and pitch angle and distance estimation This forms the spatial location information of the sound source. ).
[0091] The sound source localization strategy predefines an effective interaction area, for example: within 0.5-3 meters to the side of the vehicle and 1-5 meters behind it, with a height between 0.8-2 meters. Sound sources outside this area are considered non-target interference.
[0092] Step a25: When the sound source localization result meets the preset area conditions, and the keyword matching result or the interaction intent confidence meets the preset detection conditions, the trigger type is determined to be a passive request from a person outside the vehicle.
[0093] In this implementation, a multi-condition fusion criterion is adopted: the passive request event of the person outside the vehicle is finally determined to be triggered only when the sound source localization result meets the preset area conditions and the keyword matching is successful or the interaction intent confidence meets the preset detection conditions.
[0094] The preset region condition can refer to the fact that the three-dimensional spatial coordinates of the sound source are located within a preset interactive sensitive area of the vehicle, which is typically defined as:
[0095] Horizontal direction: within a conical area of ±60° in front of the vehicle's front;
[0096] Distance range: 0.5 meters to 3.0 meters from the outer surface of the vehicle body;
[0097] Height range: 0.8 meters to 1.8 meters above the ground (covering the mouth height of an adult standing and speaking).
[0098] The preset detection condition can be any of the following conditions being met:
[0099] Keyword matching successful: The speech recognition result contains a preset wake word or command keyword (such as "Hello, Xiao A", "Open the door", "Start the air conditioner"), and the keyword matching score is ≥ the preset first threshold (such as 0.75).
[0100] Alternatively, the confidence level of the interaction intent meets the standard: the confidence level of the interaction intent category (such as "door control", "air conditioning adjustment", "car search request") output by the Natural Language Understanding (NLU) module is greater than or equal to the preset second threshold (such as 0.80), even if the exact keyword is not hit, the semantic intent is clear.
[0101] Through the above steps, combined with energy screening, semantic verification, and spatial constraint triple filtering, the false trigger rate caused by environmental noise, non-target direction speech, or brief invalid shouts is significantly reduced, ensuring that the system only responds to real, valid, and close-range external interaction requests, thereby achieving highly reliable and low-interference passive voice perception capabilities.
[0102] Step S30: If a passive request event from an outside person is triggered, a credibility assessment is performed based on the outside voice signal and the external perception information to determine whether the system is allowed to respond.
[0103] In this embodiment, since the external interaction scenarios are complex and varied, relying solely on a single modality of voice or vision is susceptible to interference from factors such as occlusion, noise, and misidentification, which may cause the system to respond to non-genuine requests. This not only wastes resources but may also lead to misoperation or social misunderstanding. Therefore, by constructing a multimodal credibility assessment mechanism that integrates external voice signals and multi-source external perception information, invalid responses caused by single-modal misjudgment are avoided, thereby achieving highly robust and secure external voice interaction decision-making.
[0104] In one feasible embodiment, step S30 includes steps S31 to S34:
[0105] Step S31: Based on the external voice signal and / or the external perception information, obtain the sound source location confidence score, and combine it with the preset ambient sound energy weight, voice recognition confidence weight, sound source location confidence weight and sound energy to calculate the acoustic evidence score.
[0106] It should be noted that the ambient sound energy weight is used to control the contribution of ambient sound intensity to the final score; the speech recognition confidence weight is used to measure the contribution of the reliability of the speech recognition result to the final score; the sound source localization confidence weight is used to measure the reliability of the sound source spatial matching; and the sound energy reflects the prominence of the currently detected speech signal intensity relative to the background noise.
[0107] In this embodiment, an acoustic evidence score is constructed by building acoustic confidence. ):
[0108]
[0109] In the formula, This represents the final overall confidence score, used to determine whether the voice trigger is valid. The value is usually in ([0,1]) or a normalizable range. Indicates the environmental sound energy weight; ASR confidence weight; The confidence score output by the speech recognition module (ASR) is the probability of recognizing the trigger word. Assign location confidence weights to sound sources; Location reliability for sound sources; This represents normalized audio energy.
[0110] Step S32: Based on the external perception information and the sound source localization result, construct a visual confidence evidence score;
[0111] Furthermore, by calculating multi-factor confidence scores, the system determines whether the visually detected target is a real interactive object, thereby improving the accuracy of voice triggering. To this end, the system constructs a visual confidence evidence score (...). This is used to quantify the degree of spatial, behavioral, and semantic matching between the visual target and the current voice-triggered intent.
[0112] Specifically, the visual confidence evidence score Calculated using the following weighted fusion formula:
[0113]
[0114] In the formula, This represents the final visual confidence score, which measures the degree of matching between the visually detected target and the triggering intent. The value is usually in the range of ([0,1]) or normalized interval, and the higher the value, the more reliable the result. The degree of overlap between the visual inspection box and the projection of the sound source direction is used to measure the visual inspection box ( ) and sound source projection frame ( The degree of matching; weighting coefficients Control the contribution of this indicator to the overall score; Gesture confidence is calculated by using a visual model to determine whether a user's gesture matches a triggering action (such as waving or raising a hand), and the highest match confidence score is selected. Weights are assigned accordingly. The importance of controlling gesture matching; This represents the confidence score of hand keypoints, obtained through a human keypoint detection model, with the maximum value being taken; weights. The impact of key hand information on scoring; This represents the user's spatial proximity score to the vehicle, calculated using cameras or radar to determine the distance between the user and the vehicle; the closer the distance, the higher the confidence score. (Weight) Determine the contribution of spatial proximity to the overall score.
[0115] Step S33: Calculate the joint confidence score by combining the acoustic evidence score, the visual confidence evidence score, and the current vehicle context confidence score;
[0116] The current vehicle context confidence score can be calculated using vehicle status and environmental scene information. This vehicle status and environmental scene information includes, but is not limited to, vehicle speed being 0, door opening / closing, nighttime / school zone conditions, etc.
[0117] In this embodiment, a scoring fusion unit is designed through a multimodal triggering determination mechanism to generate a joint confidence score by integrating acoustic evidence scores, visual confidence evidence, and current vehicle context confidence. .
[0118] Specifically, joint confidence Defined as:
[0119]
[0120] In the formula, Score the acoustic evidence; This serves as evidence of visual confidence. The current vehicle context confidence score reflects the reasonableness of the current vehicle state and environment (e.g., vehicle speed is 0, doors are closed, in a school zone, or in night mode); α, β, and γ are adjustable weights; σ is a normalization function (e.g., sigmoid) to ensure... ∈[0,1].
[0121] The weights are dynamically adjusted using an adaptive mechanism: the acoustic weight α is reduced in high-noise or high-speed driving scenarios, and the visual weight β is reduced in scenarios with obstructed vision or at night. For example:
[0122]
[0123] In the formula, Let η be the SNR mapping function, and η be the learning rate.
[0124] Step S34: Perform a reliability assessment based on the joint confidence level and the preset confidence threshold, and determine whether to allow the system to respond based on the reliability assessment results.
[0125] Understandably, in order to balance response sensitivity and false touch suppression in complex urban scenarios, the system adopts a hierarchical decision-making mechanism, that is, by setting a high confidence threshold. With low confidence threshold It also combines auxiliary evidence such as gestures and visual orientation to dynamically determine interactive behavior.
[0126] In another feasible embodiment, step S34 may further include steps S341 to S343:
[0127] Step S341: When the joint confidence level is greater than or equal to a preset high confidence threshold, and the gesture or action matching score reaches a preset score threshold, the system is allowed to respond.
[0128] In this embodiment, if And the gesture or action matching score reaches a preset score threshold (such as gesture confidence). =1, if someone is facing the vehicle), then it is determined to be a highly reliable passive request from outside the vehicle. The ASR and large language model are then invoked to generate a response, which is then broadcast in real time through the vehicle's external speakers.
[0129] Step S342: When the joint confidence level is between the preset low confidence threshold and the high confidence threshold, and the confidence level of the hand key point is greater than the preset hand key point confidence threshold, the driver is prompted to manually confirm through the human-machine interface to determine whether to respond.
[0130] In this embodiment, if And there is confidence in the key points of the hand. (Preset confidence threshold for hand key points) If there is a potential interaction intention but insufficient evidence, the system will temporarily store the request in the manual confirmation queue and display "Someone is calling, please confirm" on the external screen or push a notification to the in-vehicle HMI, allowing the driver to decide whether to respond.
[0131] Step S343: When the joint confidence level is less than the preset low confidence level threshold, the system response is not allowed.
[0132] In this embodiment, if At this point, the system determines that the input is noise interference or invalid, ignores the event, and records relevant logs (including speech segments, perception data, and confidence components) for subsequent model iteration and optimization.
[0133] Furthermore, to avoid flickering or repeated triggering caused by signal fluctuations, the system introduces a time window and cooling strategy: once a trigger response is initiated or the confirmation process is entered, a hold time window is activated. (For example, 2 seconds), during which new requests from the same sound source are blocked; at the same time, a minimum response interval is set for the same interactive object. (For example, 5 seconds) to prevent high-frequency interference.
[0134] Through the above steps, acoustic evidence (speech energy, recognition confidence, sound source location), visual evidence (target detection, gestures, spatial proximity), and vehicle context (vehicle speed, scene mode, etc.) are integrated to construct a multi-dimensional cross-validation system. This effectively distinguishes between genuine interaction intentions and environmental noise, distant shouts, or interference from non-target personnel, significantly reducing the false trigger rate. At the same time, the joint confidence adopts an adjustable weight and normalization fusion strategy and supports dynamic adjustment of the contribution weights of each modality according to environmental conditions, enabling the system to maintain stable and reliable judgment performance in diverse scenarios such as day / night, urban / suburban, and stationary / low-speed environments.
[0135] Step S40: If a response is allowed, the corresponding speech text is generated using a preset large language model;
[0136] In this embodiment, after the credibility assessment result in step S30 confirms that the system can respond, the system inputs the original voice request into a preset Large Language Model (LLM). This LLM has been fine-tuned for external vehicle interaction scenarios and has the ability to understand and generate common requests for help, inquiries, and warnings. The LLM not only restates the user's request but also combines the current vehicle status, external environmental context, and social interaction norms to generate semantically accurate, appropriately toned, and appropriately long response text.
[0137] For example, when a person outside the vehicle calls out, "Sir, your headlights are still on!" and the vehicle is turned off, the large language model can generate, "Thank you for reminding me! I'll turn off the headlights right away," instead of mechanically repeating "Headlights are still on." This process achieves intelligent conversion from the original intent to natural, polite, and context-appropriate speech text.
[0138] Step S50: Generate corresponding voice output based on the voice text, and broadcast it through the in-vehicle screen / speaker.
[0139] In this embodiment, a text-to-speech (TTS) engine is invoked to load a sound model that matches the current interaction scenario (such as a gentle female voice, a clear male voice, or a dialect voice), and the speech text generated in step S40 is synthesized into a high-quality audio signal according to preset synthesis parameters (including speech rate, volume, and emotional style).
[0140] The audio signal is then played through the vehicle's audio system, and optionally, text content is displayed on the in-vehicle screen (such as "Responding to an external call: 'Thank you for the reminder! I will turn off the lights immediately') to enhance the driver's awareness of the interaction status.
[0141] It should be noted that the "in-vehicle announcement" here is intended to inform the driver that the system has responded to a request from outside the vehicle, rather than to make a direct external announcement.
[0142] The above-described method collects external voice signals and external perception information. Based on the external voice signals and sound source localization results, it determines whether a passive request event from an external person is triggered. If a passive request event is triggered, a credibility assessment is performed based on the external perception information to determine whether a system response is allowed. If a response is allowed, a corresponding voice text is generated using a pre-set large language model. A corresponding voice output is generated based on the voice text and broadcast through the in-vehicle screen / speaker. This solution effectively eliminates long-distance interference, false triggering due to background noise, and non-target direction voice interference by integrating a multi-dimensional criterion mechanism that combines acoustic feature detection, sound source spatial location, and environmental context credibility assessment. This improves the reliability and anti-interference capability of triggering passive request events from outside the vehicle. Simultaneously, by introducing a large language model to perform semantic understanding and context-aware natural language generation of the user's original request, the system's naturalness and user-friendly interaction in responding to requests from external persons are enhanced.
[0143] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Specifically defining step S31, the external voice interaction method further includes steps S311-S312:
[0144] Step S311: Based on the degree of matching between the sound source localization result and the direction of the pedestrian or human target detected in the external perception information;
[0145] In this embodiment, the sound source localization result (including the sound source azimuth angle) is used to locate the sound source. ), pitch angle ( ) and distance estimation ( The sound source direction projection area is formed by projecting the sound source direction projection area onto the 2D image plane of the vehicle's vision sensor (such as a forward-facing camera). Simultaneously, one or more pedestrian detection boxes are acquired from external perception information. By calculating the spatial overlap (e.g., intersection-over-union ratio) between the sound source direction projection area and each pedestrian detection box, the consistency between acoustic and visual targets is quantified, and a sound source localization confidence score is generated accordingly: the higher the overlap, the more likely the sound source corresponds to a real-world interactive object, and the higher the confidence score; if there is no effective visual target matching, the confidence score decreases.
[0146] Step S312: Based on the external voice signal, the location confidence of the sound source is calculated using the Time Difference of Arrival (TDOA) algorithm.
[0147] In this embodiment, the uncertainty indicators such as the covariance matrix, residual error, or geometric precision factor output by the TDOA (Time Difference of Arrival) sound source localization algorithm are further used to assess the increased uncertainty of the current sound source localization result. If the uncertainty increases, the system will correspondingly reduce the sound source localization confidence; otherwise, a higher confidence level will be assigned to form the final sound source localization confidence.
[0148] The above-described methods provide dual verification of sound source location quality from two dimensions: cross-modal consistency and inherent reliability of sound source localization. Firstly, the degree of matching between the sound source direction and the visually detected target determines whether the speech originates from a real pedestrian. Secondly, the uncertainty parameters output by the TDOA algorithm assess the credibility of the sound source localization result in the current acoustic environment. This confidence level serves as a key input in the subsequent calculation of the acoustic evidence score, effectively suppressing misjudgments caused by sound source localization drift, multipath interference, or visual occlusion, and significantly improving the perception accuracy and response reliability of the vehicle-to-everything (V2X) voice interaction system in complex urban scenarios.
[0149] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S32 also includes steps A11 to A15:
[0150] Step A11: Obtain at least one pedestrian detection box from the external perception information, and project the sound source localization result onto the image coordinate system to generate a sound source direction projection area;
[0151] It should be noted that the pedestrian detection box (denoted as...) A pedestrian target area (PTA) is a rectangular boundary region identified by the vehicle's visual perception module (such as a deep learning-based object detection model) in image frames captured by a camera. This PTA is used to characterize the position and scale of a pedestrian target. The boundary region is represented by the coordinates of the upper left corner, width, and height (or center point and dimensions) in the image coordinate system, and is used to identify the spatial extent of potential external interactive objects in the field of view.
[0152] The sound source localization result includes the sound source azimuth angle ( ), pitch angle ( ) and distance estimation ( ).
[0153] In this embodiment, human detection algorithms (such as YOLO and Faster R-CNN) are used to output the bounding boxes of pedestrians. (Each pedestrian corresponds to a box), represented in pixel coordinates. ,in The top left corner The bottom right corner covers the entire body or upper body of the pedestrian.
[0154] Simultaneously, the 3D spatial direction formed by the sound source localization result is projected onto the 2D image plane of the vehicle's visual sensor (such as a forward-facing camera) to form a sound source direction projection area.
[0155] Specifically, firstly, based on the intrinsic parameters (focal length) of the forward-looking or side-looking camera... , Principal point coordinates ), azimuth angle Pitch angle Pixel offset converted to the image plane:
[0156] Horizontal pixel offset: ;
[0157] Vertical pixel offset: ;
[0158] The coordinates of the center pixel of the projection are ( , ).
[0159] Furthermore, considering the impact of sound source distance on positioning accuracy, the angle of the projection area is adjusted:
[0160] When distance estimation When the depth is ≤5m, the horizontal angle is set to ±10°; when The angle was extended from >5m to ±15° to compensate for errors in locating distant sound sources; the vertical angle was fixed at ±5°, corresponding to the typical height range of pedestrians (1.2–1.8 m). Ultimately, this fan-shaped region was simplified into a rectangular bounding box. = , which is the projection area of the sound source direction in the image.
[0161] Step A12: Calculate the intersection-union ratio between the pedestrian detection box and the projection area in the direction of the sound source;
[0162] In this embodiment, each pedestrian detection box is calculated. With the projection area of the sound source The intersection-union ratio (IUU) serves as a quantitative indicator of the spatial overlap between the two entities. The calculation formula is as follows:
[0163]
[0164] In the formula, The bounding box of the pedestrian target detected by the visual perception module in the image coordinate system is used to characterize the spatial location of potential interactive objects. This represents the projection area of the sound source direction formed after projecting the sound source direction onto the same image plane based on the sound source localization result. It is used to characterize the visual corresponding region of the source of the speech signal.
[0165] If multiple pedestrian detection boxes exist Then take all and The value with the largest intersection-union ratio is used as the final spatial consistency score; if no pedestrians are detected (i.e. the detection box set is empty), the score is set to zero, indicating that there is no visual evidence to support the existence of an interactive object in the direction of the sound source.
[0166] The value of IoU_with_DOA ranges from [0, 1]. When its value is close to 1, it indicates that the pedestrian detected by vision is completely located in the area pointed to by the sound source, and the acoustic and visual information are highly consistent, which significantly enhances the credibility of the target as a real interactor. When its value is close to 0, it indicates that the pedestrian is deviating from the direction of the sound source or there is no corresponding visual target, which weakens the reliability of cross-modal matching.
[0167] In practical applications, the system can set basic thresholds based on historical data statistics (for example, an intersection-union ratio ≥ 0.5 is considered a strong spatial association, and ≥ 0.3 is considered a weak association), and dynamically adjust them according to the complexity of the scenario—for example, appropriately lowering the threshold in densely populated areas to reduce missed detections, and raising the threshold in open road sections to suppress false matches.
[0168] Through the above steps, the cross-union ratio effectively establishes a bridge between the spatial directivity of acoustic localization and the entity existence of visual detection, becoming a key criterion for verifying whether a sound source corresponds to a real person in multimodal fusion, and significantly reducing the risk of false triggering caused by relying solely on speech signals.
[0169] Step A13: Based on the pedestrian detection box, extract the head orientation information and gesture features of the corresponding pedestrian to obtain the gesture or action matching score and the confidence of the hand key points;
[0170] Furthermore, to reduce acoustic errors, the visual confidence evidence score must be based on strong cues such as "whether someone is in the direction of the specified sound source, or whether they are paying attention to vehicles / heading towards vehicles".
[0171] Specifically, first detect pedestrian / body frames ( ), calculate the distance to the vehicle ( (Using depth or visual scale and camera calibration), we obtain ( (Velocity vector).
[0172] Then, estimate each Head facing , looking up If the angle between the line-of-sight projection and the vehicle's direction is less than a preset angle threshold... (For example, ≤ 25°) is considered "vehicle-oriented" evidence. This embodiment defines a binary index:
[0173]
[0174] in, Whether the i-th pedestrian is facing the target vehicle; This is the preset threshold for the line-of-sight angle.
[0175] Take the maximum value among all pedestrians as This metric is used to characterize whether there is a strong "vehicle-oriented" interaction intent in the current scene. As a head orientation confidence score, it is incorporated into the calculation of the visual confidence evidence score to enhance the ability to distinguish real interaction objects.
[0176] Furthermore, to improve trigger accuracy in noisy environments, the system can be linked to the vehicle's camera. When the system detects that the person making the call is simultaneously making a specific gesture (such as waving continuously), it will significantly increase the trigger weight of that call, achieving accurate recognition of the combined action of "waving and calling," effectively reducing the false trigger rate.
[0177] Specifically, the system analyzes short video sequences using a human pose estimation model, extracts hand key points, identifies preset interactive actions, and outputs the confidence scores of the hand key points. ( ), as quantitative evidence of the intention behind the gesture.
[0178] Step A14: Calculate the spatial proximity score between the pedestrian and the target vehicle;
[0179] In this embodiment, the distance between the user and the vehicle is calculated based on the camera or radar, and used as a spatial proximity score. (The closer the distance, the higher the confidence level).
[0180] Step A15: Calculate the visual confidence score based on the intersection-union ratio, the interactive behavior features, and the spatial proximity score.
[0181] Based on the intersection-exchange ratio of the pedestrian detection frame and the projection area of the sound source direction. Gesture or action matching score Confidence of key points in the hand and spatial proximity score By combining relevant weighting coefficients, the visual confidence score of visual evidence is calculated. ).
[0182] The above-described methods effectively integrate four key cues: spatial alignment, orientation intent, interactive actions, and physical distance. This not only significantly improves the accuracy of interactive object recognition in noisy or complex urban scenarios but also provides highly reliable visual input for multimodal joint confidence decision-making, thereby greatly reducing the risk of false triggering and enhancing system robustness and user experience.
[0183] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vehicle-to-vehicle voice interaction method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0184] This application also provides an external voice interaction system; please refer to... Figure 4 The external voice interaction system includes:
[0185] The acquisition module 10 is used to acquire external voice signals and external perception information of the target vehicle;
[0186] Trigger module 20 is used to determine whether to trigger a passive request event from people outside the vehicle based on the external voice signal and the sound source localization result;
[0187] The evaluation module 30 is used to perform a credibility evaluation based on the external voice signal and the external perception information if a passive request event from a person outside the vehicle is triggered, and to determine whether the system response is allowed.
[0188] The text generation module 40 is used to generate corresponding speech text using a preset large language model if a response is allowed.
[0189] The voice generation module 50 is used to generate corresponding voice output based on the voice text and broadcast it through the in-vehicle screen / speaker.
[0190] The vehicle-to-vehicle voice interaction system provided in this application, employing the vehicle-to-vehicle voice interaction method described in the above embodiments, can solve the technical problem of how to balance trigger reliability, response safety, and natural expression in intelligent voice interaction when a person outside the vehicle initiates voice interaction. Compared with the prior art, the beneficial effects of the vehicle-to-vehicle voice interaction system provided in this application are the same as those of the vehicle-to-vehicle voice interaction method provided in the above embodiments, and other technical features of the vehicle-to-vehicle voice interaction system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0191] This application provides an external voice interaction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the external voice interaction method in the above embodiment 1.
[0192] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the vehicle-to-vehicle voice interaction device in the embodiments of this application. The vehicle-to-vehicle voice interaction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The external voice interaction device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0193] like Figure 5As shown, the external voice interaction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the external voice interaction device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the external voice interaction device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show external voice interaction devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0194] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0195] The vehicle-to-vehicle voice interaction device provided in this application, employing the vehicle-to-vehicle voice interaction method described in the above embodiments, can solve the technical problem of how to balance trigger reliability, response security, and natural expression in intelligent voice interaction when a person outside the vehicle initiates voice interaction. Compared with the prior art, the beneficial effects of the vehicle-to-vehicle voice interaction device provided in this application are the same as those of the vehicle-to-vehicle voice interaction method provided in the above embodiments, and other technical features of this vehicle-to-vehicle voice interaction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0196] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0197] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0198] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the vehicle-to-vehicle voice interaction method in the above embodiments.
[0199] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0200] The aforementioned computer-readable storage medium may be included in the external voice interaction device; or it may exist independently and not be installed in the external voice interaction device.
[0201] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the external voice interaction device, the external voice interaction device performs the following actions: collects external voice signals and external perception information of the target vehicle; determines, based on the external voice signals and sound source localization results, whether a passive request event from an external person is triggered; if a passive request event from an external person is triggered, performs a credibility assessment based on the external voice signals and external perception information to determine whether a system response is allowed; if a response is allowed, generates corresponding speech text using a preset large language model; generates corresponding speech output based on the speech text, and broadcasts it through the in-vehicle screen / speaker.
[0202] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0203] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0204] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0205] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described external vehicle voice interaction method. This solves the technical problem of how to balance trigger reliability, response security, and natural expression in intelligent voice interaction when a person outside the vehicle initiates voice interaction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the external vehicle voice interaction method provided in the above embodiments, and will not be repeated here.
[0206] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the vehicle-exterior voice interaction method described above.
[0207] The computer program product provided in this application solves the technical problem of how to balance trigger reliability, response safety, and natural expression in intelligent voice interaction when initiated by people outside the vehicle. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the external voice interaction method provided in the above embodiments, and will not be repeated here.
[0208] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for external voice interaction, characterized in that, The external voice interaction method includes: Collect external voice signals and external perception information of the target vehicle; Based on the external voice signal and the sound source localization result, determine whether a passive request event from an external person is triggered; If a passive request event from an outside person is triggered, a credibility assessment is performed based on the outside voice signal and the external perception information to determine whether the system is allowed to respond, including: Based on the external voice signal and / or the external perception information, the sound source location confidence is obtained, and combined with the preset environmental energy weight, voice recognition confidence weight, sound source location confidence weight and sound energy, the acoustic evidence score is calculated. Based on the external perception information and the sound source localization result, a visual confidence evidence score is constructed. The joint confidence score is calculated by combining the acoustic evidence score, the visual confidence evidence score, and the current vehicle context confidence score, wherein the current vehicle context confidence score is calculated based on the current vehicle state and environmental scene information. A reliability assessment is performed based on the joint confidence level and a preset confidence threshold, and based on the reliability assessment results, it is determined whether the system response is allowed. If a response is allowed, the corresponding speech text is generated using a pre-defined large language model; Based on the spoken text, a corresponding voice output is generated and broadcast through the in-vehicle screen / speaker.
2. The vehicle-exterior voice interaction method as described in claim 1, characterized in that, The step of determining whether a passive request event by an outside person is triggered based on the external voice signal and the sound source localization result includes: The external voice signal is subjected to voice energy detection and keyword detection, and combined with the sound source localization results, it is determined whether a passive request event by an external person is triggered.
3. The vehicle-exterior voice interaction method as described in claim 2, characterized in that, The steps of performing voice energy detection and keyword detection on the external voice signal, and combining the sound source localization results to determine whether a passive request event by an external person is triggered include: Obtain the voice energy value corresponding to the external voice signal, and determine whether the voice energy value reaches a preset energy threshold; If the voice energy value reaches a preset energy threshold, then semantic recognition is performed on the external voice signal to obtain the voice text; The voice text is matched with a preset external trigger word library to obtain keyword matching results. At the same time, semantic understanding is performed on the voice text to obtain the confidence level of the interaction intent. Based on the preset sound source localization strategy, the external voice signal is localized to obtain the sound source localization result. When the sound source localization result meets the preset area conditions, and the keyword matching result or the interaction intent confidence meets the preset detection conditions, it is determined that a passive request event from people outside the vehicle is triggered.
4. The vehicle-exterior voice interaction method as described in claim 1, characterized in that, The step of obtaining the sound source localization confidence based on the external voice signal and / or the external perception information includes: Based on the degree of matching between the sound source localization result and the direction of the pedestrian or human target detected in the external sensing information, the sound source localization reliability is calculated; or Based on the external voice signal, the location confidence of the sound source is calculated using the Time Difference of Arrival (TDOA) algorithm.
5. The vehicle-exterior voice interaction method as described in claim 1, characterized in that, The step of constructing a visual confidence evidence score based on the external perception information and the sound source localization result includes: At least one pedestrian detection box is obtained from the external sensing information, and the sound source localization result is projected onto the image coordinate system to generate a sound source direction projection region. Calculate the intersection-union ratio between the pedestrian detection box and the projection area in the direction of the sound source; Based on the pedestrian detection box, the head orientation information and gesture features of the corresponding pedestrian are extracted to obtain the gesture or action matching score and the confidence level of the hand key points; Calculate the spatial proximity score between the pedestrian and the target vehicle; The visual confidence score is calculated based on the intersection-union ratio, the gesture or action matching score, the confidence score of the hand key points, and the spatial proximity score.
6. The vehicle exterior voice interaction method as described in claim 5, characterized in that, The step of performing a reliability assessment based on the joint confidence level and a preset confidence threshold, and determining whether to allow the system response based on the reliability assessment results, includes: When the joint confidence level is greater than or equal to a preset high confidence threshold, and the gesture or action matching score reaches a preset score threshold, the system is allowed to respond; and / or When the joint confidence level is between a preset low confidence threshold and a preset high confidence threshold, and the confidence level of the hand keypoint is greater than a preset hand keypoint confidence threshold, the driver is prompted to manually confirm via the human-machine interface to determine whether to respond; and / or When the joint confidence level is less than the preset low confidence threshold, the system is not allowed to respond.
7. An external voice interaction system, characterized in that, The external voice interaction system includes: The acquisition module is used to acquire external voice signals and external perception information of the target vehicle; The triggering module is used to determine whether to trigger a passive request event from people outside the vehicle based on the external voice signal and the sound source localization result. An evaluation module is used to perform a credibility assessment based on the external voice signal and the external perception information, and determine whether to allow the system to respond, if a passive request event from an external person is triggered, including: Based on the external voice signal and / or the external perception information, the sound source location confidence is obtained, and combined with the preset environmental energy weight, voice recognition confidence weight, sound source location confidence weight and sound energy, the acoustic evidence score is calculated. Based on the external perception information and the sound source localization result, a visual confidence evidence score is constructed. The joint confidence score is calculated by combining the acoustic evidence score, the visual confidence evidence score, and the current vehicle context confidence score, wherein the current vehicle context confidence score is calculated based on the current vehicle state and environmental scene information. A reliability assessment is performed based on the joint confidence level and a preset confidence threshold, and based on the reliability assessment results, it is determined whether the system response is allowed. The text generation module is used to generate corresponding speech text using a preset large language model if a response is allowed. The voice generation module is used to generate corresponding voice output based on the voice text and broadcast it through the in-vehicle screen / speaker.
8. An external voice interaction device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle-to-everything (V2X) voice interaction method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vehicle-exterior voice interaction method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the vehicle-exterior voice interaction method as described in any one of claims 1 to 6.