A vehicle-mounted false wake-up filtering method and system based on continuity of sound source trajectory
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]但在现有技术中,难以有效区分车外快速掠过的声源与车内乘员发出的真实唤醒语音,也无法准确识别由车内多人先后发音拼接而成的近似唤醒词
候选唤醒词被划分为具有独立定位价值的音素证据单元,系统能够逐段检查声源方向是否发生切换,从而识别不同乘员先后发音所形成的拼接式误唤醒。
Smart Images

Figure CN122551797A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of vehicle-mounted voice interaction and acoustic signal processing technology, and in particular to a vehicle-mounted false wake-up filtering method and system based on the continuity of sound source trajectories. Background Technology
[0002] With the rapid development of intelligent connected vehicles, in-vehicle voice interaction systems have become the core entry point for human-machine interaction in vehicles, widely used in scenarios such as navigation control, multimedia playback, and vehicle function adjustment. Existing in-vehicle voice wake-up technologies typically employ microphone arrays to collect in-vehicle audio signals and use a voice recognition engine to detect the presence of preset wake-up words. Once a candidate wake-up word is detected, existing technologies primarily rely on recognition confidence, contextual semantic information, or secondary recognition results for filtering to eliminate interference from near-sounding words during conversation. Some existing solutions introduce sound source localization technology, treating the complete wake-up word as a whole for directional determination, or using a joint model to output the wake-up result and sound source location. Other solutions use frequency domain signal weights to weightedly locate multiple microphone signals, thereby determining whether the sound source originates from the driver, passenger, or rear seats. Furthermore, some technologies utilize particle filtering or Kalman filtering to continuously track multiple sound sources to avoid overlapping sound source trajectories.
[0003] However, existing technologies struggle to effectively distinguish between rapidly passing sound sources outside the vehicle and the actual wake-up voices emitted by occupants inside the vehicle. They also cannot accurately identify approximate wake-up words composed of multiple voices spoken sequentially inside the vehicle. Existing solutions often fail to verify whether the phonemes constituting the same wake-up word actually originate from the same sound source, leading to false wake-ups or missed detections when the vehicle passes pedestrians, toll booths, or when multiple people inside the vehicle are conversing. Summary of the Invention
[0004] This application provides a vehicle-mounted false wake-up filtering method and system based on the continuity of sound source trajectory to solve the above problems.
[0005] In a first aspect, this application provides an in-vehicle false wake-up filtering method based on the continuity of sound source trajectories, the method comprising: S1. Obtain the multi-channel audio signal corresponding to the candidate wake-up word and the vehicle pose change information during the duration of the candidate wake-up word; S2. Perform phoneme timing alignment on the candidate wake words to divide them into multiple phoneme evidence units, and determine at least one candidate sound source direction for each phoneme evidence unit. S3. Based on the vehicle pose change information, the directions of each candidate sound source are represented in the vehicle coordinate system and the road reference coordinate system respectively. The fitting result of the in-vehicle sound source is determined based on the matching error between each candidate sound source direction and the activity area of the occupants in the vehicle. The fitting result of the external sound source is determined based on the spatial aggregation error of the sound source observation rays corresponding to different vehicle positions. S4. According to the timing of each phoneme evidence unit, associate the candidate sound source direction of each phoneme evidence unit with the corresponding candidate sound source trajectory, and determine whether the key phonemes in the candidate wake-up word continuously belong to the same candidate sound source trajectory. S5. Based on the sound source attribution results of the key phonemes, the fitting results of the in-vehicle sound sources and the fitting results of the external sound sources, determine whether the candidate wake-up word is a valid wake-up or a false wake-up, and control the in-vehicle voice interaction system to output or prevent the output of the wake-up command.
[0006] Optionally, the candidate wake word is divided into multiple phoneme evidence units, including: Based on the posterior probability of the phonemes corresponding to each audio frame, audio frames that consecutively correspond to the same target phoneme are grouped into phoneme intervals. For phoneme intervals that cannot form a reliable independent sound source direction, the spectral continuity between the phoneme interval and the preceding and following phoneme intervals is calculated respectively, and the phoneme interval is merged with the adjacent phoneme intervals with higher spectral continuity to obtain the phoneme evidence unit.
[0007] Optionally, the sound source direction of each phoneme evidence unit can be determined separately, including: Based on the phoneme posterior probability of each audio frame in the phoneme evidence unit, each audio frame is weighted. Each frequency point is weighted according to the stability of the phase difference or time of arrival difference between different microphone channels in consecutive audio frames. A spatial acoustic spectrum is generated using weighted audio frames and frequency points, and at least one direction in the spatial acoustic spectrum that satisfies the local peak condition is determined as the candidate sound source direction of the phoneme evidence unit.
[0008] Optionally, the in-vehicle occupant activity area includes occupant head activity areas corresponding to each seat in the vehicle, and the occupant head activity areas are adjusted according to the fore-and-aft position and backrest angle of the corresponding seats. For each occupant's head movement area, a direction intersecting with that occupant's head movement area is selected from the candidate sound source directions of each phoneme evidence unit. The average angular deviation between the selected direction and the center direction of the occupant's head movement area, as well as the average angular change between adjacent selected directions, are calculated. The sum of these two is taken as the region matching error. The minimum value of the region matching error corresponding to each occupant's head movement area is taken as the in-vehicle sound source fitting result, and the corresponding occupant's head movement area is taken as the region constraint of the candidate sound source trajectory.
[0009] Optionally, each phoneme evidence unit can be associated with its corresponding candidate sound source trajectory, including: Based on the sound source direction and corresponding time of the associated phoneme evidence unit in the candidate sound source trajectory, predict the sound source direction of the candidate sound source trajectory at the corresponding time of the current phoneme evidence unit. Among the candidate sound source directions that intersect with the regional constraints of the candidate sound source trajectories, the sound source directions to be associated are determined in ascending order of angular deviation from the predicted sound source directions. When the difference between the minimum angle deviation and the second smallest angle deviation is not greater than the preset angle discrimination tolerance, the speaker feature similarity between the current phoneme evidence unit and the previous associated phoneme evidence unit in each candidate sound source trajectory is sorted from largest to smallest; when the difference between the two speaker feature similarities in the first sorting is not greater than the preset similarity discrimination tolerance, the candidate sound source trajectory corresponding to the current phoneme evidence unit is determined from smallest to largest according to the difference in the energy ratio of the multi-microphone channels.
[0010] Optionally, determining the fitting result of the external sound source includes: Based on the vehicle position and attitude at the corresponding time of each phoneme evidence unit, the corresponding sound source direction is converted into a sound source observation ray in the road reference coordinate system; The road reference space around the vehicle is divided into spatial grids, and the average distance from the observation ray of each sound source to the center of each spatial grid is calculated. The spatial grid with the smallest average distance is identified as the candidate external sound source region, and the fitting result of the external sound source is determined based on the ratio of the average distance to the size of the spatial grid.
[0011] Optionally, when the vehicle pose change information indicates that the vehicle's displacement during the duration of the candidate wake-up word is not greater than a preset static displacement threshold and the heading change is not greater than a preset static heading threshold, the fitting result of the external sound source is not used as a separate basis for determining false wake-up. Based on whether the sound source direction of each phoneme evidence unit continuously corresponds to the same occupant activity area inside the vehicle, and whether the key phonemes belong to the same candidate sound source trajectory, the candidate wake-up word is determined to be a valid wake-up or a false wake-up.
[0012] Optionally, the key phonemes are determined by comparing the wake-up word phoneme template with the preset confused word phoneme template position by position, and are the wake-up word phonemes at different positions of their pronunciation; When all key phonemes are associated with the same candidate sound source trajectory, it is determined that the key phonemes continuously belong to the same candidate sound source trajectory. When a key phoneme does not form an independent sound source direction, and the preceding and following key phonemes adjacent to the key phoneme are both associated with the same candidate sound source trajectory, the key phoneme is supplemented and associated according to the predicted sound source direction of the candidate sound source trajectory at the corresponding time of the key phoneme.
[0013] Optionally, the angular deviation between the associated sound source direction of each phoneme evidence unit and the center direction of the corresponding occupant's head movement area is averaged and divided by the preset maximum allowable directional deviation to obtain the in-vehicle fitting error; the average distance from each sound source observation ray in the road reference coordinate system to the center of the candidate external sound source area is divided by the feature size of the spatial grid to obtain the external fitting error. When the key phonemes continuously belong to the same candidate sound source trajectory, and the difference between the external fitting error and the internal fitting error is greater than the preset fitting discrimination tolerance, the candidate wake-up word is determined to be a valid wake-up word. When at least two key phonemes belong to different candidate sound source trajectories, or when the difference between the in-vehicle fitting error and the out-of-vehicle fitting error is greater than the preset fitting discrimination tolerance, the candidate wake-up word is determined to be a false wake-up. When the absolute value of the difference between the in-vehicle fitting error and the out-of-vehicle fitting error is not greater than the preset fitting discrimination tolerance, the speech segment after the candidate wake-up word is obtained, and the candidate wake-up word is determined to be a valid wake-up or a false wake-up based on whether the sound source direction of the speech segment can be associated with the candidate sound source trajectory corresponding to the key phoneme.
[0014] Secondly, this application provides an in-vehicle false wake-up filtering system based on the continuity of sound source trajectories, the system comprising: The information acquisition module is used to acquire the multi-channel audio signal corresponding to the candidate wake-up word and the vehicle pose change information during the duration of the candidate wake-up word; The unit division module is used to perform phoneme timing alignment on the candidate wake words, divide them into multiple phoneme evidence units, and determine at least one candidate sound source direction for each phoneme evidence unit. The internal and external sound source fitting module is used to represent the directions of each candidate sound source in the vehicle coordinate system and the road reference coordinate system respectively according to the vehicle pose change information. Based on the matching error between each candidate sound source direction and the activity area of the occupants in the vehicle, the module determines the fitting result of the internal sound source and determines the fitting result of the external sound source based on the spatial aggregation error of the sound source observation rays corresponding to different vehicle positions. The phoneme association module is used to associate the candidate sound source direction of each phoneme evidence unit with the corresponding candidate sound source trajectory according to the time sequence of each phoneme evidence unit, and to determine whether the key phonemes in the candidate wake word continuously belong to the same candidate sound source trajectory. The wake-up control module is used to determine whether the candidate wake-up word is a valid wake-up or a false wake-up based on the sound source attribution results of the key phonemes, the fitting results of the in-vehicle sound sources, and the fitting results of the external sound sources, and to control the in-vehicle voice interaction system to output or prevent the output of wake-up commands.
[0015] By adopting the above technical solution, this application achieves the following effects: Candidate wake words are divided into phoneme evidence units with independent localization value. The system can check whether the direction of the sound source has changed segment by segment, thereby identifying spliced false wake-ups formed by different passengers speaking in sequence.
[0016] Vehicle pose is used to represent the same set of directional observations in both the vehicle coordinate system and the road reference coordinate system. Occupant sound sources moving with the vehicle maintain their correspondence with the seating area in the vehicle coordinate system, while relatively fixed sound sources beside the road form a more concentrated observation ray in the road reference coordinate system, thereby distinguishing between the two types of sound sources.
[0017] Phoneme evidence units are sequentially associated with candidate sound source trajectories, and the continuity of the trajectory attribution of key phonemes is used as the criterion to avoid making judgments based solely on a single localization result of a complete wake word.
[0018] The head movement area of the occupants inside the vehicle is adjusted according to the fore-and-aft position of the seat and the angle of the backrest, so that normal directional fluctuations caused by turning the head, looking down, or changes in sitting posture can still fall within the same seat constraint range, reducing the possibility of true wake-up being falsely filtered.
[0019] Audio frames are weighted according to the posterior probability of phonemes, and frequency points are weighted according to the stability of the phase difference or time difference of arrival between channels. Adjacent short phoneme intervals with insufficient positioning information are merged to reduce the impact of low-energy phonemes and in-vehicle noise on direction estimation.
[0020] When the directional deviations corresponding to different trajectories are close, the speaker feature similarity and the energy ratio difference of the multi-microphone channels are compared in turn to make the phoneme attribution in cases where the directions are similar or the trajectories intersect verifiable.
[0021] When the vehicle displacement and heading changes are insufficient, the external observation ray lacks an effective baseline. The system no longer assigns a separate veto power to the external fitting results, but instead uses the in-vehicle region matching and key phoneme attribution relationship to complete the determination.
[0022] When the difference between the two fitting errors is within the discrimination tolerance, continuing to use the speech segment after the wake-up word to verify whether the trajectory continues can reduce boundary misjudgment caused by short-term positioning fluctuations.
[0023] This filtering method is set after the candidate wake-up results, without changing the main structure of the original wake-up word recognition model, making it easy to connect with existing in-vehicle voice interaction systems. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0025] Figure 1 A flowchart of an in-vehicle false wake-up filtering method based on the continuity of sound source trajectory is provided in one embodiment of this application; Figure 2 This is a schematic diagram of a vehicle-mounted false wake-up filtering system based on the continuity of sound source trajectory, provided as an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only some embodiments of this application, and not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application.
[0027] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0028] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0029] Example 1
[0030] This embodiment combines Figure 1This application describes the hardware requirements, coordinate definition, time synchronization, data processing, and overall execution order of the method. This embodiment employs a six-microphone array located in the front dome light control area of the vehicle. The positions of the six microphones are recorded in the microphone array coordinate system during factory calibration. The multi-channel audio sampling rate is 16 kHz, the quantization bit depth is 16 bits, and frames are divided according to a frame length of 25 ms and a frame shift of 10 ms. The vehicle pose is provided by at least one of the vehicle CAN bus, inertial measurement unit, and positioning controller, with a pose sampling frequency of 50 Hz. The seat rail position and backrest angle are provided by the seat controller. When the vehicle does not have corresponding seat sensors, the default seat position calibrated by the vehicle model is used, and the half-axis of each corresponding occupant head movement area is expanded by 10% to absorb the uncertainty of the actual seat position.
[0031] The microphone array coordinate system M is defined with its origin at the array's geometric center, x-axis pointing forward of the vehicle, y-axis pointing to the left of the vehicle, and z-axis pointing upwards. The vehicle coordinate system B has its origin at the projection of the rear axle center onto the ground, with its three axes aligned with the microphone array coordinate system. The road reference coordinate system R uses the vehicle's location point when entering the candidate wake-word detection window as its local origin, with its x-axis pointing in the road reference direction, y-axis in the horizontal plane and perpendicular to the x-axis, and z-axis pointing upwards. The external parameters for the microphone array's installation relative to the vehicle coordinate system include the rotation matrix R. MB Translation vector t MB The installation external parameters are written into the vehicle's memory during the vehicle's factory calibration.
[0032] The audio acquisition clock is used as a unified time reference. Let τ be the center time of the t-th frame of audio. t The time of the pose sampling point is τ a and τ b , and τ a ≤τ t ≤τ b The vehicle position is then obtained using linear interpolation:
[0033] The vehicle heading angle is expanded before interpolation to ensure that the difference between adjacent heading angles is between -180° and 180° before linear interpolation. If the pose sampling interval exceeds 100 ms, or the estimated deviation between the audio timestamp and the pose timestamp exceeds 20 ms, the time period is marked as pose unavailable, and the observation rays of that time period are not used to determine the external sound source separately.
[0034] Each phoneme evidence unit U i Record at least unit number i and target phoneme p. i Start and end times and Average phoneme confidence level qi Candidate sound source direction set D i The vehicle pose P at that moment i Speaker feature vector e i and the multi-microphone channel energy ratio vector r i Candidate sound source trajectory T j At least record the trajectory number j, the set of associated phoneme evidence units, and the direction state x. j State covariance P j Crew area code s j and trajectory confidence c j When the occupant area number is empty, it indicates that the trajectory is not currently constrained by the in-vehicle seating area and can be used as a candidate trajectory outside the vehicle for fitting the road reference coordinate system.
[0035] Example 2 This embodiment illustrates the process of phoneme timing alignment and the construction of phoneme evidence units. The wake-up detection unit outputs the posterior probability of each phoneme in the target wake-up word frame by frame. For the t-th frame, the phoneme with the highest posterior probability is selected as the candidate phoneme for that frame; when the highest posterior probability is not less than 0.60, the frame is recorded as a valid phoneme frame. When there is only one discontinuous frame with a probability lower than 0.60 between two consecutive phoneme intervals of the same type, and the posterior probability of the phoneme corresponding to the discontinuous frame is still not less than 0.40, the discontinuous frame is merged into the preceding and following phoneme intervals of the same type to avoid incorrect segmentation of the same phoneme due to short-term noise.
[0036] The initial phoneme intervals are checked for independent localization capability. In this embodiment, a phoneme interval is considered to form a reliable independent sound source direction when the following conditions are met simultaneously: the phoneme interval contains no less than 4 valid audio frames; the phase difference variance between channels at no less than 24 frequency points within the 300 Hz to 4000 Hz frequency band is lower than a preset stability upper limit; and the ratio of the maximum peak value to the second largest peak value in the spatial spectrum is not less than 1.15. Phoneme intervals that do not simultaneously meet the above conditions are entered into adjacent merging processing.
[0037] Let P be the average power of the i-th phoneme interval at frequency f. i (f), the average power within the effective frequency band is The spectral continuity c between the i-th phoneme interval and the adjacent j-th phoneme interval is... i,j Defined as:
[0038] Where F is the set of effective frequency points. Calculate c respectively. i,i-1 and c i,i+1The i-th phoneme interval is merged into the side with greater spectral continuity. When the absolute value of the difference between the two intervals is not greater than 0.05, the interval with the smaller time interval is preferred; if the time intervals are still the same, the interval with the higher average posterior probability of the phoneme is preferred. Intervals located at the beginning or end of the candidate wake word are only compared with the actual adjacent intervals. The merged continuous segment is treated as a phoneme evidence unit, and the original phoneme positions it covers are preserved in this unit for subsequent determination of the trajectory assignment of key phonemes.
[0039] For example, after alignment, a candidate wake word forms five initial phoneme intervals with effective frame counts of 6, 3, 7, 5, and 2, respectively. The second phoneme interval has only 3 frames, and its spectral continuity with the first and third phoneme intervals is 0.82 and 0.61, respectively. Therefore, the second phoneme interval is merged into the first phoneme interval. The fifth phoneme interval is located at the end and has only 2 frames, so it is merged with the fourth phoneme interval. This results in three phoneme evidence units, covering the original phoneme positions 1 to 2, position 3, and positions 4 to 5, respectively. If the merged unit still cannot form a reliable direction, the phoneme time range of that unit is retained, and its direction state is recorded as missing for use in the limited supplementary association in Example 7.
[0040] Example 3 This embodiment illustrates the process of determining the weighted spatial spectrum and candidate sound source directions of a phoneme evidence unit. For the t-th frame in the i-th phoneme evidence unit, the posterior probability q of the frame belonging to the target phoneme is used. t Determine the frame weight w t :
[0041] Among them, Ω i Let f be the set of audio frames contained in the i-th phoneme evidence unit. For frequency point f, count the frequency points in Ω. i Inter-channel phase difference variance of each frame The frequency weight v is determined according to the following formula. f :
[0042] Where ε is taken as 10^-6, to avoid the denominator being zero; when If the frequency exceeds the stability limit, remove that frequency from F. For each microphone pair m and n, let X... m (t,f) and X n (t,f) are the corresponding short-time Fourier transform coefficients, and the microphone positions are a and f, respectively. m and a n Given the speed of sound as c and the unit vector of the direction to be searched as u(θ,φ), the theoretical time delay corresponding to this direction is:
[0043] The spatial spectrum of the i-th phoneme evidence unit was calculated using weighted SRP-PHAT:
[0044] This embodiment searches in 2° steps within the azimuth range of -180° to 180° and in 3° steps within the elevation range of -30° to 30°. Directions with spectral values higher than their eight neighboring regions and not lower than 65% of the unit's global maximum spectral value are identified as local peak directions; the angle between any two retained directions is not less than 12°; each phoneme evidence unit retains a maximum of 3 candidate directions. The peak confidence ρ of the l-th candidate direction... i,l The peak value is determined by dividing the difference between its peak value and the mean of its local neighborhood by the global maximum spectral value. If there is no local peak value that meets the criteria, or if the confidence level of the highest peak value is less than 0.15, the phoneme evidence unit is marked as not being able to be located independently.
[0045] To eliminate differences in microphone channel sensitivity, before calculating the spatial spectrum, background audio collected when the vehicle is stationary and there are no active sound sources inside the vehicle is used to estimate the gain of each channel, and the amplitude of each channel is normalized to the same reference channel. If a channel is detected to have an energy level more than 20 dB lower than the median energy of other channels for 200 ms, the channel is marked as abnormal, and only the remaining effective microphones are used to calculate the spatial spectrum; if there are fewer than 3 effective channels, no reliable candidate directions are output.
[0046] Example 4 This embodiment illustrates the process of establishing the head movement area of vehicle occupants and calculating the fitting error within the vehicle. For seat s, its standard head movement area is represented as a three-dimensional ellipsoid in the vehicle coordinate system:
[0047] Among them, c s As the center of the standard head activity area, Q s The region is determined by the lengths of its three half-axles and its orientation. In this embodiment, the three half-axles of the front seats can be 0.22 m, 0.18 m, and 0.25 m, respectively, and the three half-axles of the rear seats can be 0.25 m, 0.20 m, and 0.25 m, respectively. The above values are used to illustrate the region modeling method, and the specific vehicle model is calibrated based on the actual vehicle cabin dimensions.
[0048] Let the displacement of the seat along the slide rail be d. s The change in the backrest relative to the calibrated angle is β. s The unit vector of the slide rail direction is h. s The rotation matrix corresponding to the backrest axis is R. s (β sThe updated region center and shape matrix They are respectively:
[0049] Among them, c 0,s Let O be the position of the backrest pivot in the vehicle coordinate system. For the position from the microphone array center o... M Departure direction: d i,l The sound source ray, let x=o M +λ d i,l Substitute the equation into the ellipsoid equation; when the resulting quadratic equation in λ has real roots of λ≥0, determine that the candidate direction intersects with the head movement area of the occupant in seat s.
[0050] For a temporary trajectory T j Given seat s, let the trajectory have N phoneme evidence units, where N miss There are no candidate directions intersecting with seat s in any of the units; the angle between the k-th valid direction and the direction of the region center is α. k The change in the angle between adjacent effective directions is β. k Then the in-vehicle fitting error E in (j,s) is defined as:
[0051] Where, N v =NN miss θ max Take 35°, Δθ max Take 20°. When N v When N is less than 2, the second term is 0; when N is less than 2, the second term is 0. v When E is 0, in (j,s) is set to 1. E is calculated for all seating areas. in (j,s), take the minimum value as the in-vehicle fitting error E of the temporary trajectory. in (j) and use the corresponding seat as the region constraint for the trajectory. If the difference in error between two corresponding seats is not greater than 0.05, then the two seat candidates are temporarily retained, and will be further distinguished in Example 5 using speaker features and channel energy ratio.
[0052] Example 5 This embodiment illustrates the initialization, prediction, association, updating, and deletion processes of candidate sound source trajectories. The directional state of each trajectory is defined as follows:
[0053] When the time interval between adjacent phoneme evidence units is Δt, a constant angular velocity model is used for state prediction:
[0054] When the trajectory is initially established, the azimuth and elevation angles of the candidate directions are written into the status, both angular velocities are set to 0, the initial standard deviations of the azimuth and elevation angles are set to 8°, and the initial standard deviation of the angular velocity is set to 20° / s. Observation noise is adjusted according to the peak confidence of the candidate directions: the higher the peak confidence, the lower the observation noise. The prediction covariance and updated covariance are calculated according to the standard Kalman filter formula.
[0055] Let the unit vector of the prediction direction be u. p The current candidate direction unit vector is u. c The angular deviation d between the two θ for:
[0056] When d θ If the angle is no greater than 18°, the current direction is listed as an associative direction for the trajectory. If the difference between the minimum and the second smallest angle deviation is greater than 4°, the direction corresponding to the minimum angle deviation is directly selected; if the difference is no greater than 4°, the speaker feature similarity is compared.
[0057] Centered on the current phoneme evidence unit, extend 150 ms forward and backward to obtain speaker feature segments, and extract the normalized speaker embedding vector e. The speaker feature similarity S between the current unit and an already associated unit on the trajectory is calculated. spk for:
[0058] When the difference between the highest and second highest speaker feature similarity in the competing trajectories is greater than 0.05, the trajectory corresponding to the highest similarity is selected; when the difference is not greater than 0.05, or when the effective speech length is less than 120 ms and the signal-to-noise ratio is less than 6 dB and reliable speaker features cannot be extracted, the energy ratio of multiple microphone channels is further compared.
[0059] Let r be the normalized energy ratio of the i-th phoneme evidence unit in the m-th microphone channel. i,m ,but:
[0060] The channel energy ratio difference D between the current unit i and the previous unit j on the trajectory E (i,j) is:
[0061] Choose D EThe least competitive trajectory is associated. If the angular deviation between the current reliable candidate direction and all existing trajectories exceeds 18°, a new trajectory is established. A trajectory with region candidates can be established for directions intersecting any in-vehicle occupant region; directions not intersecting any in-vehicle region can still be established as unconstrained trajectories to support fitting of fixed external sound sources. A trajectory enters a dormant state when two consecutive phoneme evidence units fail to be associated, and is deleted when three consecutive units fail to be associated. If a dormant trajectory meets the association conditions again before deletion, the trajectory is restored and its original trajectory number is retained.
[0062] Example 6 This embodiment illustrates the calculation process of the sound source observation ray and the vehicle-side fitting error in the road reference coordinate system. For the k-th phoneme evidence unit, let the position of the vehicle reference point in the road reference coordinate system be p. k The rotation matrix from the vehicle coordinate system to the road reference coordinate system is R. BR,k The translation vector and rotation matrix of the microphone array relative to the vehicle coordinate system are t, respectively. MB and R MB The unit vector of the associated sound source direction in the microphone array coordinate system is d. k Then observe the starting point of the ray o k and direction u k They are respectively:
[0063] The corresponding sound source observation ray is:
[0064] A three-dimensional search space is established within a range of 15 m in front and behind the vehicle's trajectory, 20 m to the left and right, and 0.5 m to 3.0 m above the ground. The space is then divided into grids with sides of 0.5 m. For the grid center g... j First, calculate its projection parameters along the direction of the k-th ray:
[0065] When λ k,j When ≥0, the shortest distance from the grid center to the ray is:
[0066] When λ k,j When d < 0, it indicates that the grid center is located in the opposite direction of the ray and is not considered as a candidate sound source location supported by that ray. k,j Let this be the diagonal length of the search space. For a temporary trajectory with N valid observation rays, calculate the average clustering distance of each grid, and select the grid with the smallest average distance as the candidate external sound source region:
[0067] Vehicle external fitting error E out Defined as:
[0068] Where h is the grid side length, and in this embodiment, h = 0.5 m. If N is not less than 5, the observation with the largest distance can be removed first, and then the remaining observations can be averaged to reduce the impact of a single incorrect localization on the vehicle-side fitting results.
[0069] The vehicle exterior fitting results are only included in the final comparison if they meet the requirements of a valid geometric baseline. This embodiment specifies that: the number of valid observation rays N is no less than 3; the vehicle position baseline B during the candidate wake-up word period is no less than 0.40 m, or the heading change range is no less than 3°; and the standard deviation of the positioning position is no greater than 0.30 m. When all three conditions are met simultaneously, the valid vehicle exterior fitting result F is marked. out Set to 1; otherwise set to 0. When F out When =0, E out It is only retained as an auxiliary record and will not trigger a false wake-up detection separately.
[0070] Example 7 This embodiment illustrates the determination of key phonemes, the determination of trajectory continuity, and the limited supplementation process for a single missing key phoneme. The wake-up word phoneme template and each preset confusion word phoneme template are aligned sequentially using dynamic programming. The alignment cost for identical phonemes is 0, while the cost for replacement, insertion, and deletion is 1. The replacement, insertion, and deletion positions are obtained by backtracking from the path with the minimum cumulative cost. The replacement position and the wake-up word phonemes adjacent to the insertion or deletion position are identified as the differing phoneme positions. After aligning multiple confusion words separately, all differing positions are merged to obtain the key phoneme set.
[0071] When the set of key phonemes contains only one location, in order to avoid judging continuity based on a single direction, another phoneme that can be reliably located and has the largest time interval with the existing key phonemes is selected from the candidate wake-up words as the trajectory anchor point; when there is no second reliable location phoneme, the candidate wake-up word is transferred to the boundary verification branch of Example 8, and the effective wake-up is not directly output based on the continuity of key phonemes.
[0072] The association results of key phonemes and trajectory anchors are traversed. When all key phonemes and trajectory anchors are associated with the same trajectory, and there are no key phonemes belonging to other trajectories in their temporal order, it is determined that the key phonemes continuously belong to the same candidate sound source trajectory. When at least two key phonemes are clearly associated with different trajectories, it is directly determined that a sound source switch exists.
[0073] Only supplementary associations are allowed for a single key phoneme that has not yet formed an independent direction. Assume that both the key phonemes before and after this key phoneme are associated with trajectory T.j Using the state prediction in Example 5, the prediction direction and prediction covariance at the missing time are obtained; when the prediction direction is still consistent with T j When the occupant region constraints intersect and the standard deviations of the predicted azimuth and pitch angles are both no greater than 8°, the key phoneme is supplementally associated with T. j Missing key phonemes located at the beginning or end, two or more consecutive missing key phonemes, and cases where the prediction uncertainty exceeds the above upper limit will not be supplemented with association.
[0074] Example 8 This embodiment illustrates the final judgment rules, the degradation processing for insufficient pose, and the boundary verification process. The difference between the vehicle's internal and external fitting values is defined as:
[0075] In this embodiment, the fitting discrimination tolerance δ is set to 0.20. This value is determined by collecting real wake-up samples in the vehicle, fixed sound source samples on the road, and multi-person spliced samples in the vehicle, and statistically analyzing the distribution of ΔE in each type of sample. For other vehicle models, δ can be recalibrated without changing the judgment logic.
[0076] When at least two key phonemes belong to different trajectories, E is no longer compared. in and E out This is directly identified as a false wake-up. When key phonemes continuously belong to the same trajectory, F... out When F = 1 and ΔE > δ, it indicates that the fit of the in-vehicle region is better than the fit of the fixed road sound source, and it is determined to be an effective wake-up; when F out When ΔE = 1 and ΔE < -δ, it indicates that the fitting of the fixed sound source on the road is better than the fitting of the in-vehicle area, and it is determined to be a false wake-up.
[0077] When F out When the value is 0, check whether the key phonemes continuously belong to the same trajectory, and whether at least 80% of the effective directions in that trajectory intersect with the same occupant's head movement area. If both conditions are met, it is determined to be a valid wake-up; if either condition is not met, it is determined to be a false wake-up. Therefore, when the vehicle is stationary, moving at low speed, or pose data is temporarily unavailable, the road ray aggregation results lacking a valid baseline are not used as a sole criterion for rejection.
[0078] When F outWhen ΔE = 1 and |ΔE| ≤ δ, 500 ms of audio is read from the buffer after the end of the candidate wake-up word. If the voice activity detection yields at least 120 ms of subsequent voice, the direction of the subsequent voice is extracted according to the methods of Examples 3 to 5, and an attempt is made to continue the original trajectory. If the subsequent voice can be associated with the trajectory where the key phoneme is located, it is determined to be a valid wake-up; if it is clearly associated with other trajectories, it is determined to be a false wake-up. If the subsequent voice does not exist or cannot form a reliable direction, the confidence of the original wake-up detection and the continuity within the vehicle are checked: if the confidence of the original wake-up detection is not less than 0.85, the key phoneme is on the same track, and at least 80% of the valid directions intersect with the same occupant area, it is determined to be a valid wake-up; otherwise, it is determined to be a false wake-up. The above branch avoids mechanically judging a user as a false wake-up when the user only says the wake-up word and does not continue speaking.
[0079] Example 9 This embodiment provides four complete numerical processing procedures to illustrate how the aforementioned formulas and judgment conditions are combined and executed. The numerical values are only used to illustrate the calculation steps and do not limit the scope of protection of this application.
[0080] The first group consists of actual wake-ups issued by the primary driver. Candidate wake-up words were divided into five phoneme evidence units, U1 to U5, representing times of 0.10 s, 0.22 s, 0.35 s, 0.49 s, and 0.63 s, respectively. During this period, the vehicle moved from the road reference coordinates (0,0,0) to (0.72,0.02,0), with a heading change of 0.8°, satisfying the valid conditions for external fitting. The azimuth angles after formal association were -29°, -27°, -30°, -26°, and -28°, and the pitch angles were 8°, 9°, 8°, 7°, and 8°, all associated with the primary driver trajectory T1. The angular deviations corresponding to the center direction of the primary driver's area were 3°, 2°, 4°, 3°, and 2°, with an average angular deviation of 2.8°; the changes in adjacent directions were 1°, 2°, 2°, and 1°, with an average change of 0.5°; the number of missing observations was 0. Therefore:
[0081] The distances from the five observation rays in the road reference coordinate system to the center of the optimal grid are 0.72 m, 0.81 m, 0.77 m, 0.86 m, and 0.89 m, respectively, with an average distance of 0.81 m. Therefore:
[0082] Therefore, ΔE = 1.62 - 0.0665 = 1.5535, which is greater than δ = 0.20, and since all key phonemes are associated with T1, a valid wake-up command is output.
[0083] The second group consists of a fixed broadcast sound source on the right side of the road as the vehicle passes. The candidate wake-up words also form five phoneme evidence units, with their vehicle coordinate system azimuth angles changing sequentially from 42° to 28°, 12°, -5°, and -21°, exhibiting a sweeping characteristic from the right front to the right rear. The angular deviations between each direction and the direction of the nearest passenger-side center are 12°, 18°, 24°, 29°, and 34°, respectively, with an average angular deviation of 23.4°; the average change in adjacent directions is 15.75°; and the number of missing observations is 0. Therefore:
[0084] The vehicle moved 1.05 m during the candidate wake-up word. After switching to the road reference coordinate system, the five observation rays converged on the same grid approximately 8.5 m to the right of the road. The average distance of the rays to the center of this grid was 0.16 m. Therefore:
[0085] Therefore, ΔE = 0.32 - 0.6039 = -0.2839, which is less than -δ. Thus, it is determined that the false wake-up is caused by a fixed sound source on the road, and the output wake-up command is prevented.
[0086] The third group consists of candidate wake-up words formed by splicing together the wake-up words of two occupants inside the vehicle. The candidate directions of U1 and U2 intersect with the driver's area and are associated with trajectory T1, while the candidate directions of U3 to U5 intersect with the passenger's area and are associated with trajectory T2. The key phonemes corresponding to U2 and U4 are determined by aligning the wake-up words with the confusion word template. Since U2 is associated with T1 and U4 is associated with T2, at least two key phonemes clearly belong to different trajectories. Therefore, the fitting error between the vehicle interior and exterior is no longer compared, and it is directly determined to be a spliced false wake-up.
[0087] The fourth group consists of real wake-up calls issued by the driver when the vehicle is stationary. During the candidate wake-up call period, the vehicle displacement was 0.06m and the heading change was 0.4°, which does not meet the valid conditions for external fitting. Therefore, F... out =0. Five of the six valid directions intersect with the head movement area of the driver / occupant, accounting for 83.3%; all key phonemes are associated with the driver's trajectory T1. Since the key phonemes are on the same track and the intersection rate in the same area is higher than 80%, a valid wake-up command is output without using the external fitting results for separate determination.
[0088] Example 10 Please see Figure 2The vehicle-mounted false wake-up filtering system 300 includes an information acquisition module 301, a unit division module 302, an internal and external sound source fitting module 303, a phoneme association module 304, and a wake-up control module 305. The system can be implemented by a software program executed by the vehicle's main processor, or by a digital signal processor performing multi-channel audio calculations and the main processor performing pose transformation and judgment control. The memory stores microphone array calibration parameters, external parameters of the microphone array installation to the vehicle coordinate system, head movement area parameters for each seat, wake-up word phoneme templates, confusion word phoneme templates, and various judgment thresholds.
[0089] The information acquisition module 301 establishes a unified timestamp cache and writes multi-channel audio frames, vehicle position, vehicle posture, seat position, and backrest angle to them respectively. When vehicle posture data is late, the information acquisition module waits for a maximum of 100 ms; after the waiting time, the corresponding posture is marked as unavailable, but the audio processing flow is still maintained. The unit division module 302 sequentially outputs phoneme evidence unit records U. i The system also saves the candidate direction set, peak confidence, and direction missing flag in the records.
[0090] The phoneme association module 304 maintains a candidate trajectory list. Upon receiving a new phoneme evidence unit, it first performs state prediction and angle threshold filtering, then resolves competing associations based on speaker characteristics and channel energy ratios; after association is complete, it updates the trajectory state, candidate trajectory regions, and trajectory confidence. The internal and external sound source fitting module 303 outputs the in-vehicle fitting error E for each temporary trajectory. in Vehicle external fitting error E out Optimal occupant region number, candidate external sound source grid, and valid external fitting flag F out Therefore, what is transmitted between modules are specific data records, rather than just result descriptions such as "high degree of matching" or "continuous trajectory".
[0091] The wake-up control module 305 reads the key phoneme set, the trajectory number of each key phoneme, and E. in E out and F out The system executes the judgment order according to Example 8. If a wake-up is determined to be valid, an output permission flag is sent to the original in-vehicle voice interaction system; if a wake-up is determined to be false, an output blocking flag is sent and the temporary trajectory of this candidate wake-up is cleared. For candidate wake-ups entering the boundary review branch, the system retains the original trajectory for 500 ms, and releases the buffer after the review is completed.
[0092] When there are fewer than three effective microphone channels, all key phonemes cannot be located, the audio and pose time deviation exceeds the synchronization tolerance, and continuous direction within the same occupant area cannot be obtained, the system marks the current result as a low-confidence result. For low-confidence results, instead of using road-fixed sound source fitting to draw a positive conclusion, the system performs the downgrade judgment in Example 8 based on the original wake-up detection confidence and in-vehicle direction continuity. Through the above module interfaces, data structures, calculation formulas, threshold conditions, and abnormal branches, those skilled in the art can reproduce the vehicle-mounted false wake-up filtering process of this application according to the same processing link.
[0093] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; without departing from the spirit and scope of the technical solutions of this application, the above modifications or substitutions still fall within the protection scope of this application.
Claims
1. A vehicle-mounted false wake-up filtering method based on the continuity of sound source trajectory, characterized in that, include: S1. Obtain the multi-channel audio signal corresponding to the candidate wake-up word and the vehicle pose change information during the duration of the candidate wake-up word; S2. Perform phoneme timing alignment on the candidate wake words to divide them into multiple phoneme evidence units, and determine at least one candidate sound source direction for each phoneme evidence unit. S3. Based on the vehicle pose change information, the directions of each candidate sound source are represented in the vehicle coordinate system and the road reference coordinate system respectively. The fitting result of the in-vehicle sound source is determined based on the matching error between each candidate sound source direction and the activity area of the occupants in the vehicle. The fitting result of the external sound source is determined based on the spatial aggregation error of the sound source observation rays corresponding to different vehicle positions. S4. According to the timing of each phoneme evidence unit, associate the candidate sound source direction of each phoneme evidence unit with the corresponding candidate sound source trajectory, and determine whether the key phonemes in the candidate wake-up word continuously belong to the same candidate sound source trajectory. S5. Based on the sound source attribution results of the key phonemes, the fitting results of the in-vehicle sound sources and the fitting results of the external sound sources, determine whether the candidate wake-up word is a valid wake-up or a false wake-up, and control the in-vehicle voice interaction system to output or prevent the output of the wake-up command.
2. The method according to claim 1, characterized in that, The candidate wake words are divided into multiple phoneme evidence units, including: Based on the posterior probability of the phonemes corresponding to each audio frame, audio frames that consecutively correspond to the same target phoneme are grouped into phoneme intervals. For phoneme intervals that cannot form a reliable independent sound source direction, the spectral continuity between the phoneme interval and the preceding and following phoneme intervals is calculated respectively, and the phoneme interval is merged with the adjacent phoneme intervals with higher spectral continuity to obtain the phoneme evidence unit.
3. The method according to claim 2, characterized in that, Determine the sound source direction of each phoneme evidence unit, including: Based on the phoneme posterior probability of each audio frame in the phoneme evidence unit, each audio frame is weighted. Each frequency point is weighted according to the stability of the phase difference or time of arrival difference between different microphone channels in consecutive audio frames. A spatial acoustic spectrum is generated using weighted audio frames and frequency points, and at least one direction in the spatial acoustic spectrum that satisfies the local peak condition is determined as the candidate sound source direction of the phoneme evidence unit.
4. The method according to claim 3, characterized in that, The occupant activity area includes occupant head movement areas corresponding to each seat in the vehicle, and the occupant head movement areas are adjusted according to the fore-and-aft position and backrest angle of the corresponding seats. For each occupant's head movement area, a direction intersecting with that occupant's head movement area is selected from the candidate sound source directions of each phoneme evidence unit. The average angular deviation between the selected direction and the center direction of the occupant's head movement area, as well as the average angular change between adjacent selected directions, are calculated. The sum of these two is taken as the region matching error. The minimum value of the region matching error corresponding to each occupant's head movement area is taken as the in-vehicle sound source fitting result, and the corresponding occupant's head movement area is taken as the region constraint of the candidate sound source trajectory.
5. The method according to claim 4, characterized in that, Associating each phoneme evidence unit with its corresponding candidate sound source trajectory, including: Based on the sound source direction and corresponding time of the associated phoneme evidence unit in the candidate sound source trajectory, predict the sound source direction of the candidate sound source trajectory at the corresponding time of the current phoneme evidence unit. Among the candidate sound source directions that intersect with the regional constraints of the candidate sound source trajectories, the sound source directions to be associated are determined in ascending order of angular deviation from the predicted sound source directions. When the difference between the minimum angle deviation and the second smallest angle deviation is not greater than the preset angle discrimination tolerance, the speaker feature similarity between the current phoneme evidence unit and the previous associated phoneme evidence unit in each candidate sound source trajectory is sorted from largest to smallest; when the difference between the two speaker feature similarities in the first sorting is not greater than the preset similarity discrimination tolerance, the candidate sound source trajectory corresponding to the current phoneme evidence unit is determined from smallest to largest according to the difference in the energy ratio of the multi-microphone channels.
6. The method according to claim 5, characterized in that, Determining the fitting result of the external sound source includes: Based on the vehicle position and attitude at the corresponding time of each phoneme evidence unit, the corresponding sound source direction is converted into a sound source observation ray in the road reference coordinate system; The road reference space around the vehicle is divided into spatial grids, and the average distance from the observation ray of each sound source to the center of each spatial grid is calculated. The spatial grid with the smallest average distance is identified as the candidate external sound source region, and the fitting result of the external sound source is determined based on the ratio of the average distance to the size of the spatial grid.
7. The method according to claim 6, characterized in that, When the vehicle pose change information indicates that the displacement of the vehicle during the duration of the candidate wake-up word is not greater than a preset static displacement threshold and the change in heading is not greater than a preset static heading threshold, the fitting result of the external sound source is not used as a separate basis for determining false wake-up. Based on whether the sound source direction of each phoneme evidence unit continuously corresponds to the same occupant activity area inside the vehicle, and whether the key phonemes belong to the same candidate sound source trajectory, the candidate wake-up word is determined to be a valid wake-up or a false wake-up.
8. The method according to claim 6, characterized in that, The key phonemes are determined by comparing the wake word phoneme template with the preset confused word phoneme template position by position, and are the wake word phonemes at different positions of their pronunciation; When all key phonemes are associated with the same candidate sound source trajectory, it is determined that the key phonemes continuously belong to the same candidate sound source trajectory. When a key phoneme does not form an independent sound source direction, and the preceding and following key phonemes adjacent to the key phoneme are both associated with the same candidate sound source trajectory, the key phoneme is supplemented and associated according to the predicted sound source direction of the candidate sound source trajectory at the corresponding time of the key phoneme.
9. The method according to claim 8, characterized in that, The in-vehicle fitting error is obtained by averaging the angular deviation between the associated sound source direction of each phoneme evidence unit and the center direction of the corresponding occupant's head movement area, and dividing it by the preset maximum allowable directional deviation. The in-vehicle fitting error is obtained by dividing the average distance from each sound source observation ray in the road reference coordinate system to the center of the candidate external sound source area by the feature size of the spatial grid. When the key phonemes continuously belong to the same candidate sound source trajectory, and the difference between the external fitting error and the internal fitting error is greater than the preset fitting discrimination tolerance, the candidate wake-up word is determined to be a valid wake-up word. When at least two key phonemes belong to different candidate sound source trajectories, or when the difference between the in-vehicle fitting error and the out-of-vehicle fitting error is greater than the preset fitting discrimination tolerance, the candidate wake-up word is determined to be a false wake-up. When the absolute value of the difference between the in-vehicle fitting error and the out-of-vehicle fitting error is not greater than the preset fitting discrimination tolerance, the speech segment after the candidate wake-up word is obtained, and the candidate wake-up word is determined to be a valid wake-up or a false wake-up based on whether the sound source direction of the speech segment can be associated with the candidate sound source trajectory corresponding to the key phoneme.
10. A vehicle-mounted false wake-up filtering system based on the continuity of sound source trajectory, characterized in that, include: The information acquisition module is used to acquire the multi-channel audio signal corresponding to the candidate wake-up word and the vehicle pose change information during the duration of the candidate wake-up word; The unit division module is used to perform phoneme timing alignment on the candidate wake words, divide them into multiple phoneme evidence units, and determine at least one candidate sound source direction for each phoneme evidence unit. The internal and external sound source fitting module is used to represent the directions of each candidate sound source in the vehicle coordinate system and the road reference coordinate system respectively according to the vehicle pose change information. Based on the matching error between each candidate sound source direction and the activity area of the occupants in the vehicle, the module determines the fitting result of the internal sound source and determines the fitting result of the external sound source based on the spatial aggregation error of the sound source observation rays corresponding to different vehicle positions. The phoneme association module is used to associate the candidate sound source direction of each phoneme evidence unit with the corresponding candidate sound source trajectory according to the time sequence of each phoneme evidence unit, and to determine whether the key phonemes in the candidate wake word continuously belong to the same candidate sound source trajectory. The wake-up control module is used to determine whether the candidate wake-up word is a valid wake-up or a false wake-up based on the sound source attribution results of the key phonemes, the fitting results of the in-vehicle sound sources, and the fitting results of the external sound sources, and to control the in-vehicle voice interaction system to output or prevent the output of wake-up commands.