A multi-occupant concurrent instruction-oriented vehicle-mounted sound source positioning conflict resolution method
Patent Information
- Application Number
- CN202611166842.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-08-03
AI Technical Summary
本申请将逐帧声源位置扩展为连续声源轨迹,并结合声纹相似度、音素序列连续性和历史归属概率确定候选语音流的发令乘员,避免仅以当前声源所在音区确定乘员归属,能够减少乘员转身、探身或座椅位置变化造成的非真实归属切换。
Smart Images

Figure CN122676829B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech signal processing technology, and in particular to a method for resolving conflicts in vehicle-mounted sound source localization for concurrent commands from multiple occupants. Background Technology
[0002] With the development of in-vehicle voice interaction systems, occupants in different seats can control functions such as windows, air conditioning, audio, navigation, and defogging via voice. When multiple occupants speak simultaneously, existing technologies typically first separate overlapping speech, then determine the occupant issuing the command based on the sound source's seat frequency range, and process conflicting voice commands according to occupant permissions or preset priorities.
[0003] However, the confined space inside a vehicle and the complex acoustic reflections mean that occupants may turn around, lean forward, or change their seat position while speaking, causing the sound source of the same occupant to cross the preset sound zone boundary, resulting in a continuous speech being incorrectly assigned to different occupants. When adjacent occupants lean forward simultaneously and their sound source trajectories approach each other, the speech separation output channel may also experience speaker identity swapping, causing the recognized text to be incorrectly associated with the occupant who gave the command.
[0004] Furthermore, existing methods typically treat adjacent recognized statements as independent commands, making it difficult to distinguish between cancellation, parameter overriding, or target replacement commands from the same occupant and competing commands from different occupants. Simultaneously, some voice commands, although targeting different objects, may collectively affect vehicle functional nodes such as compressors and blowers; comparing only the target object names can easily overlook actual control conflicts. Using fixed occupant priorities to retain or discard commands as a whole results in the abandonment of non-conflicting local requirements. Therefore, a vehicle-mounted multi-occupant concurrent voice command resolution method is needed that can simultaneously address continuous sound source attribution, channel switching suppression, semantic correction recognition, and functional conflict handling. Summary of the Invention
[0005] This application provides a method for resolving conflicts in vehicle-mounted sound source localization for concurrent commands from multiple occupants, in order to solve the above-mentioned problems. The method includes: S1. Acquire multi-channel speech signals collected by the in-vehicle microphone array, separate overlapping speech intervals to obtain multiple candidate speech streams, and generate corresponding sound source trajectories based on the arrival time difference between the channels of each candidate speech stream. S2. Extract the voiceprint features and phoneme sequence features of each candidate speech stream. Based on the spatial continuity, voiceprint similarity and phoneme sequence continuity of the sound source trajectory, determine the attribution relationship between each candidate speech stream and the occupants in the vehicle, and maintain the corresponding attribution relationship when the sound source trajectory crosses the boundary of a preset sound zone. S3. When the distance between at least two sound source trajectories is less than the preset distance, based on the continuity of sound source positions and voiceprint similarity within the preset time window, compare the cumulative matching results when maintaining the original attribution relationship and when exchanging the attribution relationship, and penalize the switching of the attribution relationship to determine the passenger attribution relationship of each candidate speech stream after the sound source trajectories are close. S4. After determining the affiliation of the crew members, perform speech recognition and semantic parsing on each candidate speech stream to obtain standardized speech commands that include the commanding crew member, target object, target area, and control action. S5. Based on whether the issuing crew members of the standardized voice commands that are adjacent in time are the same, whether the subsequent voice command inherits the target object of the preceding voice command, and whether it contains a correction expression, identify the correction relationship of the same crew member to the preceding voice command, and update the corresponding standardized voice command accordingly. S6. For multiple standardized voice commands that do not belong to the aforementioned correction relationship, determine the vehicle functional nodes that are directly controlled and indirectly affected by each standardized voice command based on the role dependency relationship between vehicle functional nodes. S7. Based on the overlapping relationship between the vehicle function nodes corresponding to different standardized voice commands and the opposing relationship of control actions on the overlapping vehicle function nodes, determine the command conflict, perform semantic reconstruction on the standardized voice commands with conflict, and output the semantic command after conflict resolution.
[0006] Optionally, the speech separation of overlapping speech regions includes: The multi-channel speech signal is framed and time-frequency transformed, and a spatial covariance matrix is constructed based on the multi-channel time-frequency signal of each speech frame; Based on the eigenvalue distribution of the spatial covariance matrix, determine the proportion of the sum of all eigenvalues except the largest eigenvalue in the sum of all eigenvalues; When the percentage is continuously greater than the preset overlap threshold, the corresponding speech frame is determined as the overlapping speech interval, and the number of active sound sources is determined according to the spatial covariance matrix. The multiple candidate speech streams are then separated according to the number of active sound sources.
[0007] Optionally, generating the corresponding sound source trajectory includes: Calculate the generalized cross-correlation phase transformation results of each candidate speech stream across different microphone channels, and determine the arrival time difference between channels based on the peak position of the generalized cross-correlation phase transformation results; By combining the positions of each microphone in the microphone array and the arrival time difference between the channels, the sound source positions of each candidate speech stream at different times are determined; The changes between the sound source position and the sound source position at adjacent times are used as position state and velocity state, respectively. The sound source state of each candidate speech stream is recursively calculated, and the recursed sound source positions are connected in chronological order to form the sound source trajectory.
[0008] Optionally, determining the attribution relationship between each candidate voice stream and the vehicle occupant includes: For the q-th candidate speech stream and the i-th passenger, the attribution matching degree is determined according to the following formula: ; in, To ensure spatial continuity based on the current sound source location and the predicted sound source location of the i-th occupant, This represents the voiceprint similarity between the current voiceprint feature and the historical voiceprint feature of the i-th vehicle occupant. To determine the phoneme continuity between the current phoneme sequence features and the historical phoneme sequence features of the i-th occupant, Let w be the crew affiliation probability at the previous time step. p w e w c w h These are the weights of the corresponding parameters, and the sum of the weights is 1; The passenger affiliation probability of the candidate voice stream is updated according to the affiliation matching degree of each passenger in the vehicle, and the passenger in the vehicle with the highest passenger affiliation probability is determined as the starting passenger of the corresponding candidate voice stream.
[0009] Optionally, when the distance between at least two sound source trajectories is less than a preset distance, the target correspondence between each candidate speech stream and the vehicle occupant is determined according to the following formula: ; In the formula, π represents the candidate correspondence between candidate speech streams and in-vehicle occupants, H represents the number of speech frames included in the preset time window, and A q,π(q) (nh) represents the attribution matching degree between the q-th candidate speech stream and the vehicle occupant in the corresponding speech frame, where N is the attribution matching degree between the q-th candidate speech stream and the vehicle occupant. SW (π) represents the number of attribution relationship changes caused by the candidate correspondence within a preset time window, λ sw To switch the penalty coefficient; The candidate correspondence that maximizes the difference between the cumulative attribution matching degree and the attribution relationship switching penalty is determined as the target correspondence.
[0010] Optionally, the standardized voice commands may also include action parameters and command confidence levels; The confidence level of the instruction is determined jointly based on the speech recognition confidence level, occupant attribution confidence level, semantic integrity, and overlapping speech separation quality of the candidate speech stream; When the confidence level of the instruction is lower than the preset instruction threshold, the corresponding standardized voice instruction is marked as an instruction to be confirmed, and the instruction to be confirmed is not directly used to output the conflict resolution result with other standardized voice instructions.
[0011] Optionally, identifying the correction relationship of the same passenger to preceding voice commands includes: Within a preset correction time window, select preceding and following voice commands that are adjacent in time, and determine the probability of the same commanding crew member based on the crew member affiliation probability distribution of the two commands. Determine whether the subsequent voice instruction omits and inherits the target object of the preceding voice instruction, or whether it contains an undo expression, parameter replacement expression, target object replacement expression, or area of effect adjustment expression; When the probability of the same commanding crew member is greater than the preset threshold for the same person, and the subsequent voice command satisfies the target object inheritance condition or includes the cancellation expression, parameter replacement expression, target object replacement expression or area of action adjustment expression, it is determined that the subsequent voice command and the preceding voice command constitute the correction relationship. Based on the correction relationship, the preceding voice command is deleted, the original action parameters are overwritten with new action parameters, the target object is replaced, or the target area is adjusted.
[0012] Optionally, determining the vehicle function nodes directly controlled and indirectly affected by each standardized voice command includes: Construct a vehicle function node dependency matrix H, where the elements in the vehicle function node dependency matrix are used to represent the degree of influence of the state change of one vehicle function node on another vehicle function node. Generate a direct target vector u based on the target object and target region of the j-th standardized speech instruction. j The corresponding control scope vector is determined according to the following formula: ; Among them, s j Let R be the control scope vector, R be the maximum propagation order of the action dependency, μ be the propagation attenuation coefficient, and clip(·,0,1) represent restricting each element in the vector to between 0 and 1. Vehicle function nodes in the control domain vector that are greater than a preset threshold are identified as vehicle function nodes that are directly controlled or indirectly affected by the corresponding standardized voice commands.
[0013] Optionally, determining instruction conflicts includes: Based on the control domain vectors corresponding to different standardized voice commands, the overlapping vehicle function nodes between two standardized voice commands are determined. When two standardized voice commands have opposite control actions on the same overlapping vehicle function node, it is determined that the two constitute a direct reverse conflict. When the set of vehicle function nodes corresponding to one standardized voice command includes the set of vehicle function nodes corresponding to another standardized voice command, and the control actions of the two on the included vehicle function nodes are opposite, it is determined that the two constitute a conflict in the scope of action. When two standardized voice commands have different direct target objects, but have the same indirect influence on the vehicle function node after being propagated through the vehicle function node action dependency matrix, and the control actions on the indirectly influenced vehicle function node are in opposite directions, it is determined that the two constitute an action dependency conflict.
[0014] Optionally, the vehicle function nodes corresponding to different standardized voice commands can be divided into non-overlapping vehicle function nodes and overlapping vehicle function nodes, and local semantic commands that can be output in parallel can be generated for the non-overlapping vehicle function nodes. If the scope of action contains conflicts, remove vehicle function nodes that conflict with standardized voice commands with smaller scopes from the standardized voice commands with larger scopes, and retain the local semantic commands corresponding to the remaining vehicle function nodes. For the direct reverse conflict or action-dependent conflict, a parameter-restricted semantic command or a semantic command to be confirmed is generated based on the occupant affiliation confidence and command confidence of the corresponding standardized voice command. The corresponding commanding crew member identifier is bound to the semantic instruction to be confirmed, so that the subsequent confirmed voice is used to update the semantic instruction to be confirmed only when its crew member affiliation matches the commanding crew member identifier.
[0015] Through the above technical solution, this application achieves the following beneficial effects: This application expands the frame-by-frame sound source location into a continuous sound source trajectory, and combines voiceprint similarity, phoneme sequence continuity and historical attribution probability to determine the commanding occupant of the candidate speech stream, avoiding the determination of occupant affiliation based solely on the current sound source's vocal range, and can reduce non-realistic affiliation switching caused by occupant turning around, leaning forward or changes in seat position.
[0016] When the sound source trajectories of different passengers are close to each other, this application compares the cumulative matching results when the original affiliation relationship is maintained and when the affiliation relationship is exchanged within a preset time window, and applies a penalty to the affiliation relationship switching. This can suppress speaker channel switching caused by short-term positioning fluctuations and overlapping speech separation errors, and improve the stability of the binding between the speech stream and the passenger identifier.
[0017] This application combines the consistency of the commanding crew, the inheritance relationship of the target object, and the modified expression to identify the cancellation, parameter overwriting, target object replacement, or adjustment of the area of action of the preceding speech command by the same crew member. It can reconstruct spoken corrections into final valid commands, reducing false conflicts and repetitive control.
[0018] Based on the role dependency relationship between vehicle functional nodes, this application extends the direct target object of voice commands to the vehicle functional nodes that are indirectly affected by them. It can identify role dependency conflicts where the direct target objects are different but actually share the same compressor, blower or other execution nodes, thereby reducing the missed detection of indirect conflicts.
[0019] This application does not simply retain or discard conflicting instructions based on fixed passenger priority, but rather performs localized semantic reconstruction of voice instructions based on conflicting and non-conflicting nodes. This avoids the shared vehicle's functional nodes being subject to opposite control while preserving the local voice needs that can be met for different passengers.
[0020] This application establishes a continuous processing chain from multi-channel overlapping speech detection and separation, sound source localization and occupant attribution, speech recognition and correction relationship judgment, to vehicle functional domain analysis and conflict semantic reconstruction, which can improve the completeness of instruction attribution, the accuracy of conflict recognition, and the retention rate of effective instructions in concurrent voice interaction among multiple occupants in vehicles. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a vehicle-mounted sound source localization conflict resolution method for concurrent commands from multiple occupants, provided as an embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0024] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0025] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0026] Example 1: Voice command attribution retention during occupant movement across voice zones
[0027] This embodiment provides a vehicle-mounted sound source localization conflict resolution method for concurrent commands from multiple occupants. It is used to solve the problem that when occupants turn around, lean forward, or change their seat position while issuing voice commands, the sound source trajectory of the same voice command crosses the boundary of a preset sound zone and is thus incorrectly attributed to different occupants.
[0028] 1. In-vehicle voice acquisition device and parameter settings
[0029] This embodiment uses a five-seat passenger vehicle as the application object, and sets up 8 microphones inside the vehicle. The 8 microphones are respectively set on the left side of the front roof, the right side of the front roof, the left side of the rear roof, the right side of the rear roof, the left side of the dashboard, the right side of the dashboard, the driver's headrest, and the passenger's headrest.
[0030] A three-dimensional coordinate system is established inside the vehicle, with the vehicle's horizontal direction as the x-axis, the vehicle's vertical direction as the y-axis, and the vehicle's height as the z-axis. The vehicle's forward direction is the positive y-axis, and the rightward direction is the positive x-axis. Example coordinates for each microphone are as follows: Table 1 Microphone Coordinate Data
[0031] Each microphone synchronously acquires in-vehicle audio signals at a sampling rate of 16kHz and a quantization bit depth of 16bit. The acquired multi-channel audio signals are processed by frame segmentation, with a frame length of 32ms, a frame shift of 16ms, and windowing using a Hamming window.
[0032] Based on the vehicle's seat positions, the interior space is initially divided into a driver's seat area, a front passenger seat area, a left rear seat area, a middle rear seat area, and a right rear seat area. These preset areas are only used to initialize the attribution relationship between candidate voice streams and occupants. During the duration of a voice command, the issuing occupant is not switched directly based on whether the sound source crosses the area boundary.
[0033] After the vehicle starts, the historical voiceprint features of the driver and front passenger are established using non-overlapping speech segments from each passenger's individual speech. If a passenger has not pre-registered their voiceprint, a temporary passenger identifier is first established based on the sound region of their initial sound source, and the corresponding historical voiceprint features are updated using subsequent stable speech segments.
[0034] 2. Setting up concurrent voice communication scenarios for multiple occupants
[0035] In this embodiment, the driver is located in the driver's seat, and the front passenger is located in the front passenger seat. The driver and front passenger simultaneously issue the following commands during certain time periods: The driver gave the voice command: "Navigate to the company." The passenger in the front seat gave a voice command: "Open my window." When issuing the above voice command, the front passenger turns to the right rear to communicate with the rear passenger, and the position of the sound source of their head gradually moves from the front passenger's voice area to the rear right voice area.
[0036] The locations of the sound sources for the front passenger at several representative moments are as follows: Table 2 Location of sound source
[0037] Under the sound zone boundary set in this embodiment, the sound source position of the front passenger occupant after 1.2 seconds has crossed the preset boundary between the front passenger sound zone and the right rear sound zone.
[0038] 3. Overlapping speech detection and speech separation
[0039] Short-time Fourier transform is performed on the multi-channel speech signals collected by eight microphones to obtain the multi-channel time-frequency signals of each speech frame, and a spatial covariance matrix is constructed based on the multi-channel time-frequency signals.
[0040] The spatial covariance matrix is decomposed into eigenvalues, and the eigenvalues are arranged in descending order. The spatial dispersion of the sound source in the corresponding speech frame is determined based on the proportion of the sum of the remaining eigenvalues (excluding the largest eigenvalue) to the sum of all eigenvalues.
[0041] When the spatial dispersion of the sound source is greater than 0.26 for five consecutive speech frames, the corresponding time period is determined as the overlapping speech interval. In this embodiment, the driver's voice and the front passenger's voice overlapped within approximately 0.4 to 1.5 seconds.
[0042] For the detected overlapping speech intervals, a multi-channel speech separation method based on time-frequency mask estimation and spatial filtering is used for processing. Specifically, the time-frequency masks corresponding to the driver and the front passenger are first estimated, and the corresponding sound source spatial covariance matrix is constructed based on each time-frequency mask. Then, the first candidate speech stream and the second candidate speech stream are output through minimum variance distortionless response beamforming processing.
[0043] After initial voice region matching, the first candidate voice stream corresponds to the driver, and the second candidate voice stream corresponds to the front passenger.
[0044] 4. Sound source trajectory generation
[0045] For each candidate speech stream, the generalized cross-correlation phase transformation results between different microphone channels are calculated, and the peak position in the generalized cross-correlation phase transformation results is determined as the arrival time difference between the corresponding microphone pairs.
[0046] By combining the positions of each microphone in the three-dimensional coordinate system inside the vehicle and the arrival time difference between different microphone pairs, the least squares localization method is used to determine the sound source position of each candidate speech stream in each speech frame.
[0047] The sound source position of the current speech frame is used as the position state quantity, and the change in sound source position between adjacent speech frames is used as the velocity state quantity. The sound source position is smoothed by a recursive filtering method. The smoothed sound source positions are connected in time order to form the first sound source trajectory and the second sound source trajectory.
[0048] Among them, the trajectory of the first sound source remains relatively stable within the driver's side sound zone, while the trajectory of the second sound source gradually extends from the passenger side sound zone to the boundary between the passenger side sound zone and the right rear sound zone.
[0049] 5. Attribution matching between candidate speech streams and passengers
[0050] Voiceprint features and phoneme sequence features are extracted from the first candidate speech stream and the second candidate speech stream, respectively.
[0051] The voiceprint features are extracted from continuous speech segments by a speaker embedding network; the phoneme sequence features are obtained from the phoneme posterior probability sequence output by the speech recognition network during frame-by-frame decoding.
[0052] For the (q)th candidate speech stream and the (i)th passenger, the affiliation matching degree is determined according to the following formula: ; in, The spatial continuity between the current sound source location of the q-th candidate speech stream and the predicted sound source location of the i-th passenger; Let be the cosine similarity between the current voiceprint feature of the (q)th candidate speech stream and the historical voiceprint feature of the (i)th passenger; The continuity between the current phoneme sequence features and the phoneme sequence features previously attributed to the i-th member; The probability that the q-th candidate speech stream belongs to the i-th passenger at the previous time step; w p w e w c w h These are the weights corresponding to spatial continuity, voiceprint similarity, phoneme sequence continuity, and historical attribution probability, respectively.
[0053] In this embodiment, the weight is set as: w p =0.35,w e =0.35,w c =0.20,w h= 0.10; and w p +w e +w c +w h =1; For spatial continuity, it is determined based on the distance between the current sound source location and the occupant sound source location predicted based on the preceding trajectory. As the distance between the two increases, spatial continuity decreases; For voiceprint similarity, the closer the voiceprint features of the current candidate speech stream are to the historical voiceprint center of the co-pilot, the higher its voiceprint similarity; For phoneme sequence continuity, it is determined based on the changes in the posterior distribution of phonemes in adjacent speech frames to avoid a continuous speech being segmented into different occupant speech at the boundary of the speech region.
[0054] After the second candidate speech stream crosses the boundary of the passenger-side voice zone, although the distance between it and the center of the passenger-side voice zone increases, its sound source position continues to move along the trajectory of the preceding sound source, and the voiceprint features and phoneme sequence do not undergo abrupt changes. Therefore, the attribution matching degree between the second candidate speech stream and the passenger-side occupant is still greater than its attribution matching degree with the right rear temporary occupant identifier.
[0055] The attribution matching results at several representative moments in this embodiment are as follows: Table 3. Matching results of several representative moments in Example 1
[0056] The above values are used to illustrate the joint processing of each feature in this embodiment. Although the sound source location is outside the passenger's voice range at 1.2s and 1.8s, the attribution matching degree between the second candidate speech stream and the passenger remains at its maximum value.
[0057] Therefore, before the entire voice command "Open my window" ends, the system maintains the attribution relationship between the second candidate voice stream and the front passenger, without establishing a new passenger attribution due to the current sound source location entering an adjacent sound zone.
[0058] 6. Speech recognition and standardized voice command generation
[0059] Speech recognition was performed on the first and second candidate speech streams after determining the passenger affiliation, and the recognized text was obtained: First candidate voice prompt: "Navigate to the company." Second candidate voice stream: "Open my car window." Further semantic analysis of the recognized text yields standardized voice commands. These standardized voice commands are represented as: C j = (U j V j Z j A j ,P j Q j ); Among them, U j For the starting crew; V j For the target object; Z j For the target area; A j To control the action; P j For action parameters; Q j This represents the confidence level of the instruction.
[0060] For the first candidate speech stream, generate the first standardized speech command: Commander: Driver; Target object: Navigation system; Target area: Complete vehicles; Control actions: Set navigation destination; Action parameters: Company; Command confidence level: 0.93.
[0061] For the second candidate voice stream, considering that the corresponding issuing occupant is the front passenger, the instruction "I'm here" in the sentence is parsed as the front passenger area, and a second standardized voice instruction is generated: Commander: First-in-command; Target object: Passenger side window; Target area: Passenger area; Control action: Open; Action parameters: Preset opening range; Command confidence level: 0.90.
[0062] The confidence level of the command is determined by the confidence level of speech recognition, the confidence level of occupant affiliation, the semantic completeness, and the quality of overlapping speech separation. In this embodiment, the confidence levels of both standardized voice commands are greater than the preset command threshold of 0.75, therefore, it is not necessary to enter the pending confirmation state.
[0063] 7. Instruction conflict detection and output
[0064] The first standardized voice command is applied to the navigation system, and the second standardized voice command is applied to the passenger-side window. The corresponding vehicle function nodes do not overlap, so there is no command conflict.
[0065] The system outputs two standardized voice commands in parallel to the in-vehicle voice interaction interface or the corresponding vehicle function interface. The output result is as follows: (1) Set the navigation destination to "Company" according to the driver's voice; (2) Open the passenger window according to the voice of the passenger in the front seat.
[0066] In this embodiment, even if the front passenger turns to the right rear and the sound source trajectory crosses the preset sound zone boundary during the speaking process, the system still continuously assigns the entire "Open my side window" voice command to the front passenger based on the spatial continuity of the sound source trajectory, voiceprint similarity, phoneme sequence continuity, and historical attribution probability. This avoids incorrectly attributing the latter half of the voice command to the right rear passenger and ensures that the target area of "my side" is correctly resolved.
[0067] 8. Verification method in this embodiment
[0068] To verify the method of this embodiment, multiple sets of multi-channel voice data can be repeatedly collected using the same occupant, the same voice content, and the same turning trajectory, and the data can be statistically analyzed respectively. (1) The identification result of the occupant who gave the command corresponding to the complete voice command; (2) The number of times passenger affiliation was not actually switched during the duration of the voice command; (3) Whether all phonemes of a voice command are continuously bound to the same correct occupant; (4) Refers to the target area parsing results corresponding to "my side".
[0069] In contrast, the continuity of sound source trajectory, phoneme sequence, and historical attribution probability can be eliminated, and the crew member who gives the command can be determined only based on the sound zone where the current sound source is located.
[0070] If the aforementioned comparison method is used, the second candidate speech stream is easily switched to the right rear passenger identifier after the sound source position of the front passenger crosses the sound zone boundary; when using the method of this embodiment, the passenger affiliation is switched only when multiple continuity features jointly indicate that the speaker has actually changed, thereby improving the stability and integrity of the voice command affiliation in the state of cross-sound zone movement.
[0071] Example 2: Speech Stream Channel Exchange Suppression under Sound Source Trajectory Proximity Conditions
[0072] This embodiment provides a vehicle-mounted sound source localization conflict resolution method for concurrent commands from multiple occupants. It is used to solve the problem that when adjacent occupants speak at the same time and perform actions such as turning, leaning forward, or leaning forward, the sound source trajectories of different occupants are close to each other, resulting in the erroneous exchange of candidate speech streams and occupant identifiers in the overlapping speech separation results.
[0073] This embodiment uses the same vehicle, microphone array, coordinate system, speech sampling parameters, overlapping speech detection method, speech separation method, voiceprint feature extraction method, and phoneme sequence feature extraction method as Embodiment 1. The difference is that this embodiment further performs a joint time window judgment on the two candidate results of maintaining the original attribution relationship and exchanging the attribution relationship for the close interval of the sound source trajectory.
[0074] 1. Setting up concurrent voice communication scenarios for multiple passengers
[0075] In this embodiment, one passenger is seated in the front passenger seat and one in the right rear seat.
[0076] As the passenger in the front seat turned around, he gave a voice command: "Turn the music down." As the passenger in the right rear seat leans forward, they simultaneously issue a voice command: "Open the right rear window." The duration of each occupant's speech was approximately 2.2 seconds, with an overlap of approximately 0.3 to 2.0 seconds.
[0077] During the sound emission, the head of the front passenger moves backward from the front passenger seat area, while the head of the right rear passenger moves forward from the right rear seat area, causing the two sound source trajectories to approach each other in the area between the front passenger seat and the right rear seat.
[0078] The locations of the sound sources for the two occupants at several representative moments are as follows: Table 4. Sound source locations at several representative moments in Example 2
[0079] In this embodiment, the preset distance for the sound source trajectories to approach each other is set to 0.25m. Therefore, within approximately 1.2 to 1.4 seconds, the distance between the two sound source trajectories is less than the preset distance.
[0080] 2. Overlapping speech separation and initial crew assignment The overlapping speech of the multi-channel speech signal was detected in accordance with the method described in Example 1, and it was determined that the speech of the front passenger and the right rear passenger overlapped in time.
[0081] Multi-channel speech separation is performed on the overlapping speech regions to obtain a first candidate speech stream and a second candidate speech stream.
[0082] Before the two sound source trajectories are close, the first candidate speech stream is assigned to the front passenger occupant and the second candidate speech stream is assigned to the right rear passenger occupant based on the initial sound source location, voiceprint similarity, and phoneme sequence continuity of each candidate speech stream.
[0083] Establish continuously updated passenger affiliation relationships for the first and second candidate speech streams respectively.
[0084] 3. Sound source trajectory proximity interval detection Generate sound source trajectories for the first and second candidate speech streams respectively, and calculate the distance between the corresponding sound source positions of the two sound source trajectories frame by frame: d 1,2 (n) = |p1(n) - p2(n)|; Where p1(n) is the location of the first candidate speech source corresponding to the nth speech frame; p2(n) is the location of the second candidate speech source corresponding to the nth speech frame.
[0085] When multiple consecutive speech frames satisfy: d 1,2 (n) < d c At that time, it is determined that two sound source trajectories enter the sound source trajectory proximity interval, where d c The preset distance is set to 0.25m in this embodiment.
[0086] Within the range where the sound source trajectories are close together, the sound source positions calculated based on the arrival time difference between channels may fluctuate in the short term because the spatial locations of the two sound sources are close. At the same time, the acoustic feature allocation of the two candidate output channels by the overlapping speech separation model may be unstable.
[0087] In this embodiment, it was detected that within 1.28 to 1.36 seconds, the instantaneous attribution matching degree between the first candidate voice stream and the right rear passenger was once higher than that between the first candidate voice stream and the front passenger. If the occupant who gave the command is determined only based on the maximum attribution matching degree of the current voice frame, the first candidate voice stream will be incorrectly switched to the right rear passenger.
[0088] 4. Generation of candidate attribution relationships Within the interval where the sound source trajectories are close, at least two candidate correspondences are generated for the first candidate speech stream and the second candidate speech stream.
[0089] The first option is to maintain the original attribution relationship: the first candidate voice stream corresponds to the front passenger; the second candidate voice stream corresponds to the right rear passenger.
[0090] The second type is the exchange of attribution relationships: the first candidate voice stream corresponds to the right rear passenger; the second candidate voice stream corresponds to the front passenger.
[0091] For each candidate correspondence, the frame-by-frame attribution matching degree is determined based on the spatial continuity between the candidate speech stream and the corresponding occupant, the voiceprint similarity, the phoneme sequence continuity, and the historical attribution probability.
[0092] This embodiment uses the attribution matching degree calculation method in Embodiment 1, wherein the weights are set as follows: w p =0.35, w e =0.35, w c =0.20, w h= 0.10.
[0093] Within the trajectory proximity range, the distinguishing ability corresponding to spatial continuity decreases, but voiceprint similarity, phoneme sequence continuity, and historical attribution probability are still used to limit unreal switching between candidate speech streams and occupant identifiers.
[0094] 5. Cumulative attribution determination within a preset time window In this embodiment, the preset time window is set to 25 audio frames. Since the audio frame shift is 16ms, the preset time window corresponds to approximately 0.4s.
[0095] Each time it is necessary to determine whether a candidate audio stream has undergone a title swap, the cumulative title matching degree is calculated for both the original title relationship and the title swap within the preset time window.
[0096] The target correspondence is determined according to the following formula: ; Wherein, π represents a candidate correspondence between candidate speech streams and passengers; H is the number of voice frames included in the preset time window; in this embodiment, H=25. A q,π(q) (nh) represents the affiliation matching degree between the q-th candidate speech stream and the assigned crew member in the corresponding speech frame; N SW (π) represents the number of times the attribution relationship is switched within a preset time window when the corresponding candidate relationship is adopted; λ sw The handover penalty coefficient is set to 0.30 in this embodiment.
[0097] The attribution switching penalty is used to avoid frequent switching of the occupant identifier corresponding to the candidate speech stream due to fluctuations in the sound source location or voiceprint extraction results of individual speech frames.
[0098] 6. Example of target correspondence calculation At the judgment time corresponding to 1.36s, the aforementioned 25 speech frames are cumulatively calculated.
[0099] While maintaining the original attribution relationship, the cumulative attribution matching degree between the first candidate voice stream and the front passenger, and between the second candidate voice stream and the right rear passenger, is: S keep =12.84; Since this correspondence is consistent with the occupant affiliation before the sound source trajectory approaches, the number of affiliation switching times is 0, and the corresponding correction result is: 12.84 - 0.30 × 0 = 12.84; When exchanging affiliations, the cumulative affiliation matching degree between the first candidate voice stream and the right rear passenger, and between the second candidate voice stream and the front passenger, is: S swap =13.02; If this correspondence is adopted, both candidate speech streams need to switch the occupant identifier once, so the number of attribution switching times is 2, and the corresponding correction result is: 13.02 - 0.30 × 2 = 12.42; Because: J keep >J swap Therefore, the original attribution relationship will be maintained as the target correspondence relationship, that is: The first candidate voice stream continues to be attributed to the front passenger; The second candidate voice stream continues to be attributed to the right rear passenger.
[0100] Therefore, even if the exchange of affiliations achieves a slightly higher uncorrected cumulative affiliation match rate within a certain time window, it is not enough to offset the penalty caused by the simultaneous affiliation switching of the two voice streams.
[0101] 7. Attribution update after the sound source trajectory leaves the proximity interval After about 1.5 seconds, the two occupants returned to their respective seats, and the distance between the two sound source trajectories gradually increased.
[0102] The system continues to update the attribution matching degree between candidate speech streams and passengers based on the continuity of sound source location, voiceprint similarity, and phoneme sequence continuity.
[0103] At this point, the spatial continuity and voiceprint similarity between the first candidate voice stream and the front passenger occupant are significantly higher than the corresponding features between the first candidate voice stream and the right rear passenger occupant; the affiliation matching degree between the second candidate voice stream and the right rear passenger occupant also increases again.
[0104] Therefore, there is no need to perform a passenger affiliation switching, and the identity of the voice stream remains consistent before and after the sound source trajectory is close.
[0105] In another application scenario, if the cumulative correction result of the swapped affiliation over multiple consecutive preset time windows is greater than the cumulative correction result of maintaining the original affiliation, and the difference between the two exceeds a preset switching threshold, then updating the occupant affiliation according to the swapped affiliation is allowed. This avoids setting the switching penalty to absolutely prohibit affiliation changes, while still being able to identify actual speaker switching or channel rearrangement.
[0106] 8. Speech recognition and standardized voice command generation Speech recognition is performed on the first and second candidate speech streams after determining the affiliation of passengers.
[0107] The first candidate speech stream was identified as the text: "Turn the music volume down". The recognized text of the second candidate speech stream is: "Open the right rear window". (1) Parse the first candidate speech stream into the first standardized speech command: Commander: First-in-command; Target audience: Car audio systems; Target area: Vehicle audio system or the corresponding audio zone on the passenger side; Control action: Lower volume; Action parameters: preset reduction amount; Command confidence level: 0.88.
[0108] (2) Parse the second candidate speech stream into the second standardized speech command: Commander: Passenger in the right rear row; Target object: Right rear window; Target area: Right rear row area; Control action: Open; Action parameters: Preset opening range; Command confidence level: 0.91.
[0109] If the occupant affiliation of two candidate speech streams is incorrectly swapped within the range where the sound source trajectories are close, the first standardized speech command will be incorrectly assigned to the right rear passenger, and the second standardized speech command will be incorrectly assigned to the front passenger, thus affecting the local target area resolution and subsequent speech permission judgment.
[0110] This embodiment maintains the binding relationship between the recognized text and the correct commanding passenger by accumulating judgments within a preset time window.
[0111] 9. Instruction conflict detection and output The first standardized voice command applies to the vehicle audio node, and the second standardized voice command applies to the right rear window node. The corresponding vehicle function nodes do not overlap, so there is no command conflict.
[0112] System parallel output: (1) Lower the volume of the car audio system according to the voice of the front passenger; (2) Open the right rear window according to the voice of the right rear passenger.
[0113] In this embodiment, the proximity of sound source trajectories only triggers stability correction for occupant affiliation, and does not directly determine the two voice commands as semantic conflicts. Semantic conflicts are still determined based on the standardized voice commands formed after speech recognition.
[0114] 10. Verification method in this embodiment The same voice content can be used, but different occupant movement trajectories and minimum distances to the sound source can be set to repeatedly collect multiple sets of multi-channel voice data. The movement trajectory includes at least: (1) The front passenger turns around while the right rear passenger remains stationary; (2) The front passenger remains still, while the right rear passenger leans forward; (3) Both passengers lean forward between the front and rear rows at the same time; (4) The two sound source trajectories approach each other and then separate again; (5) The two passengers finished speaking at different times.
[0115] Statistics for each set of voice data: (1) Number of non-true attribution switching of candidate speech streams; (2) The occupant identification is consistent before and after the sound source trajectory is close to the same speech stream; (3) The accuracy of binding between speech recognition text and the correct commanding occupant; (4) Local target region parsing accuracy.
[0116] In contrast, the preset time window cumulative judgment and the affiliation switching penalty can be cancelled, and the occupant corresponding to the candidate voice stream can be determined only based on the maximum affiliation matching degree of each voice frame.
[0117] In this comparison mode, when two sound source trajectories enter the proximity interval and the instantaneous attribution matching degree crosses, the candidate speech stream is prone to passenger identification exchange. When using the method of this embodiment, the passenger attribution is changed only when the exchanged attribution relationship continuously obtains sufficient matching gain within the preset time window and can offset the attribution switching penalty, thereby reducing the channel switching risk in the process of overlapping speech separation.
[0118] Example 3: Joint resolution of semantic modification of the same occupant and vehicle function dependency conflict This embodiment provides a vehicle sound source localization conflict resolution method for concurrent commands from multiple occupants. It is used to solve the problem that during concurrent voice interaction among multiple people, the spoken corrections of the same occupant are easily misjudged as conflicting independent commands, and that voice commands issued by different occupants, although targeting different objects, actually act on the same vehicle functional nodes, thus forming indirect conflicts.
[0119] This embodiment uses the same vehicle, microphone array, speech sampling parameters, overlapping speech detection method, speech separation method, sound source trajectory generation method, and occupant attribution method as Embodiment 1. The difference lies in that, after attributing candidate speech streams to occupants, this embodiment further identifies the correction relationship between the same occupant's voice commands and determines the indirect conflicts between different occupant voice commands based on the role dependencies between vehicle functional nodes.
[0120] 1. Setting up concurrent voice communication scenarios for multiple passengers In this embodiment, the driver and the passenger in the right rear seat issue voice commands respectively.
[0121] The driver repeatedly said, "Turn on the windshield defroster, set the temperature to 22 degrees, no, 24 degrees." The passenger in the right rear seat simultaneously said, "Turn off the air conditioning," while the driver was speaking. The driver's voice lasted approximately 3.4 seconds, while the voice of the right rear passenger lasted approximately 1.2 seconds, with an overlap between the two within approximately 0.6 to 1.8 seconds.
[0122] The driver's speech includes the following three semantic segments: (1) First semantic fragment: "Turn on the windshield defroster"; (2) Second semantic segment: "The temperature is set to twenty-two degrees"; (3) The third semantic segment: “No, twenty-four degrees.”
[0123] The third semantic segment is the driver's correction of the temperature parameter in the second semantic segment, rather than a competing command issued by another occupant.
[0124] The "turn off the air conditioner" command issued by the rear passenger on the right and the "turn on the windshield defroster" command issued by the driver have different target objects in terms of text. However, the windshield defroster function requires the compressor and blower to be activated, while turning off the air conditioner will stop or reduce the working status of the compressor and blower. Therefore, the two commands have an indirect conflict at the vehicle function node level.
[0125] 2. Overlapping speech separation and passenger attribution Based on the eigenvalue distribution of the spatial covariance matrix of the multi-channel speech signals, the overlapping speech intervals of the driver's speech and the right rear passenger's speech were detected.
[0126] The overlapping speech regions are separated to obtain a first candidate speech stream and a second candidate speech stream.
[0127] Based on the sound source trajectory, voiceprint features, phoneme sequence features, and historical attribution probability of the first and second candidate speech streams, it is determined that the first candidate speech stream belongs to the driver and the second candidate speech stream belongs to the right rear passenger.
[0128] In this embodiment, the distance between the two sound source trajectories is always greater than the preset intersection distance, thus not triggering the voice stream channel exchange correction process described in Embodiment 2.
[0129] 3. Speech recognition and initial standardized voice command generation Speech recognition and semantic parsing are performed on the first candidate speech stream and the second candidate speech stream, respectively.
[0130] The first candidate speech stream corresponds to the driver's speech. After detecting sentence boundaries and semantic object changes, the following three initial standardized speech commands are formed: (1) First initial voice command (C_1): Commander: Driver; Target object: windshield defroster; Target area: The area of the vehicle's windshield; Control action: Turn on; Action parameters: Preset defogging mode; Speech recognition confidence level: 0.95; Crew attribution confidence level: 0.96.
[0131] (2) Second initial voice command (C_2): Commander: Driver; Target object: Air conditioning temperature; Target area: Driving area; Control Actions: Settings; Action parameters: 22℃; Speech recognition confidence level: 0.94; Crew attribution confidence level: 0.96.
[0132] (3) Third initial voice command (C_3): Commander: Driver; Target object: to be inherited; Target region: to be inherited; Control action: parameter correction; Action parameters: 24℃; Corrected expression: "No"; Speech recognition confidence level: 0.92; Crew attribution confidence level: 0.95.
[0133] (4) The second candidate speech stream corresponds to the speech of the passenger in the right rear row, forming the fourth initial speech command (C_4): Commander: Passenger in the right rear row; Target object: Air conditioning system; Target area: Not explicitly defined; Control action: Close; Action parameters: None; Speech recognition confidence level: 0.90; Crew attribution confidence level: 0.93.
[0134] 4. Determining the confidence level of the instruction In this embodiment, the confidence level of each initial voice command is determined based on the confidence level of speech recognition, the confidence level of passenger affiliation, the semantic completeness, and the quality of overlapping speech separation.
[0135] For example, the confidence level of the initial speech instruction (j) can be determined using the following formula: ; in, For speech recognition confidence; Confidence level for occupant attribution; For semantic completeness; α represents speech separation quality; α, β, γ, and δ are weighting indices greater than 0.
[0136] In this embodiment, all weight indices are set to 0.25. Calculations show that the confidence scores of the first to fourth initial voice commands are all greater than the preset command threshold of 0.75, therefore they all proceed to the subsequent semantic relationship judgment.
[0137] Although the target object is omitted in the third initial speech instruction C_3, it contains explicit corrective expressions and new parameters. Its semantic completeness is further determined by the inheritance relationship of the target object of the preceding speech instruction, and it is not directly deleted because the target object is omitted.
[0138] 5. Identification of semantic modification relationships for the same occupant Within a preset correction time window, the relationship between initial voice commands that are adjacent in time is determined.
[0139] In this embodiment, the preset correction time window is set to 2.0s. The time interval between the second initial voice command C_2 and the third initial voice command C_3 is 0.35s, which meets the correction time window requirement.
[0140] Based on the occupant affiliation probability distribution of the two initial voice commands, the probability of the same commanding occupant is determined. The probability of the same commanding occupant can be expressed as: Among them, P 2,i P is the probability that the second initial voice command belongs to the i-th passenger. 3,i The probability that the third initial voice command belongs to the i-th passenger.
[0141] In this embodiment, both the second and third initial voice commands are primarily attributed to the driver, resulting in a probability of the same person issuing the command being 0.93, which is greater than the preset threshold of 0.80.
[0142] Further determine whether the third initial voice command meets the correction conditions: the third initial voice command contains the negative correction expression "no"; the parameter type of the third initial voice command is temperature value, which is consistent with the action parameter type of the second initial voice command; the target object is omitted in the third initial voice command, but the time interval between it and the second initial voice command is less than the preset correction time window; both the second and third initial voice commands belong to the driver.
[0143] Therefore, the third initial voice command is determined to be a parameter overriding command of the second initial voice command, rather than an independent competing voice command.
[0144] The target object and target area in the third initial voice command are inherited as "air conditioning temperature" and "driving area" respectively, and 24℃ is used to cover 22℃ in the second initial voice command, resulting in the corrected temperature voice command C_23: Commander: Driver; Target object: Air conditioning temperature; Target area: Driving area; Control Actions: Settings; Action parameters: 24℃.
[0145] After the correction is completed, the 22℃ parameter in the second initial voice command becomes invalid and will no longer be used as an independent voice command in subsequent multi-person command conflict determination.
[0146] Therefore, the currently valid voice commands include: (1) First voice command: Turn on the windshield defroster; (2) Corrected temperature voice command: Set the driving area temperature to 24℃; (3) Fourth voice command: Turn off the air conditioner.
[0147] 6. Construction of the Dependency Graph for Vehicle Functional Nodes To determine the actual functional relationships of different voice commands at the vehicle function level, a dependency graph of vehicle function nodes is established.
[0148] The vehicle functional nodes involved in this embodiment include: v_1: Front windshield defroster; v_2: Main function of air conditioning; v_3: Compressor; v_4: Blower; v_5: External circulation damper; v_6: Driver's area temperature control; v_7: Air vents in the driver's area; v_8: Rear air vent.
[0149] The functional dependencies between vehicle functional nodes include at least the following: (1) Turning on the windshield defroster will activate the compressor; (2) Turning on the windshield defroster will activate the blower; (3) Turning on the windshield defroster will control the external air circulation damper; (4) Turning off the main function of the air conditioner will stop or limit the compressor; (5) Turning off the main function of the air conditioner will stop or limit the blower; (6) Turning off the main air conditioning function will close the air vents in the driver's area and the rear air vents; (7) The driver's area temperature adjustment will activate the compressor, blower and driver's area air outlet.
[0150] Construct an action dependency matrix H based on the above action dependencies. Matrix elements H kl Represents the functional node v k State changes affect functional node v l The extent of the impact.
[0151] In this embodiment, the influence values of different components are as follows: Table 5. Influence values of different vehicle functional nodes
[0152] The influence value is used to determine the vehicle function nodes that are directly controlled and indirectly affected by each voice command. The specific value can be adjusted according to the air conditioning system structure of the vehicle model and the actual vehicle calibration results.
[0153] 7. Control scope vector generation A direct target vector is generated based on the target object and target region of each valid voice command.
[0154] The direct target node of the first voice command "Turn on windshield defroster" is the windshield defroster node (v_1), and the corresponding direct target vector is denoted as u1.
[0155] The direct target node for the corrected temperature voice command "Set the driving area temperature to 24℃" is the driving area temperature adjustment node v_6, and the corresponding direct target vector is denoted as u. 23 .
[0156] The direct target node of the fourth voice command "Turn off the air conditioner" is the main function node v_2 of the air conditioner, and the corresponding direct target vector is denoted as u4.
[0157] Based on the dependency relationships of vehicle function nodes, the control scope vector of each voice command is determined according to the following formula: ; Among them, s j Let be the control scope vector of the j-th voice command; R is the maximum propagation order; μ is the propagation attenuation coefficient; clip(·,0,1) means restricting each element of the vector to between 0 and 1.
[0158] In this embodiment, the maximum propagation order (R) is set to 2, the propagation attenuation coefficient (μ) is set to 0.65, and the action threshold is set to 0.35.
[0159] After transmission, the control domain corresponding to the first voice command includes at least: windshield defroster node; compressor node; blower node; and external circulation damper node.
[0160] The control domain corresponding to the revised temperature voice command includes at least the following: driver area temperature adjustment node; compressor node; blower node; and driver area air outlet node.
[0161] The control domain corresponding to the fourth voice command includes at least the following: air conditioning main function node; compressor node; blower node; driver area air vent node; and rear air vent node.
[0162] 8. Instruction Conflict Determination First, compare the first voice command with the fourth voice command.
[0163] The two voice commands have different direct targets: the windshield defroster and the main air conditioning function. However, after action-dependent propagation, both affect the compressor node and the blower node.
[0164] At the compressor node and blower node: (1) The first voice command requires the corresponding node to be in the enabled state; (2) The fourth voice command requires the corresponding node to be in a stopped or restricted state.
[0165] Therefore, the first voice command and the fourth voice command have opposite control action directions at the common indirectly affected node, thus determining that the two constitute an action dependency conflict.
[0166] Then compare the revised temperature voice command with the fourth voice command.
[0167] The revised temperature voice command requires adjusting the temperature in the driver's area through the compressor, blower, and air vents in the driver's area, while the fourth voice command requires turning off the main air conditioning function, thereby stopping or restricting the aforementioned vehicle function nodes.
[0168] Therefore, the revised temperature voice command and the fourth voice command also constitute an action dependency conflict.
[0169] Both the first voice command and the revised temperature voice command are issued by the driver, and their control actions are not contradictory. They can jointly activate the compressor and blower without conflicting with each other.
[0170] 9. Conflict Function Node Division The vehicle function nodes corresponding to the first voice command, the corrected temperature voice command, and the fourth voice command are divided.
[0171] The shared conflict nodes between the fourth voice command and the driver's two voice commands include: the compressor node; the blower node; and the air outlet node in the driver's area.
[0172] The fourth voice command corresponds to the following functional nodes that do not directly conflict with the windshield defrost and driver's area temperature adjustment: rear air vent node.
[0173] Regarding the fourth voice command, "Turn off the air conditioning," since the commanding passenger is a right rear passenger and the command does not explicitly define the target area, the system first combines the area where the commanding passenger is located and the voice context to interpret it as prioritizing the operation of the rear comfort functions corresponding to the right rear passenger.
[0174] If "turn off the air conditioner" is directly interpreted as turning off the main function of the vehicle's air conditioning, it will disrupt the shared function nodes required for windshield defogging and driver's area temperature adjustment; therefore, the fourth voice command needs to be reconstructed locally.
[0175] 10. Reconstruction of conflicting voice commands Based on the operator issuing each voice command, the confidence level of the command, the overlap of functional nodes, and the direction of control actions, the semantics of the valid voice commands are reconstructed.
[0176] The first voice command, "Turn on windshield defroster," remains unchanged to maintain the compressor, blower, and external air circulation damper required for windshield defroster operation.
[0177] The revised temperature voice command retains its final parameter of 24°C and outputs: "Set the driver's area air conditioning temperature to 24°C". For the fourth voice command "Turn off the air conditioning", remove the compressor, blower and driver area air vent nodes that conflict with the windshield defrost and driver area temperature adjustment from its control domain, retain the non-conflicting vehicle function node corresponding to the right rear passenger, and reconstruct it as: "Close the rear air vent". Alternatively, if the vehicle's air conditioning system does not support completely closing the rear exhaust vents, it can be reconfigured as: "Adjust the rear exhaust airflow to the lowest setting". This embodiment ultimately outputs the following set of semantic instructions after conflict resolution: (1) Turn on the windshield defroster; (2) Set the air conditioning temperature in the driver's area to 24℃; (3) Close the rear exhaust vent, or adjust the rear exhaust air volume to the lowest setting.
[0178] Through the above reconstruction, the voice needs of the right rear passenger were not simply discarded, nor was the vehicle's air conditioning directly turned off according to fixed identity priority. Instead, the intention to control the vehicle locally on non-conflicting functional nodes was retained.
[0179] 11. Low confidence level pending confirmation processing In one modified implementation, if the target region in the voice of the right rear passenger is unclear, and the semantic confidence of interpreting "turn off the air conditioner" as "close the rear air vents" is lower than a preset reconstruction threshold, then the reconstructed local semantic instruction is not directly output, but a semantic instruction to be confirmed is generated instead: "Should the exhaust fan be turned off?" Bind the right rear passenger identifier and the rear air vent target node to the semantic instruction to be confirmed.
[0180] If the driver and the right rear passenger respond simultaneously, the system processes the confirmation voice according to the passenger attribution method described in Embodiments 1 and 2. Only voices belonging to the right rear passenger and containing affirmative confirmation are used to update the pending semantic instruction into a valid local semantic instruction.
[0181] For example, if the passenger in the right rear seat answers "yes" and the driver simultaneously says "continue navigation," then only the passenger's "yes" will be linked to the pending confirmation semantic command "whether to turn off the rear air vents," thus preventing irrelevant voice from being mistaken for confirmation content.
[0182] 12. Verification method in this embodiment To verify the method of this embodiment, multiple scenarios can be set up, including self-correction by the same crew member and conflict of action dependencies between different crew members, including: (1) The phrases "set the temperature to 22 degrees, no, 24 degrees" and "turn off the air conditioner" appear simultaneously; (2) “Open the sunroof, or better yet, open the passenger window” and “Close all windows” appear simultaneously; (3) The words “Turn the volume up, no, turn it down” and the other passenger’s words “Keep the volume the same” appear at the same time; (4) The options "Turn on windshield defroster" and "Turn off blower" appear simultaneously; (5) "Close all windows" and "Open my window" appear at the same time.
[0183] Statistical breakdown: (1) The proportion of passengers whose parameters are corrected or whose target objects are replaced is correctly identified; (2) The proportion of the same passenger's correction statement that was incorrectly identified as a multi-person conflict; (3) The proportion of action dependency conflicts detected between different target objects; (4) The percentage of effective local speech demand retained after conflict resolution; (5) The accuracy of the binding between the voice to be confirmed and the target occupant.
[0184] As a control, the following treatment methods can be used: (1) The correction relationship of the same passenger is not recognized, and each recognition statement is regarded as an independent instruction; (2) Do not construct the dependency relationship of vehicle function nodes, only compare the direct target objects of voice commands; (3) In the event of a conflict, only one complete voice command will be selected according to the priority of the passenger's identity.
[0185] When using the aforementioned comparison method, "22℃" and "24℃" are easily regarded as two conflicting temperature commands; "turn on windshield defroster" and "turn off air conditioning" are easily mistakenly judged as being able to be executed in parallel because the target object names are different; or when the driver priority strategy is adopted, the effective demand of the right rear passenger to reduce the rear air conditioning output is completely discarded.
[0186] When using the method of this embodiment, the spoken corrections of the same passenger are first merged into the final effective semantics. Then, based on the direct and indirect interaction of vehicle functional nodes, the actual conflicts between different passenger instructions are identified, and the non-conflicting parts are reconstructed into local semantic instructions, thereby reducing the missed detection of false conflicts and indirect conflicts, and improving the effective retention of concurrent voice instructions from multiple people.
[0187] Comparative Example 1: Crew Attribution Method Based on Fixed Pitch Range and Frame-by-Frame Sound Source Location This comparative example uses the same vehicle, microphone array, speech sampling parameters, overlapping speech detection method, and multi-channel speech separation method as Examples 1 and 2. The difference lies in that this comparative example does not generate and continuously track sound source trajectories, does not combine voiceprint features, phoneme sequence features, and historical attribution probabilities to maintain the occupant affiliation of candidate speech streams, and does not employ preset time window cumulative matching and attribution switching penalties when sound source trajectories are close. Instead, it determines the commanding occupant of the candidate speech stream solely based on the audio region where the current sound source location corresponds to each speech frame.
[0188] 1. Fixed register division Based on the position of the vehicle seats, the interior space is divided into the driver's seat audio zone, the passenger's seat audio zone, the left rear audio zone, the middle rear audio zone, and the right rear audio zone.
[0189] For each candidate speech stream, calculate the sound source location corresponding to the current speech frame and determine the preset pitch range to which the sound source location belongs. Identify the rider in the corresponding preset pitch range as the commanding rider for the current speech frame.
[0190] If the sound source location of a candidate speech stream moves from one preset pitch zone to another preset pitch zone, then starting from the speech frame where the sound source location crosses the pitch zone boundary, the candidate speech stream is switched to the passenger in the other preset pitch zone.
[0191] This comparative example does not determine whether the change in the location of the sound source was caused by the same occupant turning around, leaning forward, or changing the seat position.
[0192] 2. Cross-range scenario corresponding to Example 1 Using the voice scenario in Example 1: The driver said, "Navigate to the company." As the passenger in the front seat turned to the right rear, he simultaneously said, "Open my window." During the process of the passenger speaking, the sound source location moves from the passenger's sound zone to outside the boundary between the passenger's sound zone and the right rear sound zone.
[0193] When the sound source is still located in the front passenger area, this comparison model will assign the corresponding candidate voice stream to the front passenger; after the sound source crosses the preset area boundary, this comparison model will assign the subsequent voice frames to the right rear passenger.
[0194] Therefore, the voice command "Open my car window" could be categorized as: (1) The preceding voice clip “Open my side” belongs to the front passenger; (2) The following audio segment “car window” belongs to the right rear passenger.
[0195] In another approach, if the system determines the occupant who gave the command based on the location of the sound source at the end of the voice, the entire voice error may be attributed to the occupant in the right rear row.
[0196] Therefore, when performing target region parsing on "my side", one of the following results is likely to occur: (1) The corresponding crew member for "I" cannot be determined; (2) The target object was incorrectly parsed as the right rear window; (3) The entire voice command is discarded because the occupant identification of the preceding and following voice segments is inconsistent; (4) Triggering unnecessary occupant confirmation process.
[0197] 3. The sound source trajectory approximates the scene corresponding to Example 2. Using the voice scenario in Example 2: As the passenger in the front seat turned around, he said, "Turn the music down." As the passenger in the right rear seat leaned forward, he simultaneously said, "Open the right rear window." When the sound sources of two occupants are close to each other, due to in-vehicle reflection, microphone positioning error, and fluctuations in overlapping speech separation, the sound source location of the first candidate speech stream in some speech frames may be closer to the right rear row sound area, while the sound source location of the second candidate speech stream in some speech frames may be closer to the front passenger sound area.
[0198] This comparison method determines passenger affiliation based on the current sound source location of each speech frame. When the sound source locations of two candidate speech streams briefly overlap, the passenger identifiers of the two candidate speech streams are switched respectively.
[0199] This may lead to the following: (1) The first candidate speech stream belongs to the front passenger before the sound source approaches, and belongs to the right rear passenger during the approach of the sound source; (2) The second candidate speech stream belongs to the right rear passenger before the sound source approaches, and belongs to the front passenger during the approach of the sound source; (3) The same continuous voice stream switches passenger affiliation multiple times within a short period of time; (4) The binding relationship between the voice recognition text and the commanding crew is exchanged.
[0200] For example, "turn down the music volume" might be incorrectly assigned to the right rear passenger, and "open the right rear window" might be incorrectly assigned to the front passenger, thus affecting local sound zone resolution, window target area resolution, and subsequent permission judgment.
[0201] 4. Comparative Analysis This comparative example only uses the current sound source location and fixed pitch zone boundaries to determine the commanding crew member, and cannot distinguish the following two situations: (1) The same passenger causes the sound source to cross the boundary of the sound range due to turning around or leaning forward; (2) The actual crew members who gave the order changed.
[0202] Furthermore, this comparative model cannot utilize the continuity of sound source location, voiceprint similarity, and phoneme sequence within a preset time window to correct for short-term localization fluctuations. When two sound sources are close to each other, non-realistic channel switching can easily occur due to changes in the attribution matching results of individual speech frames.
[0203] Compared with the comparative example, Example 1 maintains the passenger affiliation during cross-sound zone movement by combining sound source trajectory, voiceprint, phoneme sequence and historical affiliation probability; Example 2 reduces the possibility of non-real identity exchange in candidate voice streams by comparing the cumulative results of maintaining the original affiliation relationship and exchanging the affiliation relationship within the time window and penalizing the affiliation switch.
[0204] Comparative Example 2: Voice command conflict handling method based solely on direct target and occupant priority This comparative example uses the same vehicle, microphone array, speech sampling parameters, speech separation method, sound source localization method, occupant attribution method, and speech recognition method as Example 3. The differences are: (1) It does not recognize the relationship between cancellation, parameter overwriting, target object replacement or area of action adjustment between adjacent voice commands of the same occupant; (2) Treat each speech recognition segment as an independent speech command; (3) No dependency relationship is established between vehicle function nodes; (4) Uncertain vehicle function nodes that are directly controlled and indirectly affected by voice commands; (5) A conflict is identified only when the direct target of different voice commands is the same and the control actions are opposite; (6) When a conflict is detected, one complete instruction is retained according to the preset crew priority, and the other complete instruction is discarded.
[0205] In this comparative study, the driver is set as the first priority, the front passenger as the second priority, and the rear passenger as the third priority.
[0206] 1. Same occupant parameter correction scenario The driver's voice from Example 3 is used: "Turn on the windshield defroster and set the temperature to 22 degrees, no, 24 degrees." This comparison example generates the following independent instructions based on speech pauses and recognition results: (1) Turn on the windshield defroster; (2) Set the temperature in the driving area to 22℃; (3) Set the temperature of the driving area to 24°C.
[0207] Since the direct target of the second and third instructions is the air conditioning temperature, and the action parameters are different, this comparison example determines that the two are temperature setting conflicts.
[0208] If the instruction is processed according to the time it is received, the 22°C setting may be executed first, followed by the 24°C setting, causing unnecessary duplicate control. If the first-to-arrive instruction is prioritized, the subsequent 24°C correction result may be discarded. If the last-to-arrive instruction is prioritized, although the 24°C may be obtained, it cannot recognize that the subsequent statement is a self-correction of the preceding statement, nor can it handle other correction types such as target object omission, range adjustment, or cancellation.
[0209] Therefore, this comparative ratio cannot reliably distinguish: (1) Correction of spoken language by the same passenger; (2) Competing instructions issued by different crew members.
[0210] 2. Indirect conflict scenarios between different direct target objects Continuing with the concurrent speech in Example 3: The driver said, "Turn on the windshield defroster." The passenger in the right rear row said, "Turn off the air conditioning." This comparative example identifies the direct target objects of the two instructions as: windshield defroster and air conditioning system, respectively.
[0211] Since the direct target object names of the two instructions are different, this comparison model determines that there is no conflict between them and outputs both instructions simultaneously.
[0212] During vehicle function execution, "turn on windshield defroster" requires the compressor and blower to remain active, while "turn off air conditioning" will stop or limit the compressor and blower. When both commands are output simultaneously, the following may occur: (1) The air conditioner shuts off, causing the compressor to stop, which prevents the windshield defogging effect from being achieved normally; (2) The windshield defrosting action restarts the blower, causing the "air conditioning off" output result to be unsustainable; (3) The two vehicle function control processes alternately cover the same execution node; (4) The vehicle voice system reported that both commands had been executed, but the actual vehicle status could not satisfy both commands at the same time.
[0213] Therefore, it is evident that simply comparing the direct target objects cannot identify the dependency conflict of function nodes of shared vehicles that are different in semantics.
[0214] 3. Priority handling for fixed occupants In another comparative processing method, in order to avoid the simultaneous execution of "turn on windshield defroster" and "turn off air conditioning", this comparative example retains the driver's "turn on windshield defroster" according to the priority of passengers and discards the right rear passenger's "turn off air conditioning" when a preset conflict is detected between air conditioning related functions.
[0215] While this approach avoids shutting down the compressor and blower, it also eliminates the right rear passenger's valid request to reduce the rear air conditioning output.
[0216] In fact, the voice request from the right rear passenger can be reconfigured as "close the rear air vents" or "reduce the rear airflow," and this local command does not require stopping the compressor and blower required for the windshield defroster.
[0217] This comparative model does not distinguish between overlapping and non-overlapping vehicle function nodes of different instructions, and therefore cannot retain the non-conflicting parts of low-priority instructions.
[0218] 4. Scenarios of set instructions and local instructions To further illustrate the limitations of this comparative example, the following concurrent audio can be configured: The driver ordered, "Close all windows." The passenger in the front seat said, "Open my window." If this comparison is made only according to the string or preset category of the direct target object, "all windows" and "passenger window" may be identified as different target objects, thus missing the conflict. If both are classified into the window category and the driver is given priority, the partial requirement of closing all windows and discarding the passenger in the front seat will be executed.
[0219] This comparative model cannot split the set control range corresponding to "all windows" into the passenger side window and other windows, nor can it generate partial closing commands for other non-conflicting windows.
[0220] 5. Comparative Analysis This comparative example treats each identified segment as an independent instruction, which may easily lead to misjudgment of multiple conflicts when the parameter correction of the same occupant is made. At the same time, by only comparing the direct target objects, it is easy to miss indirect conflicts caused by the dependency relationship of vehicle functional nodes.
[0221] Even if fixed crew priority is further adopted, choices can only be made between multiple complete instructions. Conflicting instructions cannot be split into shared conflicting and non-conflicting parts, resulting in the complete discarding of effective local requirements of low-priority crew members.
[0222] Compared to the comparative example, Example 3 first identifies the semantic modification relationship of the same occupant based on the consistency of the issuing occupant, the inheritance of the target object, and the modification expression. Then, it extends each valid voice command along the dependency relationship of the vehicle functional nodes to the directly controlled and indirectly affected nodes, and generates local semantic commands for non-conflicting functional nodes. Therefore, it can simultaneously reduce false conflicts, missed detection of indirect conflicts, and the overall discarding of valid commands.
[0223] Effect evaluation
[0224] I. Purpose of Evaluation
[0225] To verify the processing effect of the present invention in a multi-occupant concurrent voice environment in a vehicle, tests were conducted on Examples 1-3, Comparative Example 1, and Comparative Example 2, respectively.
[0226] The effectiveness evaluation mainly verifies the following aspects: (1) When a passenger crosses the boundary of a vocal range during speech, does the passenger affiliation of the speech stream remain stable? (2) When the sound source trajectories of the two occupants are close to each other, does the candidate speech stream undergo non-real channel exchange? (3) Whether the cancellation, parameter modification or target object replacement of the same occupant is correctly identified; (4) Whether instructions that have different direct target objects but overlap in their effect on vehicle function nodes are identified as conflicting; (5) After the conflict is resolved, are the non-conflicting local speech requirements preserved?
[0227] II. Composition of Test Data Test data may include data collected from the actual vehicle and supplementary data generated based on the impulse response of the actual vehicle compartment.
[0228] 1. Cross-regional speech data Different passengers can issue preset voice commands from the front passenger seat, left rear seat, and right rear seat, and perform one of the following actions while issuing the commands: Turn your head toward the seat next to you; Lean forward or backward; Bending down to pick up the item; Adjust the seat position forward and backward; Change from a normal sitting posture to a leaning posture.
[0229] Each speech entry contains at least one expression referring to the target area, such as "this side of me," "this one next to me," or "the car window behind me."
[0230] 2. Sound source trajectory proximity data The front-row passengers and the adjacent rear-row passengers speak simultaneously, each leaning towards the other, so that the minimum distance between the two sound source trajectories is distributed in different intervals, for example: Greater than 0.40m; 0.25~0.40m; 0.15~0.25m; Less than 0.15m.
[0231] The voice content is directed to different vehicle functions to determine whether there is an incorrect exchange between the candidate voice stream and the occupant identification.
[0232] 3. Semantic correction data for the same occupant Semantic correction data includes at least: Parameter coverage: "22 degrees, no, 24 degrees"; Command cancelled: "Open the sunroof, forget it"; Target object replacement: "Open the sunroof, never mind, open the right-side window"; Narrowing the scope: "Close all windows, I mean just the rear windows"; Expanding the scope: "Close the right rear window, no, close all of them."
[0233] 4. Vehicle functions depend on conflicting data. It should include at least the following combination of instructions: Turn on the windshield defroster and turn off the air conditioning; Turn on the windshield defroster and turn off the blower; Adjust the driver's area temperature and turn off the vehicle's air conditioning; Lower the local volume and turn off the main speaker amplifier; Close all windows and open some windows.
[0234] Each set of test data should record the actual occupant giving the command, the actual voice and text, the actual correction relationship, the direct target object, the actual affected vehicle function nodes, and the expected conflict resolution results, as evaluation labels.
[0235] III. Evaluation Indicators
[0236] 1. Crew assignment error rate
[0237] The occupant assignment error rate represents the proportion of voice commands that are assigned to the wrong occupant, and is calculated using the following formula: ; in, N represents the number of voice commands that the occupant misidentifies. cmd The total number of voice commands that participated in the evaluation.
[0238] The lower the speaker attribution error rate, the more accurate the speaker attribution is in cross-regional and concurrent speech environments.
[0239] 2. Number of times non-true ownership has been switched
[0240] For each continuous voice stream, count the number of times its corresponding occupant identifier changes when the actual occupant giving the command remains unchanged.
[0241] The total number of non-true attribution switching is: ; Where, N stream N represents the number of continuous speech streams. sw,k The number of non-real-home handovers that occur in the k-th continuous voice stream.
[0242] 3. Command completeness rate Command completeness rate refers to the proportion of times when all valid voice segments of a voice command are consistently bound to the same correct occupant: ; In the formula, N int The number of complete voice commands that did not result in erroneous splitting or identity swapping.
[0243] 4. Improve the accuracy of relationship identification. Correction relationship recognition accuracy represents the proportion of relationships that the system correctly identifies when undoing, parameter overriding, target object replacement, and scope adjustment are performed. ; in, The number of samples required to correctly identify the correction type and obtain the correct final semantics; N rep This represents the total number of samples that include semantically modified relations.
[0244] 5. False Conflict Rate The false conflict rate represents the proportion of corrections or supplementary statements from the same crew member that are incorrectly identified as multiple conflicting instructions: ; in, N represents the number of corrected or supplementary samples that were incorrectly identified as conflicting. nonconf This represents the actual number of samples that do not constitute a conflict involving multiple people.
[0245] 6. Indirect conflict false negative rate The indirect conflict miss rate represents the proportion of instructions that, despite having different direct target objects, actually conflict after propagation through the dependency relationship of vehicle functional nodes, are not detected. ; in, The number of unidentified action-dependent conflicts; This represents the total number of samples that actually exhibit conflicting effects.
[0246] 7. Effective instruction retention rate Effective instruction retention rate represents the proportion of valid crew requirements among conflicting instructions that can be retained as complete semantic instructions or non-conflicting partial semantic instructions: ; Where, N keep N represents the number of valid semantic requirements retained after conflict resolution. valid This refers to the number of effective demands that can actually be fully or partially satisfied in a conflict scenario.
[0247] IV. Evaluation Process
[0248] 1. Comparison of Example 1 and Comparative Example 1
[0249] The cross-regional speech data was input into the methods of Example 1 and Comparative Example 1, respectively, and the following statistics were compiled: Crew assignment error rate; Number of times non-true ownership has been switched; Command integrity rate; This refers to the accuracy of the parsing of the corresponding target area.
[0250] In samples where the sound source location does not cross the boundary of the sound region, both methods can rely primarily on the sound source location to complete the initial assignment.
[0251] In samples where the sound source location crosses the boundary of the sound zone, Comparative Example 1 changes the passenger affiliation according to the sound zone where the current sound source is located, which easily increases the switching of non-true affiliation and instruction splitting; Example 1 integrates the sound source trajectory, voiceprint similarity, phoneme sequence continuity and historical affiliation probability, and only updates the affiliation when multiple features indicate that the passenger giving the command has actually changed, thus maintaining the passenger binding of complete voice commands.
[0252] 2. Comparison of Example 2 and Comparative Example 1 The sound source trajectory proximity data were input into the methods of Example 2 and Comparative Example 1, respectively, and statistical analysis was performed according to different minimum sound source distance intervals: The number of non-true attribution switching in candidate audio streams; Number of samples exchanged through the channel; The binding results of speech recognition text with the correct commanding passenger.
[0253] As the minimum distance between the two sound source trajectories decreases, the fluctuation of the frame-by-frame localization results usually increases. Comparative Example 1 directly uses the frame-by-frame maximum matching result, which easily leads to multiple switching of occupant identifiers between adjacent speech frames.
[0254] Example 2 compares the cumulative matching results of maintaining the original attribution relationship and the exchanged attribution relationship within a preset time window, and applies a penalty to the switching of attribution relationship. This can suppress channel switching caused by short-term positioning errors or separation fluctuations, while retaining the ability to perform real attribution update when the new correspondence continuously obtains sufficient matching gain.
[0255] 3. Comparison of Example 3 and Comparative Example 2 The semantic correction data of the same occupant and the vehicle function dependency conflict data were respectively input into the method of Example 3 and the method of Comparative Example 2, and the statistics were performed: Improved relationship identification accuracy; False conflict rate; Indirect conflict false negative rate; Effective instruction retention rate.
[0256] For corrected speech such as "twenty-two degrees, no, twenty-four degrees", Comparative Example 2 treats the recognition results before and after as independent commands, which is prone to false conflicts or repetitive control; Example 3 reconstructs the speech before and after into the final valid command based on the consistency of the commanding crew, the inheritance of the target object and the corrected expression.
[0257] For commands such as "turn on windshield defroster" and "turn off air conditioning" that have different target objects but share the same compressor and blower, Comparative Example 2 only compares the direct target object, which is prone to missing indirect conflicts; Example 3 determines the direct control and indirect influence nodes along the dependency relationship of vehicle function nodes, and can identify the action opposition relationship on overlapping nodes.
[0258] For voice commands that have some conflicts, Comparative Example 2 retains or discards the commands as a whole according to the fixed occupant priority; Example 3 can retain the local semantic requirements corresponding to non-conflicting vehicle function nodes, thus having a higher effective command retention capability.
[0259] V. Record Table of Measured Results This test included 240 concurrent voice commands from multiple occupants, including 120 cross-voice region voice commands and 120 voice commands indicating proximity to the sound source trajectory. Target region resolution accuracy was calculated using the 240 voice commands containing region designations such as "this side," "the one behind," and "right-side window." The results are shown in Table 6. Table 6. Measured Results of Examples and Comparative Examples 1
[0260] in: In Example 1 or Example 2, 10 out of 240 voice commands resulted in passenger assignment errors, with a passenger assignment error rate of [missing information]. ; In Comparative Example 1, 42 instances of passenger assignment errors occurred, resulting in a passenger assignment error rate of [missing information]. ; In Example 1 or Example 2, all valid voice segments of 221 instructions consistently belong to the same correct passenger, resulting in an instruction completeness rate of [missing information]. ; In Comparative Example 1, 172 instructions achieved complete attribution, resulting in an instruction complete attribution rate of [percentage missing]. ; In 120 sets of sound source trajectories that are close to speech, Example 2 shows 5 sets of non-real channel exchanges, while Comparative Example 1 shows 22 sets of non-real channel exchanges. Example 1 or Example 2 correctly parsed the target region of 227 instructions, with a target region parsing accuracy of [percentage missing]. ; Comparative Example 1 correctly parsed the target region of 189 instructions, achieving a target region parsing accuracy of %. .
[0261] The above results indicate that using a joint judgment based on the continuity of sound source trajectory, voiceprint similarity, phoneme sequence continuity, and historical attribution probability can reduce the occupant attribution error caused by sound sources crossing the sound zone boundary; and using time window cumulative attribution matching and attribution switching penalty can reduce non-real channel switching when sound source trajectories are close together.
[0262] This set of tests includes: 200 sets of semantically corrected speech samples from the same passenger; 200 sets of corrected or supplementary voice lines that do not actually constitute multi-person conflicts; 120 sets of voice messages that have different direct target objects but actually conflict in their dependence on vehicle functions; 200 effective occupant requirements that can be fully or partially met in conflict scenarios; There are 120 sets of confirmation voice messages that require passenger confirmation and may involve multiple people responding simultaneously.
[0263] The test results are as follows: Table 7 Test Results of Examples and Comparative Examples 2
[0264] in: Example 3 correctly identified 187 sets of semantic correction relations, with a correction relation identification accuracy of [percentage missing]. ; Comparative Example 2 yielded 122 groups of samples that correctly obtained the final corrected results, with the accuracy rate of corrected relationship identification being [percentage missing]. Its partially correct results mainly come from simple subsequent instruction overriding, rather than stable recognition of semantic modification relationships; Example 3 identifies 11 groups of voice errors that were actually corrected or supplemented by the same passenger as conflicts, with a false conflict rate of [missing information]. ; Comparative Example 2: 62 groups of incorrect judgments, false conflict rate ; In 120 groups of action-dependent conflict samples, Example 3 missed 9 groups, with an indirect conflict false negative rate of [missing information]. ; Comparative Example 2 had 46 missed detections, with an indirect conflict missed detection rate of 100%. ; Example 3 retains 178 out of 200 satisfiable valid requirements. These requirements include complete semantic instructions and local semantic instructions obtained after scope decomposition. The effective instruction retention rate is [percentage missing]. ; Comparative Example 2 retained 127 valid requirements, with a valid instruction retention rate of ; In Example 3, out of 120 concurrent multi-user confirmation voice messages, 113 confirmation voice messages were correctly bound to the confirmation command and the corresponding issuing passenger, achieving a confirmation voice binding accuracy rate of [percentage missing]. ; Comparative Example 2 correctly bound 92 groups; the accuracy rate of voice binding to be confirmed is... .
[0265] The above results show that, by identifying the semantic correction relationship of the same occupant through the consistency of the commanding occupant, the inheritance of the target object, and the correction expression, Example 3 can reduce the false conflict rate; by determining the indirect influence nodes of the voice command through the dependency relationship of the vehicle function nodes, the indirect conflict false detection rate can be reduced; and by retaining the local semantic commands corresponding to the non-conflicting vehicle function nodes, the effective command retention rate can be improved.
Claims
1. A method for resolving conflicts in vehicle-mounted sound source localization for concurrent commands from multiple occupants, characterized in that, include: S1. Acquire multi-channel speech signals collected by the in-vehicle microphone array, separate overlapping speech intervals to obtain multiple candidate speech streams, and generate corresponding sound source trajectories based on the arrival time difference between the channels of each candidate speech stream. S2. Extract the voiceprint features and phoneme sequence features of each candidate speech stream. Based on the spatial continuity, voiceprint similarity and phoneme sequence continuity of the sound source trajectory, determine the attribution relationship between each candidate speech stream and the occupants in the vehicle, and maintain the corresponding attribution relationship when the sound source trajectory crosses the boundary of a preset sound zone. S3. When the distance between at least two sound source trajectories is less than the preset distance, based on the continuity of sound source positions and voiceprint similarity within the preset time window, compare the cumulative matching results when maintaining the original attribution relationship and when exchanging the attribution relationship, and penalize the switching of the attribution relationship to determine the passenger attribution relationship of each candidate speech stream after the sound source trajectories are close. S4. After determining the affiliation of the crew members, perform speech recognition and semantic parsing on each candidate speech stream to obtain standardized speech commands that include the commanding crew member, target object, target area, and control action. S5. Based on whether the issuing crew members of the standardized voice commands that are adjacent in time are the same, whether the subsequent voice command inherits the target object of the preceding voice command, and whether it contains a correction expression, identify the correction relationship of the same crew member to the preceding voice command, and update the corresponding standardized voice command accordingly. S6. For multiple standardized voice commands that do not belong to the aforementioned correction relationship, determine the vehicle functional nodes that are directly controlled and indirectly affected by each standardized voice command based on the role dependency relationship between vehicle functional nodes. S7. Based on the overlapping relationship between the vehicle function nodes corresponding to different standardized voice commands and the opposing relationship of control actions on the overlapping vehicle function nodes, determine the command conflict, perform semantic reconstruction on the standardized voice commands with conflict, and output the semantic command after conflict resolution.
2. The method according to claim 1, characterized in that, The speech separation of overlapping speech regions includes: The multi-channel speech signal is framed and time-frequency transformed, and a spatial covariance matrix is constructed based on the multi-channel time-frequency signal of each speech frame; Based on the eigenvalue distribution of the spatial covariance matrix, determine the proportion of the sum of all eigenvalues except the largest eigenvalue in the sum of all eigenvalues; When the percentage is continuously greater than the preset overlap threshold, the corresponding speech frame is determined as the overlapping speech interval, and the number of active sound sources is determined according to the spatial covariance matrix. The multiple candidate speech streams are then separated according to the number of active sound sources.
3. The method according to claim 1, characterized in that, The generation of the corresponding sound source trajectory includes: Calculate the generalized cross-correlation phase transformation results of each candidate speech stream across different microphone channels, and determine the arrival time difference between channels based on the peak position of the generalized cross-correlation phase transformation results; By combining the positions of each microphone in the microphone array and the arrival time difference between the channels, the sound source positions of each candidate speech stream at different times are determined; The changes between the sound source position and the sound source position at adjacent times are used as position state and velocity state, respectively. The sound source state of each candidate speech stream is recursively calculated, and the recursed sound source positions are connected in chronological order to form the sound source trajectory.
4. The method according to claim 1, characterized in that, Determining the attribution relationship between each candidate voice stream and the vehicle occupants includes: For the q-th candidate speech stream and the i-th passenger, the attribution matching degree is determined according to the following formula: ; in, To ensure spatial continuity based on the current sound source location and the predicted sound source location of the i-th occupant, This represents the voiceprint similarity between the current voiceprint feature and the historical voiceprint feature of the i-th vehicle occupant. To determine the phoneme continuity between the current phoneme sequence features and the historical phoneme sequence features of the i-th occupant, Let w be the crew affiliation probability at the previous time step. p w e w c w h These are the weights of the corresponding parameters, and the sum of the weights is 1; The passenger affiliation probability of the candidate voice stream is updated according to the affiliation matching degree of each passenger in the vehicle, and the passenger in the vehicle with the highest passenger affiliation probability is determined as the starting passenger of the corresponding candidate voice stream.
5. The method according to claim 4, characterized in that, When the distance between at least two sound source trajectories is less than a preset distance, the target correspondence between each candidate speech stream and the vehicle occupant is determined according to the following formula: ; In the formula, π represents the candidate correspondence between candidate speech streams and in-vehicle occupants, H represents the number of speech frames included in the preset time window, and A q,π(q) (nh) represents the attribution matching degree between the q-th candidate speech stream and the vehicle occupant in the corresponding speech frame, where N is the attribution matching degree between the q-th candidate speech stream and the vehicle occupant. SW (π) represents the number of attribution relationship changes caused by the candidate correspondence within a preset time window, λ sw To switch the penalty coefficient; The candidate correspondence that maximizes the difference between the cumulative attribution matching degree and the attribution relationship switching penalty is determined as the target correspondence.
6. The method according to claim 1, characterized in that, The standardized voice commands also include action parameters and command confidence levels; The confidence level of the instruction is determined jointly based on the speech recognition confidence level, occupant attribution confidence level, semantic integrity, and overlapping speech separation quality of the candidate speech stream; When the confidence level of the instruction is lower than the preset instruction threshold, the corresponding standardized voice instruction is marked as an instruction to be confirmed, and the instruction to be confirmed is not directly used to output the conflict resolution result with other standardized voice instructions.
7. The method according to claim 1, characterized in that, The identification of the same passenger's correction relationship to preceding voice commands includes: Within a preset correction time window, select preceding and following voice commands that are adjacent in time, and determine the probability of the same commanding crew member based on the crew member affiliation probability distribution of the two commands. Determine whether the subsequent voice instruction omits and inherits the target object of the preceding voice instruction, or whether it contains an undo expression, parameter replacement expression, target object replacement expression, or area of effect adjustment expression; When the probability of the same commanding crew member is greater than the preset threshold for the same person, and the subsequent voice command satisfies the target object inheritance condition or includes the cancellation expression, parameter replacement expression, target object replacement expression or area of action adjustment expression, it is determined that the subsequent voice command and the preceding voice command constitute the correction relationship. Based on the correction relationship, the preceding voice command is deleted, the original action parameters are overwritten with new action parameters, the target object is replaced, or the target area is adjusted.
8. The method according to claim 1, characterized in that, The determination of the vehicle function nodes directly controlled and indirectly affected by each standardized voice command includes: Construct a vehicle function node dependency matrix H, where the elements in the vehicle function node dependency matrix are used to represent the degree of influence of the state change of one vehicle function node on another vehicle function node. Generate a direct target vector u based on the target object and target region of the j-th standardized speech instruction. j The corresponding control scope vector is determined according to the following formula: ; Among them, s j Let R be the control scope vector, R be the maximum propagation order of the action dependency, μ be the propagation attenuation coefficient, and clip(·,0,1) represent restricting each element in the vector to between 0 and 1. Vehicle function nodes in the control domain vector that are greater than a preset threshold are identified as vehicle function nodes that are directly controlled or indirectly affected by the corresponding standardized voice commands.
9. The method according to claim 8, characterized in that, The determination of instruction conflicts includes: Based on the control domain vectors corresponding to different standardized voice commands, the overlapping vehicle function nodes between two standardized voice commands are determined. When two standardized voice commands have opposite control actions on the same overlapping vehicle function node, it is determined that the two constitute a direct reverse conflict. When the set of vehicle function nodes corresponding to one standardized voice command includes the set of vehicle function nodes corresponding to another standardized voice command, and the control actions of the two on the included vehicle function nodes are opposite, it is determined that the two constitute a conflict in the scope of action. When two standardized voice commands have different direct target objects, but have the same indirect influence on the vehicle function node after being propagated through the vehicle function node action dependency matrix, and the control actions on the indirectly influenced vehicle function node are in opposite directions, it is determined that the two constitute an action dependency conflict.
10. The method according to claim 9, characterized in that, The semantic reconstruction of conflicting standardized voice commands includes: The vehicle function nodes corresponding to different standardized voice commands are divided into non-overlapping vehicle function nodes and overlapping vehicle function nodes, and local semantic commands that can be output in parallel are generated for the non-overlapping vehicle function nodes. If the scope of action contains conflicts, remove vehicle function nodes that conflict with standardized voice commands with smaller scopes from the standardized voice commands with larger scopes, and retain the local semantic commands corresponding to the remaining vehicle function nodes. For the direct reverse conflict or action-dependent conflict, a parameter-restricted semantic command or a semantic command to be confirmed is generated based on the occupant affiliation confidence and command confidence of the corresponding standardized voice command. The corresponding commanding crew member identifier is bound to the semantic instruction to be confirmed, so that the subsequent confirmed voice is used to update the semantic instruction to be confirmed only when its crew member affiliation matches the commanding crew member identifier.
Citation Information
Patent Citations
Multi-mode graphic image processing equipment
CN120355560A
Intelligent cabin multi-mode voice interaction system and method
CN121034318A