COMPUTER-IMPLEMENTED METHOD FOR DETECTING ENGAGEMENT ZONES BY COHERENT FOCUSING SUMMATION OVER MULTIPLE GEOMETRIC POSITIONS
Patent Information
- Application Number
- DE102024112703
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2044-05-06
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
INTRODUCTION
[0001] The present invention relates generally to a computer-implemented method for detecting zones of engagement by coherent focus summation across multiple geometric positions. More specifically, the manner in which a user interacts with a user interface of a vehicle system is shaped primarily, if not exclusively, by voice input. For example, a user may request the vehicle to perform an action that includes playing media (e.g., music or podcasts), after which the user interface responds by playing audio data that matches the user's criteria. In cases where multiple microphones are recording multiple users (e.g., a driver and a passenger) speaking in the vehicle, the vehicle may need to recognize which user spoke a desired action.
[0002] US 2015 / 0 156 587 A1 discloses a method in which a multi-channel input signal is received and speech activity is detected based on a correlation of the input signals.
[0003] CN 1 10 412 499 A discloses a method in which a cross-correlation matrix is generated from frequency subbands and a focusing matrix is calculated using the cross-correlation matrix.
[0004] From DE 10 2016 212 647 A1 it is known to detect a multi-channel input signal in a plurality of zones in a vehicle and to perform keyword recognition for each zone. SUMMARY
[0005] According to the invention, a computer-implemented method for onset zone detection using coherent focus summation over multiple geometric positions is presented, which is characterized by the features of claim 1.
[0006] When executed on computing hardware, the method causes the computing hardware to perform operations comprising receiving a multi-channel input signal containing a sequence of frames captured in an environment of a vehicle, wherein the environment of the vehicle has at least two zones.For each zone of the at least two zones in the vicinity of the vehicle, the operations also include performing speech recognition by: converting each frame in the sequence of frames of the multi-channel input signal into a plurality of frequency subbands, each frequency subband including a corresponding cross-correlation matrix (CCM), and for each respective frequency subband of the plurality of frequency subbands: applying a focusing matrix to the respective CCM to generate a corrected CCM, extracting eigenvalues from the corrected CCM, and determining an eigenvalue ratio between a highest eigenvalue extracted from the corrected CCM and a second highest eigenvalue extracted from the corrected CCM. For each zone of the at least two zones, the operations also include calculating a median of the eigenvalue ratios of the plurality of frequency subbands for each frame in the sequence of frames.The operations further include: determining a difference between the respective median values of the at least two zones of the vehicle's environment, and generating an initial speech detection indication when an absolute value of the difference between the respective median values of the at least two zones of the vehicle's environment is greater than a threshold.
[0007] Implementations of the invention may include one or more of the following optional features. In some implementations, the operations further include converting the multi-channel input signal into the sequence of frames. In some examples, the focusing matrix is initialized using a steering vector unique to a model of the vehicle.
[0008] In some implementations, the operations further comprise, for each zone of the at least two zones of the environment of the vehicle, confirming the presence of speech in each frame in the sequence of frames by: projecting the multi-channel input signal onto a steering vector of the vehicle to generate a projection, determining an average energy of the plurality of frequency sub-bands, and confirming the presence of speech in the multi-channel input signal if the average energy exceeds a directionality threshold.In these implementations, the operations may further include determining a difference between the respective projections of the at least two zones, and if the difference between the respective projections exceeds a dominance threshold, generating a speech detection indication confirmation identifying a zone of the at least two zones as the source of the speech in the multi-channel input signal. Identifying the zone of the at least two zones as the source of the speech in the multi-channel input signal may be based on the initial speech detection indication and the speech detection indication confirmation. Optionally, the steering vector is vehicle-specific.
[0009] In some examples, the plurality of frequency subbands are in the frequency domain. In some implementations, the at least two zones include a first and a second zone. In some examples, speech recognition is performed without historical audio data.
[0010] Further described is a system for detecting engagement zones using coherent focus summation across multiple geometric positions, comprising computing hardware and memory hardware in communication with the computing hardware. The memory hardware stores instructions that, when executed by the computing hardware, cause the computing hardware to perform operations including: receiving a multi-channel input signal containing a sequence of frames captured in an environment of a vehicle, wherein the environment of the vehicle has at least two zones.For each zone of the at least two zones in the vicinity of the vehicle, the operations also include performing speech recognition by: converting each frame in the sequence of frames of the multi-channel input signal into a plurality of frequency subbands, each frequency subband including a corresponding cross-correlation matrix (CCM), and for each respective frequency subband of the plurality of frequency subbands: applying a focusing matrix to the respective CCM to generate a corrected CCM, extracting eigenvalues from the corrected CCM, and determining an eigenvalue ratio between a highest eigenvalue extracted from the corrected CCM and a second highest eigenvalue extracted from the corrected CCM. For each zone of the at least two zones, the operations also include calculating a median of the eigenvalue ratios of the plurality of frequency subbands for each frame in the sequence of frames.The operations further include: determining a difference between the respective median values of the at least two zones of the vehicle's environment, and generating an initial speech detection indication when an absolute value of the difference between the respective median values of the at least two zones of the vehicle's environment is greater than a threshold.
[0011] This aspect may include one or more of the following optional features. In some implementations, the operations further include converting the multi-channel input signal into a sequence of frames. In some examples, the focusing matrix is initialized using a steering vector unique to a model of the vehicle.
[0012] In some implementations, the operations further include, for each zone of the at least two zones of the environment of the vehicle, confirming the presence of speech in each frame of the sequence of frames by: projecting the multi-channel input signal onto a steering vector of the vehicle to generate a projection, determining an average energy of the plurality of frequency sub-bands, and confirming the presence of speech in the multi-channel input signal if the average energy exceeds a directionality threshold.In these implementations, the operations may further include determining a difference between the respective projections of the at least two zones, and if the difference between the respective projections exceeds a dominance threshold, generating a speech detection indication confirmation identifying a zone of the at least two zones as the source of the speech in the multi-channel input signal. Identifying the zone of the at least two zones as the source of the speech in the multi-channel input signal may be based on the initial speech detection indication and the speech detection indication confirmation. Optionally, the steering vector is vehicle-specific.
[0013] In some examples, the plurality of frequency subbands are in the frequency domain. In some implementations, the at least two zones include a first and a second zone. In some examples, speech recognition is performed without historical audio data.
[0014] The details of one or more implementations of the invention are set forth in the accompanying drawings and the description below. Further aspects, features, and advantages will become apparent from the description and drawings, as well as from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described here are for illustrative purposes only and are intended to illustrate selected configurations. Fig. Figure 1 is a schematic diagram of an exemplary system for detecting operational zones. Fig. Figure 2 is a schematic view of exemplary components of the system of Fig. 1. Fig. Figure 3 is a schematic representation of a frequency space. Fig. 4 is a flowchart of an exemplary arrangement of operations for a method for generating an initial zone prediction for a speaker. Fig. 5 is a flowchart of an exemplary arrangement of operations for a method for generating a final zone prediction for a speaker. Fig. 6 is a flowchart of an exemplary arrangement of operations for a method for detecting operational zones.
[0016] Corresponding reference numbers indicate corresponding parts in the drawings. DETAILED DESCRIPTION
[0017] As in Fig. 1, in some embodiments, a system 100 includes a vehicle 10 and / or a remote system 60 that communicates with the vehicle 10 over a network 40. The vehicle 10 detects speech utterances 18 from one or more users (i.e., a driver and / or one or more passengers) in an environment 26 of the vehicle 10 and processes the speech utterances 18 to detect a zone 30 of the vehicle 10 in which the speaker of the speech utterance 18 is located. As described in more detail below, by detecting the zone 30 of the speaker of the utterance 18, the vehicle 10 can more accurately distinguish between speech from a driver, a front passenger, and a rear seat passenger of the vehicle 10. A user can speak the utterance 18 as a question or command to elicit a response from the vehicle 10. The vehicle 10 is configured to detect sounds from one or more users in the environment 26.In this case, the sounds may refer to a spoken utterance 18 of the user that functions as an audible request, a command to the vehicle 10, or an audible communication detected by the vehicle 10. Voice-enabled systems of the vehicle 10 or systems connected to the vehicle 10 may respond to the request for the command and / or initiate the execution of the command.
[0018] The vehicle 10 and / or the remote system 60 execute a zone of engagement detection system 200 that detects a speaker of utterance 18 in only a single frame 24. In other words, unlike conventional directional voice activity detectors (DVADs), which require historical audio data to make a decision about a current audio frame, the zone of engagement detection system 200 detects a zone 30 of a speaker without historical audio data and works well for short utterances 18 (e.g., utterances less than 200 milliseconds in length) that can be used in downstream speech processing, and generally for utterances 18 of any length that can be recorded in the vehicle 10.The deployment zone detection system 200 is configured to receive as input a multi-channel input signal 22 comprising a plurality of frames 24 captured in the environment 26 of the vehicle 10. As shown in FIG. Fig. 1, the environment 26 of the vehicle 10 generally includes the interior of the vehicle 10, in which a microphone assembly 16 is disposed within a headliner of the interior of the vehicle 10 and is located in a front portion of the vehicle 10 between a driver area and a passenger area of the vehicle 10. The environment 26 may generally be divided into two or more zones 30, with each zone 30 corresponding to a user position within the vehicle 10. As illustrated, the vehicle 10 may include four (4) zones 30, 30a-30d, with zone 30a corresponding to a driver seat, zone 30b corresponding to a front passenger seat, and zones 30c, 30d corresponding to the rear passenger seats on the left and right sides of the vehicle 10. While the examples used generally refer to the two zones 30a, 30b, the system 200 for detecting operational zones can also detect more than two zones 30a, 30b, e.g.three (3) zones 30a - 30c or any other combination of zones 30.
[0019] In the examples shown, the deployment zone detection system 200 is implemented within the vehicle 10. However, the deployment zone detection system 200 may also be implemented on other computing devices (e.g., computing devices that communicate with the vehicle 10), such as a smartphone, a tablet, a smart display, a desktop / laptop, a smart watch, a smart device, or smart glasses / headset. The vehicle 10 includes computing hardware 12 and memory hardware 14 that stores instructions that, when executed on the computing hardware 12, cause the computing hardware 12 to perform operations. As shown, the vehicle 10 is in communication with the remote system 60 via the network 40. The remote system 60 (e.g.,The system (e.g., server, cloud computing environment) also includes computing hardware 62 and storage hardware 64 storing instructions that, when executed on the computing hardware 62, cause the computing hardware 62 to perform operations. In some examples, the zone of operation detection system 200 is executed jointly by the vehicle 10 and the remote system 60.
[0020] The vehicle 10 further includes an audio subsystem 20 having the microphone array 16 for capturing and converting spoken utterances 18 in the environment 26 of the vehicle 10 into electrical signals (or is in communication therewith). Each microphone 16 in the array of microphones 16 of the vehicle 10 can record the utterance 18 on a separate dedicated channel of the multi-channel input signal 22. For example, the vehicle 10 can include two microphones 16 (also referred to as a microphone array 16), each recording the utterance 18, and the recordings of the two microphones 16 can be combined into a two-channel input signal 22 (i.e., stereophonic sound or stereo). However, the microphone array 16 can include any number of microphones 16. While the vehicle 10 in the example of Fig. 1 includes the microphone assembly 16, other examples may include additional configurations at any location in the vehicle 10, such as two (2) front microphone assemblies, four (4) microphone assemblies, etc.
[0021] The audio subsystem 20 is configured to receive the spoken utterance 18 captured by the microphone assembly 16 and to convert the utterance 18 into a corresponding digital format associated with acoustic frames 24 that can be processed by the system 200 for detecting operational zones. In the Fig. 1, the audio subsystem 20 converts the utterance 18 into a multi-channel input signal 22 containing a sequence of acoustic frames (e.g., audio data) 24, which are input to a zone detection model 202 (also referred to as model 202) of the engagement zone detection system 200. The model 202 then generates / predicts as output an initial speech detection indication 412. The initial speech detection indication 412 indicates whether a particular frame 24 contains speech, and if so, from which of the two zones 30a, 30b the speech from the particular frame 24 originates.The model 202 then receives as input a steering vector 252 of the vehicle 10 and confirms, for each zone 30a, 30b in the environment 26 of the vehicle 10, the presence of speech in each frame 24 in the sequence of frames 24 to generate as output a speech detection indication confirmation 522 indicating which zone 30a, 30b is a source of the speech in the multi-channel input signal 22. The zone detection model 202 can identify one of the zones 30a, 30b as the source of the speech in the multi-channel input signal 22 based on the initial speech detection indication 412 and the speech detection indication confirmation 522.
[0022] In relation to Fig. 1 and Fig. 2, the zone detection model 202 of the deployment zone detection system 200 includes a matrix generator 210, a matrix corrector 220, an initial zone detector model 400, and a final zone detector model 500. The deployment zone detection system 200 may access a steering vector 252 stored in a vehicle data store 250 located on the memory hardware 14 of the vehicle 10 and / or the memory hardware 64 of the remote system 60. The steering vector 252 may be unique to the vehicle 10 (e.g., the model of the vehicle 10) and is tuned offline. The steering vector 252 may be based on delays or a relative transform function (RTF). In some embodiments, the steering vector 252 includes multiple steering vectors that approximate the same zones 30.
[0023] The matrix generator 210 is configured to receive as input the multi-channel input signal 22 with the plurality of frames 24 and, for each frame 24, convert each frame 24 into a plurality of frequency subbands 212. Each frequency subband 212 includes a corresponding cross-correlation matrix (CCM) 214. The plurality of frequency subbands 212 may be in the frequency domain. Fig. 3 depicts a frequency space 300, with frames 24 on the x-axis and frequency subbands 212 on the y-axis. Unlike conventional frame classification, which evaluates a combination of frequency subbands over time, as indicated by selection 310, to detect a speaker, matrix generator 210 divides each frame 24 into the plurality of frequency subbands 212, as indicated by selection 320, to detect a speaker.
[0024] Referring again to Fig. 2, the matrix corrector 220 is configured to receive as input the plurality of frequency subbands 212 and the respective CCMs 214 output from the matrix generator 210, as well as the steering vector 252, and to generate a corrected CCM 222 for each respective frequency subband 212 of the plurality of frequency subbands 212. In particular, the matrix corrector applies a focusing matrix T for each frequency subband 212 and each zone 30 to each corresponding CCM 214. A respective focusing matrix T can be initialized for each frequency subband 212 and each zone 30 in the environment 26 of the vehicle 10, with each focusing matrix T being initialized from the steering vector A 252. The focusing matrix T can be defined as follows: T(k,d)=V(k,d)U*(k,d) where k denotes an index of each frequency interval, d denotes two potential directions (d ∈ [1,2]) (e.g. the zones 30a, 30b of the environment 26 of the vehicle 10), U the left singular vector of a singular value decomposition of C A and V the right singular vector of the singular value decomposition of C A and C A is defined as follows: CA=A(k,d)A*(k0,d) where k0 denotes the center frequency of subband 212.
[0025] For each zone 30 and for each respective frequency subband 212 of the plurality of frequency subbands 212, the matrix corrector 220 may apply a respective focusing matrix T to the respective CCM 214 to generate the corrected CCM 222. In other words, for the first zone 30a and for each respective frequency subband 212, the matrix corrector 220 applies a respective focusing matrix T to the respective CCM 214 to generate the corrected CCM 222, and for the second zone 30b and for each respective frequency subband 212, the matrix corrector 220 applies a respective focusing matrix T to the respective CCM 214 to generate the corrected CCM 222. For example, each CCM 214 is corrected using its respective focusing matrix T by: Rk0,d=1K∑k=kLkHTk,dRk,dTk,dH where R k0,d the corrected CCM 222, [k L , k H ] denotes the range of frequency averaging, K = k H- k L + 1, and k0 denotes the center frequency of the subband 212. Instead of selecting a single center interval for an entire range, the frame 24 is divided into frequency subbands 212, and a center interval is assigned to a frequency subband 212 to reduce the correction error due to the large frequency differences. In other words, the matrix corrector 220 generates an R k0,d (ie a corrected CCM 222) for each frequency subband 212.
[0026] As in Fig. 4, the initial zone detector model 400 is configured to receive the corrected CCMs 222 for each zone 30a, 30b (also referred to as Zone 1 and Zone 2) and for each respective frequency subband 212 of the plurality of frequency subbands 212 and generate the initial speech detection indication 412. In the example shown, the initial zone detector model 400 receives the corrected CCMs 222a1 - 222n1 for a first zone 30a and the corrected CCMs 222a2 - 222n2 for the second zone 30b. Thereafter, the initial zone detector model 400 extracts sorted eigenvalues 4021, 4022 from each corrected CCM 222 for each zone 30a, 30b.In particular, the initial zone detector model 400 for zone 30a extracts the highest eigenvalues 4021a1, 4021b1, 4021n1 from the respective corrected CCMs 222a1, 222b1, 222n1 and the second highest eigenvalues 4021a2, 4021b2, 4021n2 and determines a respective eigenvalue ratio 4041a, 4041b, 4041n for each of the respective corrected CCMs 222a1, 222b1, 222n1. Likewise, for zone 30b, the initial zone detector model 400 extracts the highest eigenvalues 4022a1, 4022b1, 4022n1 from the respective corrected CCMs 222a2, 222b2, 222n2 and the second highest eigenvalues 4022a2, 4022b2, 4022n2 and determines a respective eigenvalue ratio 4042a, 4042b, 4042n for each of the respective corrected CCMs 222a2, 222b2, 222n2. The respective eigenvalue ratio 404 is expressed as follows:. ρk0,d=λ1λ2 where λ1 denotes the highest eigenvalue 402 extracted from the corrected CCM 222, and λ2 denotes the second highest eigenvalue 402 extracted from the corrected CCM 222. Note that the respective eigenvalue ratio 404 indicates the rank of the corrected CCM 222 directly linked to the source (point or omni) of utterance 18.
[0027] The initial zone detector model 400 calculates for each zone 30a, 30b and for each frame 24 a median value 406 of the eigenvalue ratios 404 of the plurality of frequency subbands 212. In other words: The median value ρd˜ 406 for a given direction d (ie zones 30a, 30b) is defined as: ρd˜=med(ρk0,d)
[0028] As in Fig. For example, as shown in Figure 4, the initial zone detector model 400 calculates a median value ρ1˜ 406i, which corresponds to the first zone 30a, and a median value ρ2˜ 406ii, which corresponds to the second zone 30b, and calculates a difference 408 between the median value ρ1˜ 406i of the first zone 30a and the median value ρ2˜ 406ii of the second zone 30b. For example, the median value ρ1˜ 406i of the first zone 30a from the median value ρ2˜ 406ii of the second zone 30b. If an absolute value of the difference 408 between the median value ρ1˜ 406i of the first zone 30a and the median value ρ2˜ 406ii of the second zone 30b is less than an initial threshold, then the initial zone detector model 400 may generate an initial speech detection indication 412 indicating that the frame 24 does not contain speech. Conversely, if the absolute value of the difference 408 between the median value ρ1˜ 406i of the first zone 30a and the median value ρ2˜ 406ii of the second zone 30b is greater than the initial threshold, the initial zone detector model 400 generates an initial speech detection indication 412 indicating that the frame 24 contains speech. In this case, the initial threshold may be set during tuning and may be unique to the model of the vehicle 10.
[0029] If the initial zone detector model 400 generates an initial speech detection indication 412 indicating that frame 24 contains speech, the zone detector model 202 performs Fig. 2, the final zone detector model 500 executes to confirm or reject the initial speech detection indication 412 output by the initial zone detector model 400. In other words, for each zone 30a, 30b of the environment 26 of the vehicle 10, the final zone detector model 500 confirms or rejects the presence of speech in each frame 24 in the sequence of frames 24. The final zone detector model 500 is configured to receive the multi-channel input signal 22 and the steering vector 252 and generate a speech detection indication confirmation 522 indicating which zone 30a, 30b is the source of the speech in the multi-channel input signal 22. The final zone detector model 500 may include a subband directionality threshold, a frame directionality threshold, and a dominance threshold, each of which is set / selected during tuning for the respective model of the vehicle 10.
[0030] As in Fig. 5, at operation 510, the final zone detector model 500 receives for each zone 30a, 30b the multi-channel input signal 22 and the respective steering vector 252 and projects the multi-channel input signal 22 onto the steering vector 252 to generate a respective projection P for each zone 30a, 30b.
[0031] The projection P is expressed as follows: P=hHxxHhhHhSpur(xxH)−hHxxHhhHh where h denotes the respective steering vector 252 and x denotes the multi-channel input signal 22 in the frequency domain. As shown, the final zone detector model 500 calculates a projection P i of the multi-channel input signal 22 for the first zone 30a and a projection P ii of the multi-channel input signal 22 for the second zone 30b.
[0032] At operation 514, the final zone detector model 500 determines whether the respective projection P for each zone 30a, 30b is greater than a subband directionality threshold. If the respective projection P for each zone 30a, 30b is greater than the subband directionality threshold, the final zone detector model 500 determines an average energy of the plurality of frequency subbands 212 and, at operation 516, determines for each zone 30a, 30b whether the average energy is greater than a frame directionality threshold.If the final zone detector model 500 determines that the average energy of a plurality of subbands 212 for a particular zone 30 is not greater than the frame directionality threshold, the final zone detector model 500 rejects the presence of speech in the multi-channel input signal 22 for the particular zone 30 and generates a speech detection indication acknowledgment 522 indicating that the particular zone 30 does not contain speech for the current frame 24. Conversely, if the final zone detector model 500 determines that the average energy of a plurality of subbands 212 for a particular zone 30 is greater than the frame directionality threshold, the final zone detector model 500 may confirm that speech is present in frame 24 and proceed to operation 518 to identify which zone 30 the speech originated from.
[0033] Here, the final zone detector model 500 may calculate a difference between the projections P of the individual zones 30a, 30b. For example, at operation 520, the final zone detector model 500 subtracts the projection P ii for the second zone 30b from the projection P i the first zone 30a, and, if the difference is greater than a dominance threshold, generates the speech detection indication acknowledgment 522 that identifies the first zone 30a as the source of the speech in the multi-channel input signal 22. Conversely, if the difference is less than the dominance threshold, the final zone detector model 500 proceeds to operation 524, where the final zone detector model 500 generates the projection P i for the first zone 30a from projection P iithe second zone 30b and, if the difference is greater than the dominance threshold, generates the speech detection indication confirmation 522 that identifies the second zone 30b as the source of the speech in the multi-channel input signal 22. Conversely, the final zone detector model 500 generates the speech detection indication confirmation 522 that indicates that the zones 30a, 30b do not contain speech for the current frame 24 if, at operation 524, the difference is less than the dominance threshold.
[0034] Fig. Figure 6 shows a flowchart of an exemplary arrangement of operations for a method 600 for detecting engagement zones using coherent focus summation across multiple geometric positions. The method 600 may be described with reference to Fig. 1 - 5. The data processing hardware (e.g. data processing hardware 12, 62 of Fig. 1) can execute instructions that are stored on the memory hardware (e.g. memory hardware 14, 64 of Fig. 1) are stored to perform the exemplary arrangement of operations for method 600.
[0035] The method 600 includes, in operation 602, receiving a multi-channel input signal 22 having a sequence of frames 24 recorded in an environment 26 of a vehicle 10. Here, the environment 26 includes at least two zones 30. For each zone 30 of the at least two zones 30 of the environment 26 of the vehicle 10, the method 600 includes operations 604-606. In operation 604, the method 600 includes converting each frame 24 in the sequence of frames 24 of the multi-channel input signal 22 into a plurality of frequency subbands 212. Each subband 212 includes a corresponding cross-correlation matrix (CCM) 214.For each respective frequency subband 212 of the plurality of frequency subbands 212, the method 600 includes, in operation 606, applying a focusing matrix T to the respective CCM 214 to generate a corrected CCM 222, extracting eigenvalues 402 from the corrected CCM 222, and determining an eigenvalue ratio 404 between a highest eigenvalue 4021 extracted from the corrected CCM 222 and a second highest eigenvalue 4022 extracted from the corrected CCM 222. In operation 608, the method 600 also includes calculating a median value 406 of the eigenvalue ratios 404 of the plurality of frequency subbands 212 for each frame 24 in the frame sequence.
[0036] At operation 610, the method 600 also includes determining a difference 408 between the respective median values 406 of the at least two zones 30 of the environment 26 of the vehicle 10. If an absolute value of the difference 408 between the respective median values 406 of the at least two zones 30 of the environment 26 of the vehicle 10 is greater than a threshold, the method 600 also includes generating an initial speech detection indication 412 at operation 612.
Claims
[1] A computer-implemented method which, when executed in data processing hardware (12, 62), causes the data processing hardware (12, 62) to perform operations comprising: Receiving a multi-channel input signal (22) containing a sequence of frames (24) recorded in an environment of a vehicle (10), the environment of the vehicle (10) having at least two zones (30); Performing speech recognition for each zone (30) of the at least two zones (30) in the vicinity of the vehicle (10) by: Converting each frame (24) in the sequence of frames (24) of the multi-channel input signal (22) into a plurality of frequency subbands (212), each frequency subband (212) comprising a corresponding cross-correlation matrix (CCM) (214); for each respective frequency subband (212) of the plurality of frequency subbands (212): applying a focusing matrix (T) to the respective CCM (214) to produce a corrected CCM (222); Extracting eigenvalues (402) from the corrected CCM (222); and Determining an eigenvalue ratio (404) between a highest eigenvalue extracted from the corrected CCM (222) and a second highest eigenvalue extracted from the corrected CCM (222); Calculating a median value (406) of the eigenvalue ratios (404) of the plurality of frequency subbands (212) for each frame (24) in the sequence of frames (24); Determining a difference (408) between the respective median values (406) of the at least two zones (30) of the surroundings of the vehicle (10); and Generating an initial speech detection indication when an absolute value of the difference (408) between the respective median values (406) of the at least two zones (30) of the environment of the vehicle (10) is greater than a threshold value. [2] The method of claim 1, wherein the operations further comprise converting the multi-channel input signal (22) into the sequence of frames (24). [3] Method according to claim 1, wherein the focusing matrix (T) is initialized from a steering vector (252) unique to a model of the vehicle (10). [4] The method of claim 1, wherein the operations for each zone (30) of the at least two zones (30) of the environment of the vehicle (10) further comprise confirming the presence of speech in the individual frames (24) in the sequence of frames (24) by: Projecting the multi-channel input signal (22) onto a steering vector (252) of the vehicle (10) to generate a projection (P); Determining an average energy of the plurality of frequency subbands (212); and Confirming the presence of speech in the multi-channel input signal (22) when the average energy exceeds a directionality threshold. [5] The method of claim 4, wherein the operations further comprise: Determining a difference between the respective projections (P) of the at least two zones (30); and if the difference between the respective projections (P) exceeds a dominance threshold: generating a speech detection indication acknowledgment identifying one zone (30) of the at least two zones (30) as the source of the speech in the multi-channel input signal (22). [6] The method of claim 5, wherein identifying the zone (30) of the at least two zones (30) as the source of speech in the multi-channel input signal (22) is based on the initial speech detection indication and the speech detection indication confirmation. [7] The method of claim 4, wherein the steering vector (252) is unique to the vehicle (10). [8] The method of claim 1, wherein the plurality of frequency subbands (212) are in the frequency domain. [9] The method of claim 1, wherein the at least two zones (30) comprise a first zone and a second zone. [10] The method of claim 1, wherein the speech recognition is performed without historical audio data.
Citation Information
Patent Citations
Compressed sensing theory based broadband DOA estimation algorithm of RSS algorithm
CN110412499A
Method for operating a voice control system in an indoor space and voice control system
DE102016212647A1
Wind Noise Detection For In-Car Communication Systems With Multiple Acoustic Zones
US20150156587A1
CN000110412499A