Smart glasses speech recognition method and system based on multi-channel speech enhancement
Through multi-channel speech enhancement technology, using microphone arrays and neural network models, the problem of speech recognition in multi-person conferences was solved, and the clarity of the target speech and the user experience were improved.
Patent Information
- Application Number
- CN202511030544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
In multi-person conference scenarios, existing technologies have difficulty accurately identifying the speech content of each speaker, resulting in unclear speech enhancement direction and affecting the efficiency of meeting recording and content analysis.
Smart glasses based on multi-channel speech enhancement are used to obtain mixed signals through a microphone array, separate speech signals using time delay and phase compensation, and predict the direction of the target sound source for speech enhancement by combining context information and a neural network model.
Improves the accuracy of speech recognition and user experience in multi-person meetings, ensuring the accuracy and efficiency of meeting records.
Smart Images

Figure CN120526758B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and more particularly to a method and system for smart glasses speech recognition based on multi-channel speech enhancement. Background Art
[0002] At present, speech processing technology in multi-person conference scenarios has become a key research direction in the field of smart office. Efficient processing of conference voice information helps to improve the efficiency of meeting records and content analysis. However, in multi-person conference scenarios, due to frequent voice overlap, it is difficult to identify the voice content corresponding to each speaker, and thus it is difficult to determine the direction of speech enhancement. Therefore, the existing technology has shortcomings. Summary of the Invention
[0003] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a smart glasses speech recognition method and system based on multi-channel speech enhancement. After signal separation, the next speaker is accurately predicted through context information and a neural network model, and the speech is enhanced to improve the clarity of the target speech.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] The present invention provides a method for smart glasses speech recognition based on multi-channel speech enhancement. The smart glasses include a plurality of microphones located at different positions and facing different directions to pick up sounds, and all the microphones constitute a microphone array. The method for smart glasses speech recognition includes:
[0006] Acquire a mixed signal corresponding to each microphone in the microphone array, wherein the mixed signal includes voice signals of multiple objects to be recognized;
[0007] Determining a speech signal corresponding to each object to be identified based on a time delay corresponding to each object to be identified, wherein the time delay is determined based on a distance between each object to be identified and each microphone in the microphone array;
[0008] Obtaining current conversation information according to the voice signal and the corresponding object to be identified;
[0009] predicting the direction of the target sound source based on the current conversation information;
[0010] Speech enhancement is performed according to the target sound source direction to obtain an enhanced speech signal.
[0011] As a further improvement of the present invention, determining the speech signal corresponding to each object to be identified based on the time delay corresponding to each object to be identified includes:
[0012] selecting a reference microphone from the microphone array;
[0013] For each object to be identified, determining a phase compensation factor between each remaining microphone and the reference microphone according to the corresponding time delay;
[0014] Obtaining a plurality of phase compensation signals according to the phase compensation factor and each of the mixed signals;
[0015] Amplitude separation is performed on each of the phase-compensated signals to determine a speech signal corresponding to each object to be identified.
[0016] As a further improvement of the present invention, the amplitude separation of each phase compensation signal is performed to obtain multiple speech signals included in each mixed signal, including:
[0017] Obtaining a relative amplitude response based on the distance between the object to be identified and the microphone;
[0018] performing amplitude compensation on the phase compensation signal according to the relative amplitude response to obtain a plurality of amplitude compensation signals;
[0019] Obtaining an enhanced mixing matrix according to the amplitude compensation signal;
[0020] According to the enhanced mixing matrix, a speech signal corresponding to each object to be recognized is determined.
[0021] As a further improvement of the present invention, determining the speech signal corresponding to each object to be recognized according to the enhanced mixing matrix includes:
[0022] Calculating a covariance matrix based on the enhanced mixing matrix;
[0023] Obtaining a weight vector according to the covariance matrix and the steering vector corresponding to the speech signal;
[0024] The speech signal corresponding to each object to be recognized is determined according to the weight vector.
[0025] As a further improvement of the present invention, predicting the target sound source direction according to the current dialogue information includes:
[0026] generating a conversation graph according to the current conversation information, wherein each node in the conversation graph corresponds to a speaking party, and each edge in the conversation graph corresponds to an interaction relationship between two speaking parties;
[0027] determining the number of occurrences of each transfer pair according to the interaction relationship;
[0028] constructing a frequency matrix according to the number of times each transfer pair occurs;
[0029] The target sound source direction is predicted according to the frequency matrix.
[0030] As a further improvement of the present invention, predicting the target sound source direction according to the frequency matrix includes:
[0031] Obtain a probability matrix according to the frequency matrix, wherein each element in the probability matrix represents a transition probability corresponding to each transition pair;
[0032] Obtaining the current speaker according to the current dialogue information;
[0033] Obtaining a target sound source object according to the current speaking object and the probability matrix;
[0034] The target sound source direction is determined according to the target sound source object.
[0035] As a further improvement of the present invention, predicting the target sound source direction according to the current dialogue information includes:
[0036] Obtaining an input vector corresponding to each time step according to the current conversation information;
[0037] Inputting the input vectors into a preset neural network model in chronological order to obtain the speaking probability of each sound source object;
[0038] Predicting the target sound source object according to the speech probability and preset rules;
[0039] The target sound source direction is determined according to the target sound source object.
[0040] As a further improvement of the present invention, obtaining an input vector corresponding to each time step according to the current dialogue information includes:
[0041] Obtaining the speaker identity, language behavior, and speech content at each time step in the current conversation information;
[0042] Obtaining an identity vector, a behavior vector, and a semantic vector according to the identity of the speaker, the language behavior, the speech content, and the corresponding embedding matrix;
[0043] Obtaining a temporal feature vector according to the identity vector, behavior vector, semantic vector and position code;
[0044] According to the position of each speaking object, a spatial feature vector is obtained;
[0045] An input vector corresponding to each time step is obtained according to the time series feature vector and the spatial feature vector.
[0046] As a further improvement of the present invention, performing speech enhancement according to the target sound source direction to obtain an enhanced speech signal includes:
[0047] Calculating a steering vector corresponding to the target sound source direction according to the target sound source direction;
[0048] Obtaining a weighting coefficient according to the steering vector;
[0049] The enhanced speech signal is obtained according to the weighting coefficient.
[0050] The present invention provides a smart glasses speech recognition system based on multi-channel speech enhancement, comprising smart glasses and a server;
[0051] The smart glasses include a plurality of microphones located at different positions and facing different directions to pick up sounds, wherein all the microphones constitute a microphone array, and the microphone array is used to obtain a mixed signal, wherein the mixed signal includes voice signals of multiple objects to be recognized;
[0052] The server includes:
[0053] a separation module, configured to determine a speech signal corresponding to each object to be identified based on a time delay corresponding to each object to be identified, wherein the time delay is determined based on a distance between each object to be identified and each microphone in the microphone array;
[0054] An updating module, configured to obtain current conversation information based on the voice signal and the corresponding object to be identified;
[0055] A prediction module, configured to predict a target sound source direction based on the conversation information;
[0056] The enhancement module is used to perform speech enhancement according to the predicted target sound source direction to obtain an enhanced speech signal.
[0057] The present invention first obtains a mixed signal through a microphone array, and then performs voice separation in combination with time delay to determine the speech content corresponding to each person, thereby taking the semantically recognized speech content as the basis, and predicting the sound source direction of the next speaker of the current conversation information through the semantic association of context information, thereby directionally strengthening the sound source in this sound source direction in advance through the algorithm, which is conducive to the user's attention to the voice information of the next speaker, and improving the user experience of users using smart glasses for multi-person meetings. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic diagram of the steps of the present invention;
[0059] Figure 2 A schematic diagram of the position of the speaker and the microphone array;
[0060] Figure 3 This is a diagram of a conversation graph;
[0061] Figure 4Schematic diagram of the steps for determining the direction of the target sound source;
[0062] Figure 5 Schematic diagram of corrected speaking probability. DETAILED DESCRIPTION
[0063] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations of the technical solution of the present invention.
[0064] The term "and / or" in the following text simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.
[0065] like Figure 1 As shown, the embodiment of the present application provides a smart glasses speech recognition method based on multi-channel speech enhancement, comprising:
[0066] Acquire a mixed signal corresponding to each microphone in the microphone array, where the mixed signal includes voice signals of multiple objects to be recognized;
[0067] Determine the speech signal corresponding to each object to be identified based on the time delay corresponding to each object to be identified, where the time delay is determined based on the distance between each object to be identified and each microphone in the microphone array;
[0068] According to the voice signal and the corresponding object to be identified, the current conversation information is obtained;
[0069] Predict the direction of the target sound source based on the current conversation information;
[0070] Speech enhancement is performed according to the direction of the target sound source to obtain an enhanced speech signal.
[0071] Among them, the microphone array is located on the smart glasses, and the smart glasses include multiple microphones located in different positions and facing different directions to pick up sounds. All microphones constitute a microphone array. For example, multiple microphones can be embedded in the frame of the smart glasses to form a microphone array. The smart glasses also include a binocular camera for recording video, and voice and video should be collected synchronously.
[0072] Exemplarily, the embodiments of the present application are applicable to content recording in scenarios such as multi-person meetings. In such scenarios, each person's position is fixed. When the meeting or debate begins, one of them wears and turns on the smart glasses to start collecting voice signals and recording videos. First, at the beginning of the meeting, voice signals and videos are collected over a period of time to obtain enough conversation information for prediction. For each signal collected by each microphone, if the signal corresponds to only one speaker, the speaker is confirmed from the video based on the time corresponding to the signal, and face recognition is performed. When the identity of the participant is known, the identity information of the speaker can be associated with the voice signal to generate a record corresponding to the voice signal, including the identity of the speaker, language behavior, and speech content. The identity can be divided into marketing specialists or project managers, and the language behavior can be divided into questions or answers, etc. If the collected signal is a mixed signal, that is, when multiple people speak at the same time, the mixed signal needs to be separated to obtain voice signals of multiple objects to be identified, that is, the speaking objects, and generate text content of multiple voice signals. If the text content is related to the meeting content, it will be retained. If it is not related, it means that the voice signal is an interference signal and needs to be deleted. Then, the sound source position corresponding to each retained voice signal is determined, that is, the position of each corresponding speaking object. The identity of each speaking object is determined from the video according to the sound source position, and finally multiple records corresponding to the mixed signal are obtained. Similarly, each record includes the identity of the speaking object, language behavior and speech content. Finally, the set of each record obtained during this period is used as the current dialogue information, and each record in the current dialogue information is arranged in chronological order.
[0073] Then, the target sound source direction, that is, the direction of the next speaker, is predicted based on the current conversation signal, and the voice signal collected in that direction is enhanced to obtain an enhanced voice signal. Finally, the above steps of generating records are repeated based on the enhanced voice signal, and the record generated by the enhanced voice signal is used together with all the above records as the current conversation information and the prediction step is continued.
[0074] This embodiment collects sufficient information at the beginning of the meeting to accurately predict the direction of the target sound source, and performs speech enhancement to obtain an enhanced speech signal after avoiding interference signals, so that in the subsequent step of generating records, the speech content corresponding to the next speaker can be accurately identified without signal separation. The collaborative solution of signal separation, speech prediction and speech enhancement provided by the present invention can avoid interference from other signals, and the accurate recognition of the target speech signal can be maintained at all times during the meeting, ultimately obtaining accurate meeting records.
[0075] Furthermore, this embodiment provides a step of determining a speech signal corresponding to each object to be identified based on the time delay corresponding to each object to be identified, including:
[0076] Selecting a reference microphone from the microphone array;
[0077] For each object to be identified, determining a phase compensation factor between each remaining microphone and the reference microphone according to the corresponding time delay;
[0078] A plurality of phase compensation signals are obtained according to the phase compensation factor and each mixed signal;
[0079] Each phase-compensated signal is amplitude-separated to determine the speech signal corresponding to each object to be identified.
[0080] For example, when multiple people speak at the same time, each microphone will receive a mixed signal, which can be represented by the sum of multiple voice signals. First, it is necessary to identify the multiple objects to be identified who are speaking at the same time based on the time when the mixed signal is received and the video, and determine the corresponding time delay based on the position of each object to be identified.
[0081] Specifically, based on the parallax principle of the binocular camera, the actual distance between each object to be identified and each microphone can be obtained. Then, one of the multiple microphones is randomly selected as the reference microphone, which is recorded as the first The distance corresponding to the reference microphone is called the reference distance, and the other microphones are called the remaining microphones. The remaining microphones are numbered in a certain position order and recorded as To The distances corresponding to the remaining microphones are called To Then, for each object to be identified, the distance difference between each residual distance and the reference distance is determined, and a total of Divide each distance difference by the speed of sound , get the corresponding delays, one for each remaining microphone.
[0082] Then, the mixed signal of each microphone is converted into a frequency domain signal through short-time Fourier transform to obtain the frequency domain signal corresponding to each microphone, and the following steps are repeated for the frequency domain signal corresponding to each remaining microphone. frequency domain signal For example, The frequency domain signal is the The mixed signal corresponding to each microphone is converted into a frequency domain signal, ,in is the frequency, is a frame index, and one of the speaking objects is selected from multiple speaking objects, for example, the first Speaker, from its corresponding The delay is determined by The delay corresponding to each microphone , and calculate the Phase compensation factor of the microphone and the reference microphone ,in Is an imaginary unit.
[0083] Then according to the phase compensation factor right Perform phase compensation to obtain Phase compensation signal , then according to The microphone and The distance between the speaker , and the relative amplitude response is obtained , and Perform amplitude compensation to obtain Amplitude compensation signal , repeat the above steps for the frequency domain signals corresponding to each microphone of the remaining s, and we get Amplitude compensation signal , where the speaker selected each time is repeated is the same, and then for the The frequency domain signal corresponding to the microphone is obtained by the same amplitude compensation method as above. Amplitude compensation signal, now we get Amplitude compensation signals are arranged in sequence to obtain an enhanced mixing matrix. for:
[0084]
[0085] The covariance matrix calculated according to the enhanced mixing matrix is: ,in represents the conjugate transpose, , then calculate the weight vector according to the covariance matrix and the enhanced mixing matrix: , for The inverse matrix of For the The steering vector of the speech signal corresponding to the speaking object is used to describe the phase difference between the speech signal and each microphone. Since phase compensation and amplitude compensation have been performed, it can be assumed that is a full 1 vector, and then the first The frequency domain representation of the speech signal corresponding to each speaker is: , and will eventually By converting it back to the time domain through inverse short-time Fourier transform, we can get The speech signal corresponding to the speaker is used to perform speech recognition and obtain the The speech content corresponding to the speaking object, wherein in order to further enhance the accuracy of speech recognition, the signal can be denoised by spectral subtraction before recognition. Spectral subtraction is a conventional method used by those skilled in the art and will not be described in detail in this embodiment.
[0086] Through the above steps, the The speech signal corresponding to each speaker is obtained, and then other speakers are selected to repeat the above steps to separate the speech signal corresponding to each speaker, and obtain the record corresponding to each speech signal, thereby obtaining the current speech information.
[0087] The principle of the phase compensation factor in this embodiment is that when the sound wave is transmitted to the microphone array, due to the different positions of different microphones, a time delay will be generated during the propagation of the sound wave, such as Figure 2 As shown. Speaker, The first microphone is relative to the The delay of each microphone is , this delay will cause the microphone and The microphone receives the There will be a phase difference between the speech signals of the speakers , in order to make the The voice signal received by the first microphone is The voice signals received by the microphones are aligned to eliminate time delay. In this embodiment, phase compensation is applied to them in the frequency domain through the phase compensation factor. In addition, this embodiment only sets the relative amplitude response for amplitude compensation based on the principle that the signal is inversely proportional to the amplitude signal. However, this embodiment is not limited to this. In actual application, in order to improve the accuracy of amplitude compensation, the influence of medium absorption, reflection or scattering on the amplitude can be further considered. After phase compensation and amplitude compensation, The amplitude and phase of the speech signal of each speaker in the corresponding mixed signal of each microphone are consistent, while the signals of other sound sources still have differences. Therefore, this embodiment first uses the covariance matrix to represent the spatial correlation of each signal. The speech signal of the speaker is expressed as the outer product of all-1 vectors in the covariance matrix, while the covariance matrix components of other signals are orthogonal to the all-1 vectors because of inconsistent amplitudes. Based on this, this embodiment further constructs a filter that matches only the all-1 vectors through the MVDR (minimum variance distortionless response) algorithm. This filter is the weight vector in this embodiment, thereby separating the first The speech signal of a speaker.
[0088] Based on the above analysis, it can be summarized that the steps provided in this embodiment of performing amplitude separation on each phase-compensated signal to determine the speech signal corresponding to each object to be recognized include:
[0089] Obtaining a relative amplitude response based on the distance between the object to be identified and the microphone;
[0090] performing amplitude compensation on the phase compensation signal according to the relative amplitude response to obtain a plurality of amplitude compensation signals;
[0091] Obtaining an enhanced mixing matrix according to the amplitude compensation signal;
[0092] According to the enhanced mixing matrix, the speech signal corresponding to each object to be recognized is determined.
[0093] The step of obtaining multiple speech signals included in each mixed signal according to the enhanced mixing matrix provided in this embodiment includes:
[0094] According to the enhanced mixing matrix, the covariance matrix is calculated;
[0095] Obtain a weight vector according to the covariance matrix and the steering vector corresponding to the speech signal;
[0096] According to the weight vector, multiple speech signals included in each mixed signal are obtained.
[0097] This embodiment determines the phase compensation factor and relative amplitude response through time delay, and compensates the signal to be separated so that its amplitude and phase in the corresponding mixed signal of each microphone are consistent, and then sets a matching weight vector to separate them. Compared with other signal separation methods, this embodiment does not need to rely on prior information and independence assumptions, and directly performs separation based on spatial information, reducing the amount of calculation.
[0098] Furthermore, this embodiment provides a step of predicting the direction of a target sound source based on current dialogue information, including:
[0099] Generate a conversation graph based on the current conversation information. Each node in the conversation graph corresponds to a speaker, and each edge in the conversation graph corresponds to the interaction between two speakers.
[0100] Determine the number of occurrences of each transfer pair based on the interaction relationship;
[0101] Construct a frequency matrix based on the number of times each transfer pair occurs;
[0102] Predict the target sound source direction based on the frequency matrix.
[0103] For example, Figure 3As shown, assuming that there are three speaking objects in the current dialogue information, each node corresponds to a speaking object. After the speech Then they speak, and establish a boundary between them. , the edge weight represents After the speech When two nodes have multiple conversations with each other or need to further express the interaction relationship and time sequence, multiple edges can be established between the two nodes. After the speech If the number of times you speak is 3, you can Add an edge between , with a weight of 3; or for example First time When asking questions, generate edges , First reply Generate edges The conversation graph can be used to intuitively display the interaction relationship in the current conversation. Each time a new record is added, a new edge can be directly added to the conversation graph. When the conversation frequency needs to be obtained later, it can be directly called based on the conversation graph without traversing all the current conversation information.
[0104] According to the speaking object and the direction of each edge, the transfer pair can be determined, such as Figure 3 The transfer pairs present in 、 、 and Then, the frequency matrix is constructed based on the number of times each transfer pair appears in the conversation graph. Each element in the frequency matrix corresponds to the number of times a transfer pair appears. The number of occurrences is After the speech The number of speeches that followed.
[0105] Then, the transition probability corresponding to each transition pair is calculated based on the frequency. For example, for the transition pair ,statistics After each speech or The number of times you speak is called the total number of times, followed by The number of times the next speech is divided by the total number of times is the transfer pair The corresponding transition probability. The last speaker in the current conversation is then determined, i.e., the current speaker. The transition pair with the highest transition probability is selected based on the transition probability matrix. The other speaker in this transition pair is used as the target sound source, and its direction is used as the target sound source direction.
[0106] Based on the above analysis, it can be summarized that the steps provided in this embodiment for predicting the target sound source direction based on the frequency matrix include:
[0107] The probability matrix is obtained based on the frequency matrix, and each element in the probability matrix represents the transition probability corresponding to each transition pair;
[0108] Get the current speaker based on the current dialogue information;
[0109] According to the current speaking object and probability matrix, the target sound source object is obtained;
[0110] Determine the target sound source direction based on the target sound source object.
[0111] This embodiment determines the transition probability matrix based on frequency, more intuitively inferring the next speaker, i.e., the target sound source object. This approach has low computational cost and high efficiency, making it suitable for conversation scenarios with strong regularity. However, this method only captures short-range dependencies, i.e., it is determined based only on the last speaker in the current conversation information, and cannot perceive semantic changes in the conversation content, which affects the accuracy of the prediction.
[0112] Furthermore, in order to enhance the accuracy of prediction, Figure 4 As shown, this embodiment provides another step of predicting the target sound source direction based on the current dialogue information, including:
[0113] Get the input vector corresponding to each time step based on the current dialogue information;
[0114] Input the input vectors into the preset neural network model in chronological order to obtain the speaking probability of each sound source object;
[0115] Predict the target sound source object based on speech probability and preset rules;
[0116] Determine the target sound source direction based on the target sound source object.
[0117] Among them, the preset neural network model can be a Transformer model, and the model used in this embodiment is a trained model. The training of the model is a conventional technical means and will not be described in detail in this embodiment.
[0118] Specifically, each record in the current conversation information corresponds to a time step. For each record, obtain the corresponding speaker identity , mapping it into a one-hot vector , and according to the embedding matrix corresponding to the identity of the speaker Get the identity vector , similarly, language behavior Through its corresponding embedding matrix Get the behavior vector Then the speech content is input into the trained BERT model and the output CLS position vector is used as the semantic vector , and generate the position code corresponding to the time step , the position encoding implies the order of speaking. For example, the sine-cosine function can be selected for encoding. Then, the identity vector, behavior vector, semantic vector and position encoding are unified in dimension and added together to obtain the time series feature vector corresponding to the record, that is, the time series feature vector corresponding to the time step. Then, the position coordinates of the speaking object corresponding to the record are obtained, and the spatial feature vector is obtained by the embedding matrix. Then, the dimension of the spatial feature vector is unified with the time series feature vector, and finally the main elements of the two are spliced or added element by element to obtain the input vector corresponding to the time step. Among them, the step of unifying the dimension can be achieved through a linear layer, and each embedding matrix is obtained according to model training.
[0119] Repeating the above steps yields the input vector corresponding to each time step, which is then fed into the Transformer model in chronological order. The Transformer model consists of at least an encoder and a fully connected layer (output layer). The encoder first extracts features using an attention mechanism, then linearly transforms the attention mechanism's output using a feedforward neural network to further enhance the features' expressiveness. The fully connected layer derives the speaking probability of each sound source based on the encoder's output. Specifically, the encoder outputs a global feature vector, which abstractly represents all records in the current conversation, incorporating information such as chronological order, speaker identity, and speech content. The fully connected layer maps this vector to a target space whose dimensions are the total number of sound sources (i.e., the total number of participants). Ultimately, the speaking probability of each sound source is obtained. The sound source with the highest speaking probability is designated as the target sound source, and its direction is designated as the target sound source direction.
[0120] Based on the above analysis, it can be summarized that the steps provided in this embodiment for obtaining the input vector corresponding to each time step based on the current conversation information include:
[0121] Obtain the speaker identity, language behavior, and speech content at each time step in the current conversation information;
[0122] According to the speaker's identity, language behavior, speech content and corresponding embedding matrix, the identity vector, behavior vector and semantic vector are obtained;
[0123] According to the identity vector, behavior vector, semantic vector and position encoding, a temporal feature vector is obtained;
[0124] According to the position of each speaking object, a spatial feature vector is obtained;
[0125] According to the time series feature vector and the spatial feature vector, the input vector corresponding to each time step is obtained.
[0126] This embodiment is based on a preset neural network model, which can accurately integrate temporal, spatial, and semantic multimodal information, and simultaneously capture the direct switching of neighboring rounds and indirect dependencies across multiple rounds through the attention mechanism. It is suitable for predicting speaking objects in complex scenarios and improves the accuracy of prediction.
[0127] However, the above method is only used when there is enough data. When the meeting time is short, there is not enough time to collect information, resulting in less data about certain speakers, which affects the accuracy of the prediction. To solve this problem, this embodiment further provides a step for predicting the target sound source direction based on the current dialogue information. Specifically, first, based on the less dialogue information and the above Transformer model, according to the above steps, the speaking probability of each sound source is obtained. If there is a speaker whose speaking probability and the number of speeches recorded in the current dialogue information are both less than or equal to the preset value, and there may be multiple speakers, the reasoning mechanism based on the generative diffusion model is triggered. Then, based on the identity, location and current dialogue information of the speaker, a virtual speech feature is generated, and the similarity between the virtual speech feature and the feature of the last collected mixed signal is calculated. If the similarity is high, it means that the speaker is more likely to speak. It is necessary to correct the speaking probability output by the Transformer model to obtain the final speaking probability. Finally, based on the final speaking probability of each speaker, the speaker with the highest probability is selected as the target sound source object, and its direction is used as the target sound source direction. For example Figure 5 As shown, when the Transformer model outputs a speech object 、 、 The corresponding speaking probability 、 、 Afterwards, if If the probability of speaking and the number of times spoken recorded in the current dialogue information are both less than or equal to the preset value, the diffusion model is generated to obtain the virtual speech features and correct get , then according to 、 、 The speaking object corresponding to the highest probability value is selected as the target sound source object.
[0128] The preset values for the number of speeches and the probability of speaking can be determined based on historical meeting situations. For example, after analyzing the data of 50 30-minute business meetings of the same type, it is found that the average number of speeches per participant is 5. If the estimated duration of this meeting is also 30 minutes and the current conversation information includes the first 5 minutes of the conversation content of this meeting, the preset value is , that is, objects with 1 or 0 speeches should be selected; or historical meeting data can be input into the Transformer model, and statistical analysis can be performed on the predicted probabilities given by the Transformer model under different speeches, and a probability distribution curve can be drawn. If statistics show that in a 30-minute meeting with a participant who speaks very few times (such as 1 time), the Transformer model's output probability for such participants in the range of 0-0.1 accounts for 80%, then 0.1 can be used as a preset value. The method of selecting the preset value provided in this embodiment is only an example. Other methods can be selected in actual applications, and this embodiment does not limit this.
[0129] Then, the generative diffusion model is used to generate virtual voice features based on the identity, location and current conversation information of the speaker. The core principle of the generative diffusion model is to learn and reconstruct the data distribution, which can be divided into two key processes: forward diffusion and reverse denoising. In the forward diffusion process, the model gradually adds Gaussian noise to real data (such as voice features in meetings). As time goes on, the data gradually loses its original features and eventually approaches a simple Gaussian distribution. The reverse denoising process is to learn to gradually remove noise from noisy data close to a Gaussian distribution and restore the original data. In essence, it is the inverse process of learning data distribution, that is, by continuously predicting and correcting noise, the model can generate samples that conform to the distribution of the original data. This embodiment primarily utilizes the reverse denoising process of a generative diffusion model. Specifically, the identity vector and spatial feature vector of the speaker are first obtained. The specific steps are the same as described above. Semantic analysis is then performed on the current conversation information to determine the meeting theme and specific discussion direction. For example, a keyword extraction algorithm (such as TextRank) is first used to extract high-frequency and representative keywords from the current conversation information. For example, keywords such as "marketing promotion," "marketing strategy," and "target customers" may be extracted. These keywords are then matched with a preset thesaurus to determine that the current meeting theme is "marketing promotion strategy." A pre-trained semantic model, such as BERT, is then used to obtain the semantic vector corresponding to the current conversation information. A text classification model is then used to determine the specific content direction of the current discussion (such as budget allocation, channel selection, or effect evaluation). Finally, based on the speaker's identity, a corresponding language behavior is generated for the speaker. For example, if the speaker is a marketing specialist, the current meeting theme is marketing promotion strategy, and the discussion involves target customer groups, a language behavior such as "What are the needs of the target customer group?" is generated based on the keywords and discussion content, and this language behavior is encoded as a feature vector. Among them, the keyword extraction algorithm and semantic model selected in this embodiment are only examples, and their purpose is to determine the meeting theme and specific discussion content. In actual application, other appropriate methods can be selected to determine, and this embodiment does not limit this. In addition, generating relevant text based on specific keywords and content is also a conventional technical means, which will not be elaborated in this embodiment.
[0130] After obtaining the identity vector, spatial feature vector, and feature vector, they are also dimensionalized through a linear layer. They are then concatenated to obtain a conditional vector, which is input into the generative diffusion model together with the initial noise vector sampled from the standard Gaussian distribution. In each iteration, the model predicts the current noise through the internal neural network, and then iteratively calculates based on the current noise to obtain the virtual speech feature. The specific calculation formula is:
[0131]
[0132] in, Indicates the current number of iterations. hour, is the initial noise vector, is the conditional vector, and is the preset noise scaling factor, is the variance control parameter, The current noise is predicted by the neural network. The neural network has been trained to output noise that meets specific conditions based on the input condition vector. , Specifically, As the number of iterations increases, it gradually decreases and is used to control the direction of denoising. It is used to determine the amplitude of denoising. The variance control parameter is multiplied by the normal distribution term. It can retain a certain degree of randomness in the process of generating virtual speech features to avoid generating noisy data each time. The similarity is too high. The multiplication here is not a simple numerical multiplication, but it means sampling a noise from the standard normal distribution and multiplying the noise with the square root of the variance control parameter. Finally, the multiplication result is compared with For example, in order to ensure that the noise changes evenly at each step, a linear attenuation strategy can be used to determine :
[0133]
[0134] in express The initial value of Indicates the end value, preferably, Set as , indicating that the final noise is extremely weak, Set as , indicating that the initial noise is close to Gaussian distribution, Indicates the total number of iterations, for example, setting .
[0135] Finally, through the iterative steps, a virtual speech feature (vector) can be obtained, which contains the speech acoustic characteristics and semantic information that meet the set conditions, such as the speech frequency characteristics related to the "target user group consumption habits" and the semantic features corresponding to the keywords. Then, the last mixed signal collected is obtained and converted into a vector with the same dimension as the virtual speech feature (vector). Then, the cosine similarity between them is calculated. If the similarity is higher than the preset threshold, it means that the speaker is more likely to speak. It is necessary to correct the speaking probability output by the Transformer model to obtain the final speaking probability. The cosine similarity-based method given in this embodiment is only an example. Other similarity methods can be selected in actual applications, and this embodiment does not limit this.
[0136] Corrected probability ,in represents the speaking probability output by the Transformer model, Represents the weight adjustment coefficient, which is used to control the influence of cosine similarity on probability correction. For example, setting After the probability correction is performed, the object with the highest probability is selected as the target sound source object according to the probabilities corresponding to all speaking objects, and its direction is taken as the target sound source direction.
[0137] This embodiment takes into account that when the amount of data is insufficient, some participants may speak less frequently, which in turn affects the accuracy of the predicted speaking probability. Therefore, this embodiment is based on the reverse denoising strategy of the generative diffusion model. The conditional vector guides the generation process to obtain virtual voice features, filling the data gap. The virtual voice features represent the voice features that the speaker may make in the current meeting. If the similarity between the virtual voice features and the features of the last collected mixed signal is high, it means that the speaker has a high probability of speaking in the current meeting. However, due to the short signal collection time and insufficient data volume, this information cannot be captured in time. Therefore, this embodiment performs probability correction based on similarity to obtain a more accurate target sound source object.
[0138] Furthermore, this embodiment provides a step of performing speech enhancement according to the target sound source direction to obtain an enhanced speech signal, including:
[0139] According to the direction of the target sound source, the steering vector corresponding to the target sound source direction is calculated;
[0140] Obtaining a weighting coefficient according to the steering vector;
[0141] The target sound source object is obtained according to the weighting coefficient.
[0142] Specifically, after determining the target sound source, if the microphone array still collects mixed signals, the signals collected by each microphone need to be weighted according to the weighting coefficient to obtain an enhanced speech signal. The weighting coefficient is calculated by first calculating the steering vector corresponding to the target sound source direction. The steering vector describes the time delay and phase relationship of the sound reaching each microphone. Then, the steering vector is normalized to obtain the weighting coefficient. The weighting coefficient is used to adjust the amplitude and phase of the signal received by each microphone, so that the speech signal in the target direction is enhanced after weighted summation, and the target sound source object is obtained.
[0143] The embodiment of the present application provides a smart glasses speech recognition method and system based on multi-channel speech enhancement. The method first obtains a mixed signal through a microphone array, and then performs speech separation in combination with time delay to determine the speech content corresponding to each person. The semantically recognized speech content is used as a basis, and the semantic association of context information is used to predict the sound source direction of the next speaker of the current dialogue information, so as to strengthen the sound source in this sound source direction in advance through the algorithm, which is conducive to the user's attention to the voice information of the next speaker and improves the user experience of using smart glasses for multi-person meetings.
[0144] This embodiment provides a smart glasses speech recognition system based on multi-channel speech enhancement, including smart glasses and a server;
[0145] The smart glasses include a microphone array for acquiring a mixed signal including voice signals of multiple objects to be recognized;
[0146] The server includes:
[0147] a separation module for determining a speech signal corresponding to each object to be identified based on a time delay corresponding to each object to be identified, where the time delay is determined based on a distance between each object to be identified and each microphone in the microphone array;
[0148] An update module is used to obtain current conversation information based on the voice signal and the corresponding object to be identified;
[0149] A prediction module, used to predict the direction of the target sound source based on the conversation information;
[0150] The enhancement module is used to perform speech enhancement according to the predicted target sound source direction to obtain an enhanced speech signal.
[0151] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0153] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0154] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A smart glasses speech recognition method based on multi-channel speech enhancement, characterized in that: The smart glasses include a plurality of microphones located at different positions and facing different directions to pick up sounds, and all the microphones constitute a microphone array. The smart glasses voice recognition method includes: Acquire a mixed signal corresponding to each microphone in the microphone array, wherein the mixed signal includes voice signals of multiple objects to be recognized; Determining a speech signal corresponding to each object to be identified based on a time delay corresponding to each object to be identified, wherein the time delay is determined based on a distance between each object to be identified and each microphone in the microphone array; Obtaining current conversation information according to the voice signal and the corresponding object to be identified; predicting the direction of the target sound source based on the current conversation information; Performing speech enhancement according to the target sound source direction to obtain an enhanced speech signal; The step of determining the speech signal corresponding to each object to be identified based on the time delay corresponding to each object to be identified includes: selecting a reference microphone from the microphone array; For each object to be identified, determining a phase compensation factor between each remaining microphone and the reference microphone according to the corresponding time delay; Obtaining a plurality of phase compensation signals according to the phase compensation factor and each of the mixed signals; Performing amplitude separation on each of the phase-compensated signals to determine a speech signal corresponding to each object to be identified; The step of performing amplitude separation on each phase-compensated signal to determine a speech signal corresponding to each object to be identified includes: Obtaining a relative amplitude response based on the distance between the object to be identified and the microphone; performing amplitude compensation on the phase compensation signal according to the relative amplitude response to obtain a plurality of amplitude compensation signals; Obtaining an enhanced mixing matrix according to the amplitude compensation signal; Determining a speech signal corresponding to each object to be recognized according to the enhanced mixing matrix; The method of determining the speech signal corresponding to each object to be recognized according to the enhanced mixing matrix includes: Calculating a covariance matrix based on the enhanced mixing matrix; Obtaining a weight vector according to the covariance matrix and the steering vector corresponding to the speech signal; Determining the speech signal corresponding to each object to be identified according to the weight vector; Predicting the target sound source direction according to the current conversation information includes: generating a conversation graph according to the current conversation information, wherein each node in the conversation graph corresponds to a speaking party, and each edge in the conversation graph corresponds to an interaction relationship between two speaking parties; determining the number of occurrences of each transfer pair according to the interaction relationship; constructing a frequency matrix according to the number of times each transfer pair occurs; predicting the target sound source direction according to the frequency matrix; Predicting the target sound source direction according to the frequency matrix includes: Obtain a probability matrix according to the frequency matrix, wherein each element in the probability matrix represents a transition probability corresponding to each transition pair; Obtaining the current speaker according to the current dialogue information; Obtaining a target sound source object according to the current speaking object and the probability matrix; The target sound source direction is determined according to the target sound source object.
2. The method for smart glasses speech recognition based on multi-channel speech enhancement according to claim 1, characterized in that: Predicting a target sound source direction according to the current conversation information includes: Obtaining an input vector corresponding to each time step according to the current conversation information; Inputting the input vectors into a preset neural network model in chronological order to obtain the speaking probability of each sound source object; Predicting the target sound source object according to the speech probability and preset rules; The target sound source direction is determined according to the target sound source object.
3. The method for smart glasses speech recognition based on multi-channel speech enhancement according to claim 2, characterized in that: The input vector corresponding to each time step is obtained according to the current dialogue information, including: Obtaining the speaker identity, language behavior, and speech content at each time step in the current conversation information; Obtaining an identity vector, a behavior vector, and a semantic vector according to the identity of the speaker, the language behavior, the speech content, and the corresponding embedding matrix; Obtaining a temporal feature vector according to the identity vector, behavior vector, semantic vector and position code; According to the position of each speaking object, a spatial feature vector is obtained; An input vector corresponding to each time step is obtained according to the time series feature vector and the spatial feature vector.
4. The method for smart glasses speech recognition based on multi-channel speech enhancement according to claim 1, characterized in that: Performing speech enhancement according to the target sound source direction to obtain an enhanced speech signal includes: Calculating a steering vector corresponding to the target sound source direction according to the target sound source direction; Obtaining a weighting coefficient according to the steering vector; The enhanced speech signal is obtained according to the weighting coefficient.
5. A smart glasses speech recognition system based on multi-channel speech enhancement, characterized in that: Including smart glasses and servers; The smart glasses include a plurality of microphones located at different positions and facing different directions to pick up sounds, wherein all the microphones constitute a microphone array, and the microphone array is used to obtain a mixed signal, wherein the mixed signal includes voice signals of multiple objects to be recognized; The server includes: a separation module, configured to determine a speech signal corresponding to each object to be identified based on a time delay corresponding to each object to be identified, wherein the time delay is determined based on a distance between each object to be identified and each microphone in the microphone array; An updating module, configured to obtain current conversation information based on the voice signal and the corresponding object to be identified; A prediction module, configured to predict a target sound source direction based on the conversation information; an enhancement module, configured to perform speech enhancement according to the predicted target sound source direction to obtain an enhanced speech signal; The step of determining the speech signal corresponding to each object to be identified based on the time delay corresponding to each object to be identified includes: selecting a reference microphone from the microphone array; For each object to be identified, determining a phase compensation factor between each remaining microphone and the reference microphone according to the corresponding time delay; Obtaining a plurality of phase compensation signals according to the phase compensation factor and each of the mixed signals; Performing amplitude separation on each of the phase-compensated signals to determine a speech signal corresponding to each object to be identified; The step of performing amplitude separation on each phase-compensated signal to determine a speech signal corresponding to each object to be identified includes: Obtaining a relative amplitude response based on the distance between the object to be identified and the microphone; performing amplitude compensation on the phase compensation signal according to the relative amplitude response to obtain a plurality of amplitude compensation signals; Obtaining an enhanced mixing matrix according to the amplitude compensation signal; Determining a speech signal corresponding to each object to be recognized according to the enhanced mixing matrix; The method of determining the speech signal corresponding to each object to be recognized according to the enhanced mixing matrix includes: Calculating a covariance matrix based on the enhanced mixing matrix; Obtaining a weight vector according to the covariance matrix and the steering vector corresponding to the speech signal; Determining the speech signal corresponding to each object to be identified according to the weight vector; Predicting the target sound source direction according to the current conversation information includes: generating a conversation graph according to the current conversation information, wherein each node in the conversation graph corresponds to a speaking party, and each edge in the conversation graph corresponds to an interaction relationship between two speaking parties; determining the number of occurrences of each transfer pair according to the interaction relationship; constructing a frequency matrix according to the number of times each transfer pair occurs; predicting the target sound source direction according to the frequency matrix; Predicting the target sound source direction according to the frequency matrix includes: Obtain a probability matrix according to the frequency matrix, wherein each element in the probability matrix represents a transition probability corresponding to each transition pair; Obtaining the current speaker according to the current dialogue information; Obtaining a target sound source object according to the current speaking object and the probability matrix; The target sound source direction is determined according to the target sound source object.
Citation Information
Patent Citations
Multi-channel speech enhancement method and device
CN113030862A
Multi-channel speech enhancement method based on reference microphone optimization
CN113257270A
Cited By
A conference voice interaction method and system based on AI glasses
CN122738514A