A multimodal wake-up-free system and method based on voice and sight
By combining multimodal recognition technology of voice and sight, the problem of false wake-up of smart home appliances is solved, more accurate device response is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202210381839.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-04-12
AI Technical Summary
In the prior art, smart home appliance wake-up based solely on voice recognition is prone to false wake-up, causing trouble to users.
A multimodal wake-up-free system based on voice and vision is adopted. The voice and video signals are obtained simultaneously through the acquisition module, and the voice and vision multimodal recognition model is used for recognition. The response is made based on the voice and vision recognition results.
This prevents the device from waking up accidentally, improves the user experience, and ensures that the device responds only when the voice command and eye gaze are consistent, reducing incorrect operations.
Smart Images

Figure CN114999458B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart homes, and in particular to a multimodal wake-up-free system and method based on voice and sight. Background Art
[0002] Multimodality is a key area of research in artificial intelligence and a key development trend in human-computer interaction. Multimodal interaction can improve the accuracy and reduce the difficulty of human-computer interaction. Furthermore, by leveraging the connections between multiple modalities, it enables mutual supervision, mutual calibration, and adaptive learning between different modalities.
[0003] As voice recognition technology matures, it's being incorporated into more and more smart home appliances. Before voice recognition can begin, a wake-up word is needed to activate the voice recognition function of smart home appliances, such as the "Xiao Ai" smart speaker and the "Tmall Genie" wake-up word. Electronic devices typically use voice activity detection to detect voice activity from real-time recorded data. To improve the user experience, many electronic devices currently use a wake-up-free mode for voice interaction. Wake-up-free mode involves issuing commands (i.e., command words) directly, without the need for a wake-up word. For example, if I want to watch "XXX," I only need to issue commands like "Open XXX for me," "Play in full screen," "Fast forward 20 seconds," or "Pause."
[0004] However, in actual use, various noises, chats, etc. may cause smart home appliances to wake up by mistake through voice recognition alone, causing trouble to users. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the defect in the prior art that smart home appliances are awakened only based on voice recognition, which easily leads to false awakening and causes trouble to users, thereby providing a multimodal wake-up-free system and method based on voice and vision.
[0006] The embodiment of the present invention provides a multimodal wake-up-free system based on voice and sight, including: an acquisition module, a calculation module and a response module;
[0007] The acquisition module is used to collect voice signals and video signals, and transmit the voice signals and video signals to the calculation module;
[0008] The calculation module is used to preprocess the voice signal and the video signal, input the preprocessed voice signal and the video signal into the voice and sight multimodal recognition model, and generate a voice recognition result and a sight recognition result;
[0009] The response module is used to obtain the voice recognition result and the sight line recognition result, and respond based on the voice recognition result and the sight line recognition result.
[0010] Optionally, the calculation module includes: a first encoding submodule, a second encoding submodule and an identification submodule;
[0011] The first encoding submodule is used to extract acoustic features from the speech signal and encode the acoustic features to generate speech time series features;
[0012] The second encoding submodule is used to extract the sight line features in the video signal and encode the sight line features to generate video timing features;
[0013] The recognition submodule is used to input the speech timing features and the video timing features into the multimodal recognition model, and output a speech recognition result and a sight line recognition result.
[0014] Optionally, the voice recognition result is compared with a preset voice command, and the line of sight recognition result is compared with a preset line of sight command, and a response is performed when the voice result matches the preset voice command and the line of sight recognition result matches the preset line of sight command.
[0015] Optionally, the first encoding submodule includes: a processing unit, an extraction unit and a first encoding unit;
[0016] The processing unit is used to perform acoustic processing on the speech signal to generate an acoustically processed speech signal;
[0017] The extraction unit is used to extract acoustic features from the acoustically processed speech signal;
[0018] The first encoding unit is used to encode the acoustic features of different time series to generate speech time series features.
[0019] Optionally, the second encoding submodule includes: a detection unit and a second encoding unit;
[0020] The detection unit is used to perform face detection on the video signal, generate a face image, and extract sight features from the face image;
[0021] The second encoding unit is used to encode the sight features of different time series to generate video timing features.
[0022] In the second aspect of the present application, a multimodal wake-up-free method based on voice and sight is also proposed, comprising the following steps:
[0023] Collect voice and video signals;
[0024] Preprocessing the voice signal and the video signal, inputting the preprocessed voice signal and the video signal into a voice-gaze multimodal recognition model to generate a voice recognition result and a gaze recognition result;
[0025] Respond based on the voice recognition result and the line of sight recognition result.
[0026] Optionally, preprocessing the voice signal and the video signal, inputting the preprocessed voice signal and the video signal into a voice-gaze multimodal recognition model, and generating a voice recognition result and a gaze recognition result includes:
[0027] Extracting acoustic features from the speech signal and encoding the acoustic features to generate speech time series features;
[0028] Extracting sight features from the video signal and encoding the sight features to generate video timing features;
[0029] The speech timing features and the video timing features are input into the multimodal recognition model, and a speech recognition result and a sight line recognition result are output.
[0030] Optionally, responding based on the voice recognition result and the sight line recognition result includes:
[0031] The voice recognition result is compared with a preset voice command, and the sight line recognition result is compared with a preset sight line command, and a response is performed when the voice result matches the preset voice command and the sight line recognition result matches the preset sight line command.
[0032] Optionally, extracting acoustic features from the speech signal and encoding the acoustic features to generate speech time series features includes:
[0033] Performing acoustic processing on the speech signal to generate an acoustically processed speech signal;
[0034] Extracting acoustic features from the acoustically processed speech signal;
[0035] Encode the acoustic features of different time series to generate speech time series features.
[0036] Optionally, extracting the sight line features from the video signal and encoding the sight line features to generate video timing features includes:
[0037] Performing face detection on the video signal to generate a face image, and extracting sight line features from the face image;
[0038] The sight line features of different time series are encoded to generate video temporal features.
[0039] In the third aspect of the present application, a computer device is also proposed, including a processor and a memory, wherein the memory is used to store a computer program, the computer program includes a program, and the processor is configured to call the computer program to execute the method of the above-mentioned first aspect.
[0040] In a fourth aspect of the present application, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect above.
[0041] The technical solution of the present invention has the following advantages:
[0042] The present invention provides a multimodal wake-up-free system based on voice and sight, which simultaneously collects voice signals and video signals, uses a voice and sight multimodal recognition model to recognize the voice signals and video signals, responds with the voice recognition results and uses both the voice recognition results and the sight recognition results as necessary factors for waking up the terminal, thereby avoiding the trouble caused to users by false awakening of the device and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 This is a functional block diagram of a multimodal wake-up-free system based on voice and sight in Example 1 of the present invention;
[0045] Figure 2 This is a flow chart of a multimodal wake-up-free system based on voice and sight in Example 1 of the present invention;
[0046] Figure 3 This is a schematic diagram of training a multimodal recognition model according to Example 1 of the present invention;
[0047] Figure 4 This is a flowchart of a multimodal wake-up-free method based on voice and sight in Example 2 of the present invention;
[0048] Figure 5 This is a flowchart of step S402 in Example 2 of the present invention;
[0049] Figure 6 This is a flowchart of step S4021 in Example 2 of the present invention;
[0050] Figure 7 This is a flowchart of step S4022 in embodiment 2 of the present invention. DETAILED DESCRIPTION
[0051] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0052] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components; wireless connections or wired connections. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0053] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0054] Example 1
[0055] This embodiment provides a multi-modal wake-up-free system based on voice and sight, such as Figure 1-2 As shown, it includes: an acquisition module 1, a calculation module 2 and a response module 3;
[0056] The acquisition module 1 is used to acquire voice signals and video signals, and transmit the voice signals and video signals to the calculation module 2 .
[0057] Among them, the voice signal is obtained through the voice sensor, and the video information is obtained through the image sensor.
[0058] The calculation module 2 is used to preprocess the voice signal and the video signal, input the preprocessed voice signal and the video signal into the voice and sight multimodal recognition model, and generate a voice recognition result and a sight recognition result.
[0059] Specifically, the multimodal recognition model uses transformer (a deep learning model based on the self-attention mechanism), and the training network composed of transformer encoder network and fully connected network is used to train speech and video signals.
[0060] Furthermore, if Figure 3 As shown in the figure, the training process of the multimodal recognition model is as follows: the pre-collected and stored voice signal is input into the Token Embedding (marker embedding) for encoding, and the encoding is performed into the feature information of each frame (v1, v2...V6); the pre-collected and stored video signal is input into the Token Embedding for encoding, and the encoding is performed into the feature information of each frame (I1, I2...I6), (v1, v2...V6) and (I1, I2...I6) are input into the transformer encoder (position encoding) network, and the encoded feature information is position-encoded based on the time series of the voice signal and the video signal. After the encoder network, the feature information after position encoding is input into FC (fully connected network), and the speech fully connected network FC1 is used to output label1 (speech recognition result), and the sight fully connected network FC2 is used to output label2 (sight recognition result); among them, label1 and label2 can be the speech output vector value and the video output vector value. Then, based on the speech output vector value, the video output vector value and the true value label (for example, "turn on the air conditioner", "turn off the air conditioner", "raise the temperature", "lower the temperature", etc.), the classification loss function is calculated. Based on the classification loss function, the batch gradient descent is used to update the multimodal recognition model parameters until the classification loss function curve converges, and the training of the multimodal recognition model is completed.
[0061] Among them, the classification loss function The calculation formula is as follows:
[0062]
[0063]
[0064] In the above formula, x i represents the output vector value (i.e., the voice output vector value or the video output vector value), i = 1, 2, ... n, j = 1, 2, ... n, and n represents the classifier category (wherein, the voice classifier includes categories such as "turn on the air conditioner", "turn off the air conditioner", "raise the temperature", and "lower the temperature", and the video classifier includes two categories, namely "looking at the device" and "not looking at the device"). represents the exponential function, y i Indicates the true value label corresponding to the command, f(x i) represents the normalized exponential function.
[0065] The response module 3 is used to obtain the voice recognition result and the sight line recognition result, and respond based on the voice recognition result and the sight line recognition result.
[0066] Specifically, if the voice recognition result is consistent with the preset voice command and the line of sight is facing the terminal to be commanded, the terminal responds to the command.
[0067] The above-mentioned multimodal wake-up-free system based on voice and vision collects voice signals and video signals at the same time, and uses the voice and vision multimodal recognition model to recognize the voice signals and video signals, responds to the voice recognition results and the vision recognition results, and uses both the voice recognition results and the vision recognition results as necessary factors for waking up the terminal, thereby avoiding the trouble caused to the user by the false awakening of the device and improving the user experience.
[0068] Preferably, the calculation module 2 includes: a first encoding submodule 21, a second encoding submodule 22 and an identification submodule 23;
[0069] The first encoding submodule 21 is used to extract acoustic features from the speech signal, and encode the acoustic features to generate speech time series features.
[0070] Specifically, the acoustic features are input into Token Embedding for encoding, which is then encoded into acoustic feature information for each frame. The acoustic feature information of each frame is then input into the Transformer Encoder network for position encoding to generate speech time series features.
[0071] The second encoding submodule 22 is used to extract the sight line features in the video signal, and encode the sight line features to generate video timing features.
[0072] Specifically, the line of sight features are input into Token Embedding for encoding, which is then encoded into video feature information for each frame. The video feature information for each frame is then input into the transformer encoder network for position encoding to generate video temporal features.
[0073] The recognition submodule 23 is configured to input the speech timing features and the video timing features into the multimodal recognition model, and output a speech recognition result and a sight line recognition result.
[0074] Preferably, the response module 3 includes:
[0075] Compare the above voice recognition result with the preset voice command (i.e., the command word), and compare the above line of sight recognition result with the preset line of sight command (i.e., whether the line of sight is looking at the terminal). When the above voice result matches the above preset voice command and the above line of sight recognition result matches the above preset line of sight command, respond.
[0076] Specifically, based on the content of the voice signal, determine the marking information (i.e., the preset voice command), which command word or non-command word it is; based on the content of the video signal, determine the marking information (i.e., the preset line of sight command), whether to look at the device, and use the command word and whether to look at the device as a response request, and the device responds accordingly based on the command word.
[0077] For example, when the air conditioner is in use, the user issues a command to the air conditioner, "adjust the temperature to 22 degrees", and looks at the air conditioner, then the air conditioner responds to this command. If the user does not look at the air conditioner, the air conditioner command will not respond. If the user only looks at the air conditioner without a command word, there will be no response.
[0078] Furthermore, when there are multiple similar voice-controlled products in the room, such as voice-controlled lamps and voice-controlled air conditioners, the user's gaze is used to determine which device responds.
[0079] Preferably, the first encoding submodule 21 includes: a processing unit 211, an extraction unit 212 and a first encoding unit 213;
[0080] The processing unit 211 is configured to perform acoustic processing on the speech signal to generate an acoustically processed speech signal.
[0081] Among them, the original speech signal usually contains various noises, which will cause great interference to the speech signal. In order to improve the accuracy of subsequent acoustic feature extraction, the acquired original speech signal needs to be subjected to speech noise reduction processing. Speech noise reduction processing can adopt spectral subtraction: the power spectrum of the pre-estimated noise is subtracted from the noisy speech; or a statistical model-based method is adopted to classify the speech noise reduction problem into a statistical estimation framework, such as Wiener filtering, minimum mean square error (MMSE) method and maximum a posteriori (MAP) method; or based on subspace, it is assumed that the clean speech signal subspace and the noise subspace are orthogonal, thereby performing denoising.
[0082] The extraction unit 212 is configured to extract acoustic features from the acoustically processed speech signal.
[0083] Specifically, fbank (filter bank) and mfcc (Mel-frequency cepstral coefficients) are extracted from the processed speech signal. In this embodiment, the extracted acoustic features can be set as needed.
[0084] Furthermore, the fbank feature extraction process is as follows: the speech signal of indefinite length is divided into small segments of fixed length (i.e., frames), with 10-30ms as a frame. In order to avoid the omission of the signal due to the window boundary, there must be frame overlap when the frame is offset (a part of the frames needs to overlap), and the frame overlap is selected as 10ms; pre-enhancement is performed in frames; and each frame is substituted into the window function (such as square window, Hamming window, etc.), and the value outside the window is set to 0 to eliminate the signal discontinuity that may be caused at both ends of each frame; the time domain signal is converted into a frequency domain signal using Fourier transform; since the frequency domain signal is obtained, the energy of each frequency band range is different, and the energy spectrum of different phonemes is different. The energy in each filter band is superimposed to obtain the energy spectrum output by each filter, and the output power of each filter is taken logarithmically to obtain the logarithmic power spectrum of the corresponding frequency band.
[0085] Furthermore, the MFCC feature extraction process is as follows: the speech signal of indefinite length is divided into small segments of fixed length (i.e., frames), with 10-30ms as a frame. In order to avoid the omission of the signal at the window boundary, when the frame is offset, there must be frame overlap (a part of the frame needs to overlap), and the frame overlap is selected as 10ms; pre-enhancement is performed in units of frames; and each frame is substituted into the window function (such as square window, Hamming window, etc.), and the value outside the window is set to 0 to eliminate the signal discontinuity that may be caused at both ends of each frame; the time domain signal is converted into a frequency domain signal using Fourier transform; since What is received is a frequency domain signal. The energy of each frequency band range is different, and the energy spectra of different phonemes are different. The energy in each filter band is superimposed to obtain the energy spectrum of each filter output; the upper and lower frequency limits are set to shield unnecessary or noisy frequency ranges; the output power of each filter is taken logarithmically to obtain the logarithmic energy spectrum of the corresponding frequency band, and an inverse discrete cosine transform is performed to obtain multiple MFCC coefficients; the MFCC eigenvalues are calculated and used as static features. The first-order and second-order differences of the static features are then performed to obtain the corresponding dynamic features.
[0086] The first encoding unit 213 is used to encode the acoustic features of different time series to generate speech time series features.
[0087] Among them, such as Figure 3 As shown in the figure, the acoustic features of different time series (i.e., fbank and mfcc in the speech signal) are input into Token Embedding for encoding, and encoded into feature information (v1, v2...V6) of each frame; (v1, v2...V6) is input into the transformer Encoder (position encoding) network, and the encoded feature information is positionally encoded based on the time series of the speech signal and video signal to generate speech time series features.
[0088] Preferably, the second encoding submodule 22 includes: a detection unit 221 and a second encoding unit 222;
[0089] The detection unit 221 is configured to perform face detection on the video signal, generate a face image, and extract sight line features from the face image.
[0090] Specifically, each frame of video signal is converted into an image, which can be a picture or a video frame in a video. Then, a face detection algorithm is used to obtain a face image. Facial feature template matching, or MTCNN (Multi-task convolutional neural network), or R-CNN (Region CNN) can be used, and a face feature point positioning algorithm is used to locate facial feature points in the facial area (pre-set facial feature points, such as the eye area and the mouth area). The eye image is extracted according to the feature point position, and the average facial feature point is calculated; the iris precise positioning center point is calculated from the eye image, and the offset vector between the iris precise positioning center point and the average facial feature point is used as the eye movement feature. The gaze point position of the user's line of sight is calculated, and the position of the user's gaze point is used as the line of sight feature.
[0091] Furthermore, the specific calculation process for calculating the gaze point position of the user's line of sight is as follows: grayscale the eye image to obtain the eye grayscale image; binarize the eye grayscale image to obtain the eye binary map; erode and dilate the eye binary map to remove the interference of eyelashes to obtain the iris binary map; in the iris binary map, obtain the longest horizontal distance from left to right, i.e., the longest horizontal line, and the longest vertical distance from top to bottom, i.e., the longest vertical line; and use the intersection of the longest horizontal line and the longest vertical line as the center of the iris coarse positioning; in the eye grayscale image, use the iris as the center of the iris coarse positioning. The center of coarse positioning is the center of the circle, and the star ray method is uniformly diverged outward at every angle. The image gradient is used to calculate the iris edge point on each star ray; based on the iris edge point, the RANSAC algorithm is used to fit the iris ellipse model, and the model parameters are optimized by the least squares method to obtain the optimal iris ellipse. The center of the ellipse is the center point of the iris precise positioning; the offset vector between the average feature point of the face and the center point of the iris precise positioning is calculated to obtain the eye movement feature vector, and the feature vector is used to calculate the user's gaze point position, and the position of the user's gaze point is used as the line of sight feature.
[0092] Furthermore, if there are multiple faces in an image, the largest face image is selected and the image is resized to a certain size (for example, 128*128) to extract the gaze feature. If there is no face in the image, the 128*128 size in the image is randomly selected as the gaze feature.
[0093] The second encoding unit 222 is configured to encode the sight line features of different time series to generate video temporal features.
[0094] Among them, such as Figure 3 As shown in the figure, the above-mentioned sight features of different time series are input into Token Embedding for encoding, and the encoding is performed into the feature information (I1, I2…I6) of each frame. (I1, I2…I6) is input into the transformer encoder (position encoding) network, and the encoded feature information is position-encoded based on the time series of the speech signal and the video signal to generate the video time series features.
[0095] Example 2
[0096] This embodiment provides a multi-modal wake-up-free method based on voice and sight, such as Figure 4 As shown, the following steps are included:
[0097] S401: The acquisition module acquires voice signals and video signals.
[0098] Among them, the voice signal is obtained through the voice sensor, and the video information is obtained through the image sensor.
[0099] S402: The calculation module preprocesses the voice signal and the video signal, inputs the preprocessed voice signal and the video signal into a voice-sight multimodal recognition model, and generates a voice recognition result and a sight recognition result.
[0100] Specifically, the multimodal recognition model uses transformer (a deep learning model based on the self-attention mechanism), and the training network composed of transformer encoder network and fully connected network is used to train speech and video signals.
[0101] Furthermore, if Figure 3As shown in the figure, the training process of the multimodal recognition model is as follows: the pre-collected and stored voice signal is input into the Token Embedding (marker embedding) for encoding, and the encoding is performed into the feature information of each frame (v1, v2...V6); the pre-collected and stored video signal is input into the Token Embedding for encoding, and the encoding is performed into the feature information of each frame (I1, I2...I6), (v1, v2...V6) and (I1, I2...I6) are input into the transformer encoder (position encoding) network, and the encoded feature information is position-encoded based on the time series of the voice signal and the video signal. After the encoder network, the feature information after position encoding is input into FC (fully connected network), and the speech fully connected network FC1 is used to output label1 (speech recognition result), and the sight fully connected network FC2 is used to output label2 (sight recognition result); among them, label1 and label2 can be the speech output vector value and the video output vector value. Then, based on the speech output vector value, the video output vector value and the true value label (for example, "turn on the air conditioner", "turn off the air conditioner", "raise the temperature", "lower the temperature", etc.), the classification loss function is calculated. Based on the classification loss function, the batch gradient descent is used to update the multimodal recognition model parameters until the classification loss function curve converges, and the training of the multimodal recognition model is completed.
[0102] Among them, the classification loss function The calculation formula is as follows:
[0103]
[0104]
[0105] In the above formula, x i represents the output vector value (i.e., the speech output vector value or the video output vector value), i = 1, 2, ... n, j = 1, 2, ... n, n represents the classifier category (wherein, the speech classifier includes categories such as "turn on the air conditioner", "turn off the air conditioner", "increase the temperature", and "lower the temperature", and the video classifier includes two categories, namely "looking at the device" and "not looking at the device"). xi represents the exponential function, y i Indicates the true value label corresponding to the command, f(x i ) represents the normalized exponential function.
[0106] S403: The response module responds based on the voice recognition result and the sight line recognition result.
[0107] Specifically, if the voice recognition result is consistent with the preset voice command and the line of sight is facing the terminal to be commanded, the terminal responds to the command.
[0108] The above-mentioned multimodal wake-up-free method based on voice and vision simultaneously collects voice signals and video signals, and uses the voice and vision multimodal recognition model to recognize the voice signals and video signals, responds to the voice recognition results and the vision recognition results, and uses both the voice recognition results and the vision recognition results as necessary factors for waking up the terminal, thereby avoiding the trouble caused to the user by the false awakening of the device and improving the user experience.
[0109] Preferably, if Figure 5 As shown, in the above step S402, the above voice signal and the above video signal are preprocessed, and the preprocessed voice signal and the above video signal are input into the voice-gaze multimodal recognition model to generate a voice recognition result and a gaze recognition result, including:
[0110] S4021. The first encoding submodule extracts acoustic features from the above-mentioned speech signal, and encodes the above-mentioned acoustic features to generate speech time series features.
[0111] Specifically, the acoustic features are input into Token Embedding for encoding, which is then encoded into acoustic feature information for each frame. The acoustic feature information of each frame is then input into the Transformer Encoder network for position encoding to generate speech time series features.
[0112] S4022: The second encoding submodule extracts the sight line features from the video signal and encodes the sight line features to generate video timing features.
[0113] S4023. The recognition submodule inputs the above-mentioned speech timing features and the above-mentioned video timing features into the above-mentioned multimodal recognition model, and outputs the speech recognition results and the line of sight recognition results.
[0114] Preferably, responding based on the voice recognition result and the sight line recognition result in step S403 includes:
[0115] The above-mentioned voice recognition result is compared with the preset voice command, and the above-mentioned sight line recognition result is compared with the preset sight line command, and a response is performed when the above-mentioned voice result is consistent with the above-mentioned preset voice command and the above-mentioned sight line recognition result is consistent with the above-mentioned preset sight line command.
[0116] Specifically, based on the content of the voice signal, determine the marking information (i.e., the preset voice command), which command word or non-command word it is; based on the content of the video signal, determine the marking information (i.e., the preset sight line command), whether to look at the device.
[0117] For example, when the air conditioner is in use, the user issues a command to the air conditioner, "adjust the temperature to 22 degrees", and looks at the air conditioner, then the air conditioner responds to this command. If the user does not look at the air conditioner, the air conditioner command will not respond. If the user only looks at the air conditioner without a command word, there will be no response.
[0118] Furthermore, when there are multiple similar voice-controlled products in the room, such as voice-controlled lamps and voice-controlled air conditioners, the user's gaze is used to determine which device responds.
[0119] Preferably, if Figure 6 As shown, in step S4021, the extraction of acoustic features from the speech signal and encoding of the acoustic features to generate speech time series features include:
[0120] S40211. The processing unit performs acoustic processing on the above-mentioned voice signal to generate an acoustically processed voice signal.
[0121] Among them, the original speech signal usually contains various noises, which will cause great interference to the speech signal. In order to improve the accuracy of subsequent acoustic feature extraction, the acquired original speech signal needs to be subjected to speech noise reduction processing. Speech noise reduction processing can adopt spectral subtraction: the power spectrum of the pre-estimated noise is subtracted from the noisy speech; or a statistical model-based method is adopted to classify the speech noise reduction problem into a statistical estimation framework, such as Wiener filtering, minimum mean square error (MMSE) method and maximum a posteriori (MAP) method; or based on subspace, it is assumed that the clean speech signal subspace and the noise subspace are orthogonal, thereby performing denoising.
[0122] S40212. The extraction unit extracts acoustic features from the speech signal after the acoustic processing.
[0123] Specifically, fbank (filter bank) and mfcc (Mel-frequency cepstral coefficients) are extracted from the processed speech signal. In this embodiment, the extracted acoustic features can be set as needed.
[0124] Furthermore, the fbank feature extraction process is as follows: the speech signal of indefinite length is divided into small segments of fixed length (i.e., frames), with 10-30ms as a frame. In order to avoid the omission of the signal due to the window boundary, there must be frame overlap when the frame is offset (a part of the frames needs to overlap), and the frame overlap is selected as 10ms; pre-enhancement is performed in frames; and each frame is substituted into the window function (such as square window, Hamming window, etc.), and the value outside the window is set to 0 to eliminate the signal discontinuity that may be caused at both ends of each frame; the time domain signal is converted into a frequency domain signal using Fourier transform; since the frequency domain signal is obtained, the energy of each frequency band range is different, and the energy spectrum of different phonemes is different. The energy in each filter band is superimposed to obtain the energy spectrum output by each filter, and the output power of each filter is taken logarithmically to obtain the logarithmic power spectrum of the corresponding frequency band.
[0125] Furthermore, the MFCC feature extraction process is as follows: the speech signal of indefinite length is divided into small segments of fixed length (i.e., frames), with 10-30ms as a frame. In order to avoid the omission of the signal at the window boundary, when the frame is offset, there must be frame overlap (a part of the frame needs to overlap), and the frame overlap is selected as 10ms; pre-enhancement is performed in units of frames; and each frame is substituted into the window function (such as square window, Hamming window, etc.), and the value outside the window is set to 0 to eliminate the signal discontinuity that may be caused at both ends of each frame; the time domain signal is converted into a frequency domain signal using Fourier transform; since What is received is a frequency domain signal. The energy of each frequency band range is different, and the energy spectra of different phonemes are different. The energy in each filter band is superimposed to obtain the energy spectrum of each filter output; the upper and lower frequency limits are set to shield unnecessary or noisy frequency ranges; the output power of each filter is taken logarithmically to obtain the logarithmic energy spectrum of the corresponding frequency band, and an inverse discrete cosine transform is performed to obtain multiple MFCC coefficients; the MFCC eigenvalues are calculated and used as static features. The first-order and second-order differences of the static features are then performed to obtain the corresponding dynamic features.
[0126] S40213. The first encoding unit encodes the acoustic features of different time series to generate speech time series features.
[0127] Among them, the acoustic features of different time series (i.e., fbank and mfcc in the speech signal) are input into TokenEmbedding (token embedding) for encoding, and encoded into feature information (v1, v2...V6) of each frame; (v1, v2...V6) is input into the transformer Encoder (position encoding) network, and the encoded feature information is positionally encoded based on the time series of the speech signal and video signal to generate speech time series features.
[0128] Preferably, if Figure 7As shown, in step S4022, the extraction of sight features from the video signal and encoding of the sight features to generate video timing features include:
[0129] S40221. The detection unit performs face detection on the video signal, generates a face image, and extracts sight features from the face image.
[0130] Specifically, each frame of video signal is converted into an image, which can be a picture or a video frame in a video. Then, a face detection algorithm is used to obtain a face image. Facial feature template matching, or MTCNN (Multi-task convolutional neural network), or R-CNN (Region CNN) can be used, and a face feature point positioning algorithm is used to locate facial feature points in the facial area (pre-set facial feature points, such as the eye area and the mouth area). The eye image is extracted according to the feature point position, and the average facial feature point is calculated; the iris precise positioning center point is calculated from the eye image, and the offset vector between the iris precise positioning center point and the average facial feature point is used as the eye movement feature, the gaze point position of the user's line of sight is calculated, and the position of the user's gaze point is used as the line of sight feature.
[0131] Furthermore, the specific calculation process for calculating the gaze point position of the user's line of sight is as follows: grayscale the eye image to obtain the eye grayscale image; binarize the eye grayscale image to obtain the eye binary map; erode and dilate the eye binary map to remove the interference of eyelashes to obtain the iris binary map; in the iris binary map, obtain the longest horizontal distance from left to right, i.e., the longest horizontal line, and the longest vertical distance from top to bottom, i.e., the longest vertical line; and use the intersection of the longest horizontal line and the longest vertical line as the center of the iris coarse positioning; in the eye grayscale image, use the iris as the center of the iris coarse positioning. The center of coarse positioning is the center of the circle, and the star ray method is uniformly diverged outward at every angle. The image gradient is used to calculate the iris edge point on each star ray; based on the iris edge point, the RANSAC algorithm is used to fit the iris ellipse model, and the model parameters are optimized by the least squares method to obtain the optimal iris ellipse. The center of the ellipse is the center point of the iris precise positioning; the offset vector between the average feature point of the face and the center point of the iris precise positioning is calculated to obtain the eye movement feature vector, and the feature vector is used to calculate the user's gaze point position, and the position of the user's gaze point is used as the line of sight feature.
[0132] Furthermore, if there are multiple faces in an image, the largest face image is selected and the image is resized to a certain size (for example, 128*128) to extract the gaze feature. If there is no face in the image, the 128*128 size in the image is randomly selected as the gaze feature.
[0133] S40222. The second encoding unit encodes the above-mentioned sight line features of different time series to generate video timing features.
[0134] The above-mentioned line of sight features of different time series are input into Token Embedding for encoding, and are encoded into feature information (I1, I2…I6) of each frame. (I1, I2…I6) is input into the transformer encoder (position encoding) network, and the encoded feature information is position-encoded based on the time series of the speech signal and video signal to generate video temporal features.
[0135] Example 3
[0136] This embodiment provides a computer device, including a memory and a processor, wherein the processor is configured to read instructions stored in the memory to perform the following operations:
[0137] Acquire a command data set, perform corpus processing on the command data set, and generate syntactic templates and similar sentence pairs;
[0138] Training a preset text generation model based on the syntactic template and the similar sentence pairs to generate a similar text generation model;
[0139] The command data set and the syntactic template are input into the above-mentioned text generation model to generate similar command texts.
[0140] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0142] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0144] Example 4
[0145] This embodiment provides a computer-readable storage medium storing computer-executable instructions capable of executing a method for generating similar command text in any of the above-mentioned method embodiments. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the storage medium may also include a combination of the above-mentioned types of memory.
[0146] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A multimodal wake-up-free system based on voice and sight, characterized in that: include: Acquisition module, calculation module and response module; The acquisition module is used to collect voice signals and video signals, and transmit the voice signals and video signals to the calculation module; The computing module is used to preprocess the voice signal and the video signal, input the preprocessed voice signal and the video signal into a multimodal recognition model, and generate a voice recognition result and a sight line recognition result; The response module is used to obtain the voice recognition result and the sight line recognition result, and respond based on the voice recognition result and the sight line recognition result; The calculation module includes: a first encoding submodule, a second encoding submodule and an identification submodule; The first encoding submodule is used to extract acoustic features from the speech signal and encode the acoustic features to generate speech time series features; The second encoding submodule is used to extract the sight line features in the video signal and encode the sight line features to generate video timing features; The recognition submodule is used to input the speech timing features and the video timing features into the multimodal recognition model, and output the speech recognition results and the line of sight recognition results; wherein, the training process of training the multimodal recognition model is: inputting the pre-collected and stored speech signal into Token Embedding for encoding, and encoding it into the feature information (v1, v2...V6) of each frame; inputting the pre-collected and stored video signal into Token Embedding for encoding, and encoding it into the feature information (I1, I2...I6) of each frame, and inputting (v1, v2...V6) and (I1, I2...I6) into the transformer encoder network, and performing position encoding on the encoded feature information based on the time series of the speech signal and the video signal, and passing through the transformer. After the encoder network, the position-encoded feature information is input into FC. The speech fully connected network FC1 is used to output label1, and the sight fully connected network FC2 is used to output label2, where label1 and label2 can be the speech output vector value and the video output vector value. The classification loss function is calculated based on the speech output vector value, the video output vector value and the true value label. Based on the classification loss function, batch gradient descent is used to update the multimodal recognition model parameters until the classification loss function curve converges, completing the training of the multimodal recognition model. The response module is specifically used to: The voice recognition result is compared with a preset voice command, and the sight line recognition result is compared with a preset sight line command, and a response is performed when the voice result matches the preset voice command and the sight line recognition result matches the preset sight line command.
2. The multimodal wake-up-free system based on voice and sight according to claim 1, characterized in that: The first encoding submodule includes: a processing unit, an extraction unit and a first encoding unit; The processing unit is used to perform acoustic processing on the speech signal to generate an acoustically processed speech signal; The extraction unit is used to extract acoustic features from the acoustically processed speech signal; The first encoding unit is used to encode the acoustic features of different time series to generate speech time series features.
3. The multimodal wake-up-free system based on voice and sight according to claim 1, characterized in that: The second encoding submodule includes: a detection unit and a second encoding unit; The detection unit is used to perform face detection on the video signal, generate a face image, and extract sight features from the face image; The second encoding unit is used to encode the sight features of different time series to generate video timing features.
4. A multimodal wake-up-free method based on voice and sight, characterized in that: The steps include: Collect voice and video signals; Preprocessing the voice signal and the video signal, inputting the preprocessed voice signal and the video signal into a voice-gaze multimodal recognition model to generate a voice recognition result and a gaze recognition result; Responding based on the voice recognition result and the sight line recognition result; The preprocessing of the voice signal and the video signal, inputting the preprocessed voice signal and the video signal into a multimodal recognition model, and generating a voice recognition result and a sight line recognition result includes: Extracting acoustic features from the speech signal and encoding the acoustic features to generate speech time series features; Extracting sight features from the video signal and encoding the sight features to generate video timing features; The speech timing features and the video timing features are input into the multimodal recognition model, and the speech recognition results and the line of sight recognition results are output; wherein, the training process of the multimodal recognition model is as follows: the pre-collected and stored speech signal is input into the Token Embedding for encoding, and the encoding is performed into the feature information (v1, v2…V6) of each frame; the pre-collected and stored video signal is input into the Token Embedding for encoding, and the encoding is performed into the feature information (I1, I2…I6) of each frame, (v1, v2…V6) and (I1, I2…I6) are input into the transformer Encoder network, and the encoded feature information is position-encoded based on the time series of the speech signal and the video signal, and the feature information is passed through the transformer. After the encoder network, the position-encoded feature information is input into FC. The speech fully connected network FC1 is used to output label1, and the sight fully connected network FC2 is used to output label2, where label1 and label2 can be the speech output vector value and the video output vector value. The classification loss function is calculated based on the speech output vector value, the video output vector value and the true value label. Based on the classification loss function, batch gradient descent is used to update the multimodal recognition model parameters until the classification loss function curve converges, completing the training of the multimodal recognition model. The responding based on the voice recognition result and the sight line recognition result includes: The voice recognition result is compared with a preset voice command, and the sight line recognition result is compared with a preset sight line command, and a response is performed when the voice result matches the preset voice command and the sight line recognition result matches the preset sight line command.
5. A computer device, characterized in that: The system comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is configured to call the computer program to execute the steps of the method according to claim 4.
6. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to claim 4 are implemented.
Citation Information
Patent Citations
Multi-task classification model training method and device and multi-task classification method and device
CN110728298A
Multi-mode voice wake-up method and device and computer readable storage medium
CN114220420A