Voice instruction accurate identification method and system suitable for intelligent mirror

By constructing a multimodal speech recognition model, combining weather forecast and image information, and using weighted fusion loss function training model, the problem of inaccurate speech recognition on the smart mirror and long time is solved, and a fast and accurate dressing plan generation is achieved.

CN120279913AActive Publication Date: 2025-07-08博洛尼智能科技(青岛)有限公司
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510764378.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing smart mirror speech recognition method cannot accurately identify the user's body shape, posture, current dress, designated dress parts and weather conditions, resulting in low speech recognition accuracy and inaccurate dressing schemes of output. The existing method requires cascading multiple models, resulting in long time and large errors.

Method used

A multimodal speech recognition model is constructed, speech features are extracted through the first encoder, combined with weather forecast and image information, and feature fusion is fusion using the second and third encoders, and the model is trained using a weighted fusion loss function to shorten the recognition time and improve accuracy.

Benefits of technology

It realizes fast and accurate speech recognition on the smart mirror, generates the dressing scheme required by the user, avoids the problems of error transmission and excessive time, and improves the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279913A_ABST
    Figure CN120279913A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, in particular to a voice instruction accurate recognition method and system suitable for an intelligent mirror, and the method comprises the steps: a second encoder of a voice recognition model outputs a simplified instruction recognition result, and a third encoder outputs a complete instruction recognition result; the parameter quantity of the third encoder is smaller than that of the first encoder; after each voice sample is input into the voice recognition model, obtaining a first difference between a complete instruction recognition result and a dressing label, and obtaining a second difference between the complete instruction recognition result and a simplified instruction recognition result; and carrying out weighted fusion on the second difference and the first difference to obtain a loss function, and training a speech recognition model and carrying out speech recognition by using the data set and the loss function. According to the invention, the speech recognition time is shortened, and the recognition accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for accurate voice command recognition applicable to intelligent mirrors. Background Art

[0002] An intelligent mirror is a mirror that uses artificial intelligence technology to interact with users and provide intelligent feedback. For example, users can generate dressing and makeup plans in a specified scenario by looking at the mirror and having voice interactions, thereby improving the user experience. The existing implementation method is to use a microphone to collect user voices and use a voice recognition model (such as ESPnet, WeNet, etc.) to obtain dressing and makeup plans; however, this method cannot perform voice recognition based on the current user's body shape, posture, current clothing, and specified clothing parts (such as upper body clothes, lower body clothes, shoes and socks), as well as weather conditions, resulting in low voice recognition accuracy and the output dressing and makeup plans not meeting the user's needs.

[0003] An easily conceivable solution is to use a large model for image-to-text and text-to-speech conversion to convert the user's body image and weather forecast into voice signals, and output these voice signals together with the user voice to a voice recognition model to obtain dressing and makeup plans. Although this method can further accurately output the dressing and makeup plans required by the user, it requires cascading and stacking multiple artificial intelligence models, resulting in long voice recognition time and obvious latency. Moreover, due to error transmission (such as transmission model hallucinations) between models and obvious error amplification problems, some dressing and makeup plans may be significantly incorrect or unreasonable.

[0004] Therefore, when users generate dressing and makeup in a specified scenario by looking at the mirror and having voice interactions, there is a lack of a fast and accurate voice recognition method. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method and system for accurate voice command recognition applicable to intelligent mirrors.

[0006] The method and system for accurate voice command recognition applicable to intelligent mirrors of the present invention adopt the following technical solutions: An embodiment of the present invention provides a method for accurate voice command recognition applicable to intelligent mirrors, the method comprising the following steps: The constructed voice recognition model includes: inputting the collected voice into a first encoder, the output of the first encoder being respectively used as the input of a second encoder and a third encoder, the second encoder outputting a refined instruction recognition result; using a fourth encoder to encode the weather forecast and image collected under the time instruction and body part instruction contained in the voice, and also using it as the input of the third encoder; the third encoder outputting a complete instruction recognition result; the number of parameters of the third encoder being less than that of the first encoder; The constructed dataset contains a number of speech samples and the dressed-up labels corresponding to the speech samples; after each speech sample is input into the speech recognition model, the first difference between the complete instruction recognition result and the dressed-up label is obtained, and the second difference between the complete instruction recognition result and the streamlined instruction recognition result is obtained; the second difference and the first difference are weighted and fused to obtain a loss function, and the speech recognition model is trained using the dataset and the loss function and speech recognition is performed. Among them, when weighted and fused, the weight of the second difference is positively correlated with the first difference. When the weight obtained after the same speech sample is input into the speech recognition model twice in succession increases, the change in the word vector in the complete instruction recognition result is recorded as the attention coefficient of the word vector, and the first difference is obtained from the attention coefficient.

[0007] Preferably, the specific steps for obtaining the first difference are as follows: Both the complete instruction recognition result and the dressed-up label are vector sequences, and the vector sequences contain a number of word vectors. Obtain the Euclidean distance of the word vectors in the same order in the complete instruction recognition result and the dressed-up label, and use the attention coefficient of each word vector in the complete instruction recognition result to perform weighted summation on the Euclidean distances of all word vectors to obtain the first difference.

[0008] Preferably, the specific steps for the second difference are as follows: Both the complete instruction recognition result and the streamlined instruction recognition result are vector sequences, and the vector sequences contain word vectors. Obtain the Euclidean distance of the word vectors in the same order in the complete instruction recognition result and the streamlined instruction recognition result; the mean value of the Euclidean distances of all word vectors in the complete instruction recognition result is recorded as the second difference.

[0009] Preferably, when the weight obtained after the same speech sample is input into the speech recognition model twice in succession increases, the change in the word vector in the complete instruction recognition result is recorded as the attention coefficient of the word vector, and the specific steps include the following: Each time a speech sample is input into the speech recognition model, it means that each speech sample is trained once. For the same speech sample, and at the current training of this speech sample, obtain the weight obtained at the first training before the current training, denoted as , and the complete instruction recognition result obtained is denoted as ; Obtain the weight obtained at the second training before the current training of this speech sample, denoted as , and the complete instruction recognition result obtained is denoted as ; and The i-th word vector in is denoted as 、 ; The sum of the attention coefficients and is positively correlated with the difference between, and negatively correlated with the similarity between 、 .

[0010] Preferably, the loss function obtained by weighted fusion of the second difference and the first difference includes the following specific formula: where S represents the loss function, represents the first difference between the complete instruction recognition result R1 and the dressing label R, represents the second difference between the complete instruction recognition result R1 and the refined instruction recognition result R2; represents the weight of the second difference, represents the loss function of weather forecast and image.

[0011] Preferably, the specific steps for obtaining the loss function of the weather forecast and image are as follows: Each speech sample contains a time instruction and a body part instruction; Mark the body part represented by the body part instruction in each speech sample, denoted as the body part label m2; mark the time represented by the time instruction in each speech sample, denoted as the time label m1; After inputting each speech sample in the dataset into the speech recognition model, the first encoder outputs the time, denoted as M1, and at the same time the first encoder outputs the body part, denoted as M2; where represents the Euclidean distance between M1 and m1, represents the function value of the cross-entropy loss function between M2 and m2.

[0012] Preferably, the steps for encoding the weather forecast and image collected under the time instruction and body part instruction in the speech by using the fourth encoder and also using it as the input of the third encoder are as follows: The speech contains a time instruction and a body part instruction; After the voice is input into the first encoder, the first encoder outputs the time described by the time instruction and the body part described by the body part instruction; then, the weather forecast is performed using the time to obtain a weather voice instruction, and a full-body image of the user is collected. Then, the weather voice instruction and the body part in the full-body image are input into the fourth encoder to obtain a voice-assisted recognition feature, and the voice-assisted recognition feature is used as the input of the third encoder.

[0013] Preferably, the performing voice recognition includes the following specific steps: After the voice recognition model is trained, the second encoder is deleted. The user stands in front of the smart mirror, and the smart mirror collects the user's voice and inputs the voice into the trained voice recognition model. The first encoder outputs the time and body part in the voice, and the first feature map; then, the weather voice instruction at the time is obtained, and at the same time, a full-body image of the user is collected. The weather voice instruction and the body part in the full-body image are input into the fourth encoder, and a voice-assisted recognition feature is obtained. Then, the first feature map and the voice-assisted recognition feature are stacked together and input into the third encoder to obtain a complete instruction recognition result. All the vocabulary corresponding to the word vectors in the complete instruction recognition result is used as the dressing plan required by the user.

[0014] Preferably, the weight of the second difference is positively correlated with the first difference during the weighted fusion, and the specific formula is as follows: where L0 represents a preset normalization coefficient, w represents the weight of the second difference, represents the first difference between the complete instruction recognition result R1 and the dressing label R.

[0015] Another embodiment of the present invention provides a voice instruction precise recognition system applicable to a smart mirror. The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned voice instruction precise recognition method applicable to a smart mirror are implemented.

[0016] The beneficial effects of the technical solution of the present invention are: The multimodal speech recognition model of the present invention is trained and used as a whole, initially avoiding the problems of overly long speech recognition time and inaccurate speech recognition caused by error propagation. Additionally, the present invention allows changing the number of parameters of the third encoder and the fourth encoder to further shorten the speech recognition time. To further shorten the speech recognition time while ensuring the accuracy of speech recognition, the present invention first obtains the first difference between the complete instruction recognition result and the dressing label, and then obtains the second difference between the complete instruction recognition result and the simplified instruction recognition result. The first difference and the second difference are fused, and the weight during the fusion of the second difference is positively correlated with the first difference. When there is a large training error (i.e., when the first difference is large), this process allows the simplified instruction recognition result to be closer to the complete instruction recognition result. Since the simplified instruction recognition result skips the feature extraction process of the image and weather and obtains the instruction recognition result in advance, this enables the first encoder to learn and extract some speech features in advance, avoiding the situation where when most speech features and all image and weather features (i.e., speech-assisted recognition features) are fused in the third encoder with a relatively small number of parameters, the feature extraction or feature learning ability for speech, image, and weather is insufficient, resulting in a large error. This makes the speech recognition result of this embodiment more accurate (i.e., the generated dressing plan is more reliable).

[0017] Furthermore, when the weight obtained after the same speech sample is input into the speech recognition model twice in succession increases, the change in the word vector in the complete instruction recognition result is recorded as the attention coefficient of the word vector, and the first difference is obtained from the attention coefficient. After enabling the first encoder to learn and extract some speech features in advance, to a certain extent, it avoids the situation where the speech features in the third encoder are submerged in the speech-assisted recognition features, resulting in the third encoder being unable to fuse and learn the speech features and the speech-assisted recognition features. This further improves the accuracy of speech recognition while reducing the speech recognition time.

[0018] In summary, when the user generates the dressing in a specified scenario through mirroring and voice interaction in the present invention, a fast and accurate speech recognition method is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1The flowchart of the steps of the voice command accurate recognition method applicable to an intelligent mirror provided by an embodiment of the present invention; Figure 2 It shows a schematic structural diagram of a multi-modal voice recognition model provided by an embodiment of the present invention. Detailed implementation manners

[0021] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and effects of the voice command accurate recognition method and system applicable to an intelligent mirror proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0023] The following specifically describes the specific solutions of the voice command accurate recognition method and system applicable to an intelligent mirror provided by the present invention with reference to the accompanying drawings.

[0024] Comparative example: Install a voice acquisition module (such as a microphone) on the intelligent mirror to collect the user's voice, which contains time commands, body part commands, and scene commands. The time command and the scene command represent the scene at the corresponding time, and the body part command represents the body part to be dressed up; for example, "What pants should I wear when going out shopping with friends this afternoon? How should I apply makeup?", "What clothes should I wear on the upper body when starting work tomorrow morning?", in these voices, "this afternoon" and "tomorrow morning" are time commands, "shopping" and "going to work" are scene commands, and "what pants to wear", "applying makeup", and "upper body" represent body part commands.

[0025] One implementation method is to input the voice into an existing large voice model (such as in ESPnet) to generate a dressing-up plan.

[0026] However, this implementation method cannot obtain the weather information under the time command, as well as the body shape, posture, and existing dressing information under the body part command, resulting in the generated dressing-up plan not being what the user needs.

[0027] Therefore, in this embodiment, a weather forecast module is installed on the smart mirror. This module uses a speech large model (such as ESPnet) to recognize the time in the speech, and then obtains the weather forecast at that time (the weather broadcast in speech format, briefly recorded as the weather speech instruction). In addition, a camera is installed on the smart mirror to collect the user's full-body image; then, a speech large model (such as ESPnet) is used to recognize the body part in the speech, and an image description generation model (such as the BLIP / BLIP-2 model) is used to generate the description speech of the body part in the full-body image, briefly recorded as the body part speech instruction.

[0028] The user's speech, along with the above-mentioned weather speech instruction and body part speech instruction, is concatenated together and then input into the speech large model, thereby generating a dressing plan.

[0029] Specifically, when the time instruction and body part instruction are not recognized in the user's collected speech, the time and body part are defaulted to "today" and "upper body".

[0030] In the above process, the speech large model, image description generation model, etc. are all well-known technologies and there are already mature existing applications. This embodiment will not elaborate on them one by one. The weather forecast module is also an existing technology and will not be elaborated in this embodiment.

[0031] The problems existing in this comparative embodiment are as follows: on the one hand, it is necessary to cascade and stack multiple models (or cascade the same model multiple times), resulting in too long speech recognition time. On the other hand, due to the error transmission between models (such as transmission model hallucinations) and obvious error amplification problems, the speech instruction recognition result may be significantly incorrect or unreasonable.

[0032] Embodiment 1: As Figure 1 shown, this embodiment provides a method for accurately recognizing speech instructions applicable to a smart mirror, including: Step S101, the constructed speech recognition model includes: a first encoder, a second encoder, a third encoder, and a fourth encoder. The second encoder outputs a refined instruction recognition result; the third encoder outputs a complete instruction recognition result.

[0033] In this implementation, a multi-modal speech recognition model is constructed, which includes a first encoder, a second encoder, a third encoder, and a fourth encoder. As Figure 2 shown, the specific principle is as follows: (1) Collect the user's speech. This speech is input into the first encoder, and the first encoder extracts features from the speech. The feature map output by the first encoder is denoted as the first feature map; the first feature map is input into the second encoder, and the second encoder further extracts features from the speech and outputs a dressing plan, denoted as the refined instruction recognition result.

[0034] The recognition result of the concise instructions represents the dressing plan obtained when not relying on weather instructions and body part instructions.

[0035] (2) The first encoder also outputs the time and body parts represented by the weather instructions and body part instructions in the speech.

[0036] After the time and body parts are output, the weather forecast and image acquisition functions are triggered to predict the weather speech (i.e., the weather speech instruction) at that time and collect the user's full-body image. Then, the weather speech instruction and the body parts in the full-body image are input into the fourth encoder for encoding compression (or feature extraction) to obtain the speech-assisted recognition feature (which is essentially also a feature map).

[0037] (3) Input the first feature map and the speech-assisted recognition feature into the third encoder together; for example, make the speech-assisted recognition feature and the first feature map have the same size, and stack the speech-assisted recognition feature and the first feature map together and then input them into the third encoder.

[0038] The third encoder outputs the complete instruction recognition result, which represents the dressing plan generated considering the weather, body type, and existing clothing. This dressing plan is the dressing plan that the smart mirror in this embodiment needs to generate and that the user requires.

[0039] In this step, multiple models used in the comparative example are combined into a multi-modal speech recognition model, and this multi-modal speech recognition model is trained and used as a whole, initially avoiding the problem of too long speech recognition time and inaccurate speech recognition caused by error transmission.

[0040] As an example, the first encoder, the second encoder, the third encoder, and the fourth encoder all use the Transformer network. The Transformer network uses a multi-layer self-attention mechanism to achieve in-depth feature learning or feature extraction of speech or image data, which is a well-known technology. For example, the Transformer network is used in ESPnet. Each layer of the self-attention mechanism in the first encoder, the second encoder, the third encoder, and the fourth encoder in this embodiment can use the self-attention mechanism in ESPnet; in addition, there are application examples of the Transformer network in the GPT4 large model and the BLIP / BLIP-2 model, which can be used in this embodiment. Here, the specific network structures of the first encoder, the second encoder, the third encoder, and the fourth encoder are not specifically described in this embodiment.

[0041] Furthermore, this embodiment takes into account that after the multi-modal speech recognition model receives voice input and obtains the first feature map, the third encoder cannot immediately output the complete instruction recognition result; instead, it is necessary to wait for the completion of processes such as weather collection, image collection, and the processing of the fourth encoder, which results in a relatively long speech recognition time. To further reduce the speech recognition time, the number of parameters of the fourth encoder and the third encoder in this embodiment cannot be set too large.

[0042] As an example, the Transformer network in the first encoder includes 5 layers of self-attention mechanisms, and the Transformer networks of the fourth encoder and the third encoder both include 2 layers of self-attention mechanisms, and the second encoder includes 1 layer of self-attention mechanism.

[0043] In other examples, the number of parameters of the fourth encoder and the third encoder can be set by other methods, so that the number of parameters of the fourth encoder and the third encoder is relatively small, especially ensuring that the number of parameters of the third encoder is less than that of the first encoder.

[0044] Step S102: The constructed data set includes several speech samples and the dressed-up labels corresponding to the speech samples are marked.

[0045] Collect the voices of different users, denoted as speech samples, and each speech sample includes a time instruction, a body part instruction, and a scene instruction.

[0046] Collect the full-body color images of the users. For the body parts described by the body part instructions, only the pixel points in the body parts in the full-body images are retained (the rgb values of the pixel points in the areas outside the body parts are all set to 0), and this full-body image is denoted as the full-body image of the speech sample. Randomly obtain a weather report (such as including temperature, humidity, weather, etc.) at a moment from the historical weather forecast, denoted as the weather voice instruction.

[0047] Using the method of the above comparative embodiment, generate body part voice instructions according to the full-body image, generate a dressed-up plan (in text format) according to the speech sample, the weather voice instruction, and the body part voice instruction, and use the word2vec technology to generate a word vector for each word (including punctuation marks) in the dressed-up plan (in this embodiment, all word vectors are unitized into unit vectors with a dimension of 128). The word vectors of all words in the dressed-up plan form a vector sequence, denoted as the dressed-up label of the speech sample.

[0048] It should be noted that, in order to unify the length of the vector sequence, the length of the vector sequence is set to 1000 in this embodiment. If all the words in the dressing plan are less than 1000, the positions in the vector sequence without word vectors are filled with zero vectors of 128 dimensions (this zero vector represents an empty character that does not need to be displayed).

[0049] So far, the data set contains several voice samples, and each voice sample corresponds to a weather voice command, a full-body image (a full-body image that only retains the pixel points in the body part), and a dressing label.

[0050] In some other embodiments, the voice samples in the data set are manually screened. For example, the voice samples with obvious dressing plan errors are deleted, or the dressing plan is fine-tuned and modified according to the user's existing clothes.

[0051] Step S103: After each voice sample is input into the speech recognition model, obtain the first difference between the complete instruction recognition result and the dressing label, and obtain the second difference between the complete instruction recognition result and the streamlined instruction recognition result; fuse the second difference and the first difference with weights to obtain a loss function, and use the data set and the loss function to train the speech recognition model.

[0052] The training process includes: (1) Traverse all the voice samples in the data set. For the currently traversed voice sample, input the voice sample into the multi-modal speech recognition model (that is, the first encoder), input the weather voice command and the full-body image corresponding to the voice sample into the fourth encoder, and then obtain the streamlined instruction recognition result (denoted as R2) and the complete instruction recognition result (denoted as R1) output by the multi-modal speech recognition model.

[0053] It should be noted that the streamlined instruction recognition result or the complete instruction recognition result in this embodiment is also a vector sequence. Each vector in this vector sequence is a word vector, and each word vector represents a word. The words corresponding to all the word vectors in a vector sequence describe the dressing plan required by the user.

[0054] (2) Obtain the loss function S according to the difference between the complete instruction recognition result R1 and the dressing label of each voice sample.

[0055] As a comparative example, the formula for obtaining the loss function S is as follows: where represents the loss function composed of the weather forecast and the body part corresponding to the time instruction and the body part instruction in the voice sample, briefly recorded as the loss function of the weather forecast and the image. R represents the vector sequence corresponding to the dressing label of each voice sample. Indicates the difference between R1 and R.

[0056] As an example, , where represents the i-th word vector in R, represents the i-th word vector in R1; N represents the number of word vectors in R (in this embodiment, N = 1000), represents and of the Euclidean distance.

[0057] Considering that in this embodiment, in order to further reduce the time of speech recognition, it is necessary to ensure that the number of parameters of the third encoder is relatively small (as described in step S101 in detail), but this will lead to insufficient ability of the third encoder to extract or learn features of speech. Especially when the third encoder needs to fuse the first feature map and speech auxiliary recognition features at the same time, while reducing the time of speech recognition, directly using the loss function S provided by the above comparison example cannot guarantee the accuracy of speech recognition.

[0058] As a preferred example, the acquisition formula of the loss function S is as follows: Obtain the difference between R1 and the dressing label R of each speech sample, denoted as the first difference, expressed as , and at the same time obtain the difference between R1 and R2, denoted as the second difference . Perform weighted fusion on the first difference and the second difference, and the weight of the second difference is positively correlated with the first difference during weighted fusion.

[0059] That is . Where w represents the weight of the second difference.

[0060] In this process, when the difference between R1 and the dressing label R of each speech sample (that is, the first difference ) is large, it means that the recognition error of the multi-modal speech recognition model for each speech sample is large, or in other words, the multi-modal speech recognition model does not deeply learn or extract the accurate features of each speech sample. The reason may be that in order to further reduce the time of speech recognition, the number of parameters of the third encoder is relatively small, resulting in insufficient ability of the third encoder to extract or learn features of speech, images, and weather when fusing the first feature map and speech auxiliary recognition features.

[0061] The calculation formula of is the same as that of.

[0062] At this time (when the first difference is large), w is large, that is, pay more attention to the size of the second difference . In this embodiment, by paying more The size allows the result of the reduced instruction recognition to be closer to the result of the complete instruction recognition. Since the result of the reduced instruction recognition skips the feature extraction process of the image and weather, and obtains the result in advance (i.e., the dressing plan that the user may need), this enables the first encoder to learn and extract some speech features in advance, avoiding the situation where when most speech features and all image and weather features (i.e., speech-assisted recognition features) are fused in the third encoder with fewer parameters, the feature extraction or feature learning ability of speech, image, and weather is insufficient, resulting in a large error.

[0063] As an example, when performing weighted fusion, the weight of the second difference is positively correlated with the first difference, and the formula included is: Where L0 represents the normalization coefficient. In this embodiment, is set to 5. In other embodiments, it can be set to other values, preferably values greater than 1. This embodiment does not make a limitation.

[0064] As another example, the acquisition formula of the loss function S is as follows: During the first 20 training processes of the same speech sample, the loss function S is obtained using the method of the above comparison example. During the subsequent training process of the same speech sample, the loss function S is obtained using the method of the above preferred example.

[0065] This example can ensure that the first encoder and the third encoder pre-learn some speech features by pre-training the same speech sample, avoiding the situation where when using the method of the above preferred example to obtain the loss function S subsequently, the second encoder learns and stores too many speech features and cannot enable the first encoder to learn and extract speech features in advance.

[0066] (3) Using the above dataset, and the above loss function S, and training the multi-modal speech recognition model using the stochastic gradient descent algorithm.

[0067] The specific training process is well-known and will not be specifically described in this embodiment.

[0068] Step S104: Perform speech recognition using the trained multi-modal speech recognition model.

[0069] Obtain the trained multi-modal speech recognition model, delete the second encoder, and run the model on the computer of the smart mirror (such as an embedded development board like Raspberry Pi). The user stands in front of the smart mirror, and the smart mirror collects the user's speech and inputs the speech into the trained multi-modal speech recognition model. The model first outputs (i.e., the output of the first encoder) the time and body parts in the speech, as well as the first feature map. Then, obtain the weather speech command at the said time, and at the same time collect the user's full-body image, and only retain the body parts in the full-body image. Input the weather speech command and the full-body image into the fourth encoder to obtain the speech-assisted recognition features, and then input the first feature map and the speech-assisted recognition features into the third encoder to obtain the complete instruction recognition result. All the words corresponding to the word vectors in the complete instruction recognition result are used as the dressing plan required by the user.

[0070] So far, this embodiment ends.

[0071] The multi-modal speech recognition model of this embodiment is trained and used as a whole, initially avoiding the problem of too long speech recognition time and inaccurate speech recognition caused by error transmission. In addition, this embodiment allows changing the number of parameters of the third encoder and the fourth encoder to further shorten the speech recognition time. In order to further shorten the speech recognition time while ensuring the accuracy of speech recognition, this embodiment first obtains the first difference between the complete instruction recognition result and the dressing label, and then obtains the second difference between the complete instruction recognition result and the simplified instruction recognition result. The first difference and the second difference are fused, and the weight when the second difference is fused is positively correlated with the first difference; when there is a large training error in this process (i.e., when the first difference is large), the simplified instruction recognition result is allowed to be closer to the complete instruction recognition result. Since the simplified instruction recognition result is the instruction recognition result obtained in advance by skipping the feature extraction process of the image and the weather, this enables the first encoder to learn and extract some speech features in advance, avoiding the situation where when most of the speech features and all the image and weather features (i.e., speech-assisted recognition features) are fused in the third encoder with a relatively small number of parameters, the feature extraction or feature learning ability of speech, image, and weather is insufficient, resulting in a large error, making the speech recognition result of this embodiment more accurate (i.e., the generated dressing plan is more reliable).

[0072] Embodiment 2: This embodiment further considers that in order to further reduce the speech recognition time in Embodiment 1, in addition to ensuring that the number of parameters of the third encoder is relatively small. It is also necessary to ensure (or allow) that the number of parameters of the fourth encoder is small (as described in step S101 in detail), which results in the fourth encoder being unable to deeply extract or learn the feature of the weather modality data and the image modality data, resulting in a large noise in the speech-assisted recognition features (i.e., containing too many features useless for speech recognition).

[0073] When obtaining the loss function S according to the method in step S103 of Embodiment 1, the following situation may occur: Due to the introduction of the weight w, the third encoder extracts and learns more abstract and specific high-dimensional speech features. However, the noise of the speech auxiliary recognition features is large and they contain a lot of useless information, which causes the speech features in the third encoder to be submerged in the speech auxiliary recognition features, resulting in the third encoder ignoring the speech features and over-learning and extracting dressing schemes from the weather and images, making the speech recognition result inaccurate (for example, the generated dressing scheme is only for the weather, the user's body type, and the current clothing, and does not target the content described in the scene instructions in the user's speech).

[0074] To solve this problem and further improve the accuracy of speech recognition while reducing the time of speech recognition, this embodiment provides another method for obtaining the loss function S, and the method includes: After a same speech sample is trained several times (for example, 50 times) using the method in Embodiment 1 (that is, after the speech sample is input into the multi-modal speech recognition model 50 times): For a same speech sample and during the current training of the speech sample (that is, when the speech sample is currently input into the multi-modal speech recognition model), obtain the weight obtained during the first training before the current training (that is, when the speech sample was last input into the multi-modal speech recognition model) , and the complete instruction recognition result .

[0075] Obtain the weight obtained during the second training before the current training of the speech sample (that is, when the speech sample was input into the multi-modal speech recognition model the time before last), denoted as , and the complete instruction recognition result obtained is denoted as . And Any word vector in and is denoted as

[0076] When is larger than , and at the same time and When the similarity is smaller (or the difference is larger), it means that before the current training, after the first encoder learns and extracts some speech features in advance, the word vectors in the complete instruction recognition result have mutated. This may be due to the speech features in the third encoder being submerged in the speech auxiliary recognition features. At this time, the feature extraction and learning process of the i-th word vector needs to be placed in the third encoder, that is, the third encoder needs to focus on learning the relevant features of the i-th word vector.

[0077] Therefore, when is greater than and there exists and whose similarity is less than the first preset threshold th1, the loss function S is re-obtained: where represents the second difference between the complete instruction recognition result R1 and the refined instruction recognition result R2 during the current training. represents the first difference between the complete instruction recognition result R1 and the dressing label R during the current training. represents the weight of the second difference during the current training.

[0078] where the first difference , represents the i-th word vector in the dressing label of the speech sample, represents the i-th word vector in R1. represents the word vector and the word vector 's Euclidean distance, and N represents the number of word vectors in the dressing label.

[0079] where represents the attention coefficient of the i-th word vector in the complete instruction recognition result, and is positively correlated with the difference between and is negatively correlated with the similarities between and respectively.

[0080] This embodiment is described by taking th1 = 0.45 as an example.

[0081] As an example, , represents and 's cosine similarity, where describes the change of the word vectors in the complete instruction recognition result during two adjacent trainings (that is, the last and the penultimate trainings). The larger this value is, the more serious the mutation of the word vectors is. It describes the changes in the word vectors in the complete instruction recognition result when the weight of the second difference increases. As described above, the larger this value is, it indicates that during the current training, after the first encoder learns and extracts a part of the speech features in advance, it causes a mutation in the word vectors in the complete instruction recognition result. This may be due to the speech features in the third encoder being submerged in the speech auxiliary recognition features. At this time, the feature extraction and learning process of the i-th word vector needs to be placed in the third encoder, that is, it is necessary to make the third encoder focus on learning the relevant features of the i-th word vector.

[0082] Specifically, when it is less than or equal to 0.05, let , and its purpose is to avoid the value being too small, resulting in the denominator being 0.

[0083] When is not greater than , or and the cosine similarity of the word vectors in are both not less than the first preset threshold th1, the loss function S is obtained according to the method in Embodiment 1. At this time, it is equivalent to

[0084] w represents the weight of the second difference, .

[0085] It should be noted that this embodiment starts to run after several (for example, 50 times) of training using the method in Embodiment 1. The weights of the second differences obtained during the previous and the penultimate training (that is, w and ) are used to obtain the first difference during the current training, and then the weight of the second difference during the current training is obtained, and then the loss function during the current training is obtained. During this process, specifically, when the number of times of the previous or the penultimate training is less than 50, then the first difference and the weight of the second difference are both obtained according to the method in Embodiment 1.

[0086] Generally speaking, the first difference in this embodiment is different from the first difference in Embodiment 1. The in this embodiment compared with the , by introducing the attention coefficient of the word vector, the former enables the first encoder to learn and extract part of the speech features in advance, to a certain extent avoiding the situation where the speech features in the third encoder are submerged in the speech auxiliary recognition features, resulting in the inability of the third encoder to fuse and learn the speech features and the speech auxiliary recognition features. While reducing the time of speech recognition, it further improves the accuracy of speech recognition.

[0087] Embodiment 3: This embodiment provides a method for obtaining the loss functions of weather forecasts and images , and this method includes: In this embodiment, after collecting the user's speech and inputting it into the multi-modal speech recognition model (i.e., the first encoder), the weather forecast function and the image acquisition function need to be automatically triggered. Therefore, it is necessary to enable the first encoder to have the ability to recognize time and body parts from the speech. The specific method is to construct a loss function when training the multi-modal speech recognition model using the data set , so that the multi-modal speech recognition model can learn the time and body parts in the speech. The specific method is as follows: (1) For each speech sample in the data set, manually label the time represented by the time instruction in each speech sample, denoted as the time label; the time in this embodiment is a vector composed of the date and hour. For example, the time [a, b, c, d] represents a o'clock on the bth day of the a month, where d represents a special time mark. When the time instruction includes relative times such as "today", "tomorrow", "the day after tomorrow", etc., d is equal to 0, 1, 2 respectively (at this time, let a and b both be equal to -1); when it does not include the above relative times, d = -1. Specifically, when the time instruction includes general times such as "morning", "forenoon", "noon", "afternoon", "evening", "today", etc., let c be equal to 7, 10, 12, 15, 20, 12.

[0088] In other embodiments, other methods can be used to label the time represented by the time instruction. The specific labeling method is common knowledge and will not be specifically described and limited in this embodiment.

[0089] In addition, the body parts indicated by the body part instructions in each voice sample are manually marked and denoted as body part tags. The marked body part instruction regions are divided into multiple categories, which are represented by one-hot encodings as 0001, 0010, 0100, and 1000 respectively. Among them, 0001 represents the head (face), 0010 represents the upper body, 0100 represents the legs, and 1000 represents the feet. In addition, the one-hot encodings between the above four categories are combined with different categories through "OR operation" to obtain other categories. For example, when 0001 and 0010 are subjected to "OR operation", 0011 is obtained, indicating the head and the upper body. Another example is that when the four one-hot encodings of 0001, 0010, 0100, and 1000 are subjected to "OR operation", 1111 is obtained, indicating the whole body.

[0090] In other embodiments, the voice input can be fed into an existing large voice model (such as in ESPnet), and the large voice model is used to automatically identify and generate the above-mentioned time tags and body part tags.

[0091] (2) After traversing each voice sample in the dataset and inputting the voice sample into the first encoder, the time output by the first encoder is denoted as M1, and the body part output is denoted as M2. The time tag and body part tag corresponding to each voice sample are denoted as m1 and m2 respectively; where represents the Euclidean distance between M1 and m1, represents the function value of the cross-entropy loss function between M2 and m2.

[0092] Thus, the loss function is obtained, and then the method in step S103 of Embodiment 1 is used to train the multi-modal voice recognition model.

[0093] As described in step S104 of Embodiment 1: The user stands in front of the smart mirror, and the smart mirror collects the user's voice and inputs the voice into the trained multi-modal voice recognition model. The model first outputs (i.e., the first encoder outputs) the time and body part in the voice, as well as the first feature map. Then, the weather voice instruction at the obtained time is acquired, and at the same time, the full-body image of the user is collected, and only the body part in the full-body image is retained. The weather voice instruction and the full-body image are input into the fourth encoder to obtain the voice-assisted recognition feature, and then the first feature map and the voice-assisted recognition feature are input into the third encoder to obtain the complete instruction recognition result. All the vocabulary corresponding to the word vectors in the complete instruction recognition result is used as the dressing plan required by the user.

[0094] Among them, the model first outputs (i.e., the first encoder outputs) the time and body part in the voice, specifically including: The time in the voice output by the model (i.e., the output of the first encoder) is in vector format, still represented by [a, b, c, d], where a, b, c, and d are all rounded to the nearest integer. When d is less than 0, the time in the voice represents the a-th month, b-th day, and c-th hour; when d is greater than or equal to 0, the date represented by a month and b days is the date corresponding to d days after the current date.

[0095] The body part in the voice output by the model (i.e., the output of the first encoder) is in one-hot encoding format. Map this encoding format to the body parts described above. After collecting the full-body image of the user, use image recognition technology (such as YOLOV3 technology) to identify and frame the body parts in the full-body image (the rgb values of the pixel points in the area outside this body part are all set to 0).

[0096] Finally, input the weather voice command and the full-body image at a month, b days, and c hours into the fourth encoder to obtain voice-assisted recognition features. Then, input the first feature map and the voice-assisted recognition features into the third encoder to obtain the complete command recognition result. All the vocabulary corresponding to the word vectors in the complete command recognition result is used as the dressing plan required by the user.

[0097] Example 4: This example provides a method for inputting the weather voice command and the full-body image into the fourth encoder. The method includes: First, use a Transformer network with 2 layers of attention mechanisms to encode the weather voice command into a feature map (denoted as the weather feature map). Then, use another Transformer network with 2 layers of attention mechanisms to encode the full-body image into a feature map (denoted as the body feature map). These two feature maps have the same size. Stack these two feature maps together to obtain a concatenated feature map. Finally, use another Transformer network with 1 layer of attention mechanisms to encode the concatenated feature map into the voice-assisted recognition features.

[0098] The Transformer network with a total of 5 layers of attention mechanisms used above is the fourth encoder described in this example.

[0099] Example 5: This embodiment provides a voice command precise recognition system applicable to a smart mirror. The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps included in all the above embodiments. Additionally, the system further includes a smart mirror, which is equipped with a microphone for collecting user voices, a camera for collecting the full-body images of the user, and a weather forecast module for weather forecasting. The memory and the processor are installed behind the smart mirror. Additionally, a display screen is installed in the mirror area of the smart mirror for displaying the generated dressing plans.

[0100] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for accurate recognition of voice commands applicable to a smart mirror, characterized in that, The method includes the following steps: The constructed speech recognition model includes: inputting the collected speech into a first encoder, and the output of the first encoder is respectively used as the input of a second encoder and a third encoder. The second encoder outputs a recognition result of a concise instruction; by using a fourth encoder, after encoding the weather forecast and image collected under the time instruction and body part instruction included in the speech, it is also used as the input of the third encoder; the third encoder outputs a recognition result of a complete instruction; the number of parameters of the third encoder is less than that of the first encoder. The constructed dataset contains a number of speech samples and the dressing labels corresponding to the speech samples are marked; when each speech sample is input into the speech recognition model, a first difference between the recognition result of the complete instruction and the dressing label is obtained, and a second difference between the recognition result of the complete instruction and the recognition result of the concise instruction is obtained; the second difference and the first difference are weighted and fused to obtain a loss function, and the speech recognition model is trained and speech recognition is performed by using the dataset and the loss function. Among them, when performing weighted fusion, the weight of the second difference is positively correlated with the first difference. When the weight obtained after the same speech sample is input into the speech recognition model twice in succession increases, the change generated by the word vectors in the recognition result of the complete instruction is recorded as the attention coefficient of the word vectors, and the first difference is obtained from the attention coefficient.

2. The method for accurately recognizing voice commands applicable to a smart mirror according to claim 1, characterized in that, The specific steps for obtaining the first difference are as follows: Both the recognition result of the complete instruction and the dressing label are vector sequences, and the vector sequences contain a number of word vectors. Obtain the Euclidean distance of the word vectors in the same order in the recognition result of the complete instruction and the dressing label, and use the attention coefficient of each word vector in the recognition result of the complete instruction to perform weighted summation on the Euclidean distances of all the word vectors to obtain the first difference.

3. The method for accurately recognizing voice commands applicable to a smart mirror according to claim 1, characterized in that, The specific steps for the second difference are as follows: Both the recognition result of the complete instruction and the recognition result of the concise instruction are vector sequences, and the vector sequences contain word vectors. Obtain the Euclidean distance of the word vectors in the same order in the recognition result of the complete instruction and the recognition result of the concise instruction; the mean value of the Euclidean distances of all the word vectors in the recognition result of the complete instruction is recorded as the second difference.

4. The method for accurately recognizing voice commands applicable to a smart mirror according to claim 1, characterized in that, When the weight obtained after the same speech sample is input into the speech recognition model twice in succession increases, the change generated by the word vectors in the recognition result of the complete instruction is recorded as the attention coefficient of the word vectors, and the specific steps included are as follows: Each time a speech sample is input into the speech recognition model, it means that each speech sample is trained once. For the same voice sample, and when the voice sample is trained this time, obtain the weights obtained during the first training before the current training, denoted as , and the obtained complete instruction recognition result is denoted as ; Obtain the weight obtained during the second training before the current training of the voice sample, denoted as , and denote the obtained complete instruction recognition result as ; and the i-th word vector in and ; the attention coefficient sum is positively correlated with the difference between and and is negatively correlated with the similarity.

5. The voice command precise recognition method applicable to a smart mirror according to claim 1, wherein The step of weighted fusion of the second difference and the first difference to obtain a loss function includes the following specific formula: Among them, S represents the loss function, represents the first difference between the complete instruction recognition result R1 and the dressing label R, represents the second difference between the complete instruction recognition result R1 and the reduced instruction recognition result R2; represents the weight of the second difference, represents the loss function of weather forecast and images.

6. The voice command precise recognition method applicable to an intelligent mirror according to claim 5, wherein, The specific steps for obtaining the loss function of the weather forecast and the image are as follows: Each speech sample contains a time instruction and a body part instruction. Mark the body part represented by the body part instruction in each speech sample, which is recorded as the body part label m2; the time represented by the time instruction in each speech sample is recorded as the time label m1. After each speech sample in the dataset is input into the speech recognition model, the first encoder outputs the time, which is recorded as M1, and at the same time the first encoder outputs the body part, which is recorded as M2. where represents the Euclidean distance between M1 and m1, represents the function value of the cross-entropy loss function between M2 and m2.

7. The voice command precise recognition method applicable to a smart mirror according to claim 1, characterized in that, After encoding the weather forecast and image collected under the time instruction and body part instruction included in the speech by using the fourth encoder, it is also used as the input of the third encoder, and the specific steps are as follows: The speech includes a time instruction and a body part instruction; After the speech is input into the first encoder, the first encoder outputs the time described by the time instruction and the body part described by the body part instruction; then, the weather forecast is made by using the time to obtain a weather speech instruction, and the full-body image of the user is collected. Then, the body part in the weather speech instruction and the full-body image is input into the fourth encoder to obtain a speech-assisted recognition feature, and the speech-assisted recognition feature is used as the input of the third encoder.

8. The method for accurate voice command recognition applicable to a smart mirror according to claim 1, wherein The speech recognition includes the following specific steps: After the speech recognition model is trained, the second encoder is deleted. The user stands in front of the smart mirror, and the smart mirror collects the user's speech and inputs the speech into the trained speech recognition model. The first encoder outputs the time and body part in the speech, and the first feature map; then, the weather speech instruction at the time is obtained, and at the same time, the full-body image of the user is collected. The body part in the weather speech instruction and the full-body image is input into the fourth encoder, and the speech-assisted recognition feature is obtained. Then, the first feature map and the speech-assisted recognition feature are stacked together and input into the third encoder to obtain a complete instruction recognition result. All the words corresponding to the word vectors in the complete instruction recognition result are used as the dressing plan required by the user.

9. The voice command precise recognition method applicable to a smart mirror according to claim 1, characterized in that, When performing the weighted fusion, the weight of the second difference is positively correlated with the first difference, and the specific formula is as follows: where L0 represents a preset normalization coefficient, and w represents the weight of the second difference, represents the first difference between the complete instruction recognition result R1 and the dressing label R.

10. A voice command precise recognition system applicable to a smart mirror, the system comprising: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method for accurately recognizing speech instructions applicable to a smart mirror according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Household intelligent fitting mirror system

    CN107451896A

  • Augmented mirror

    CN108431730A

  • Mirror based on smart house and control method thereof

    CN108682045A

  • Handshake interaction method and system based on intelligent mirror and storage medium

    CN110751951A

  • Clothes management method and device and smart dressing mirror

    CN110859047A