Speech recognition method, speech recognition system and electrical appliance
By simultaneously obtaining voice and lip information, generating and synthesizing final statements, the problem of insufficient accuracy of speech recognition methods in the prior art in the noisy environment is solved, and more accurately obtaining user intentions is achieved and user experience is improved.
Patent Information
- Application Number
- CN202010485180.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-06-01
AI Technical Summary
The existing speech recognition methods are prone to problems in electrical equipment that cannot accurately obtain the user's true intentions, especially in noise environments. In the prior art, errors are prone to occur only by single judgment of lip or voice information.
By simultaneously obtaining the voice information and lip information, the first and second statements are generated respectively, and the intention words are selectively retained according to the semantic similarity and the ambient noise size, and the final statement is synthesized to represent the user's true intention.
It significantly improves the accuracy of obtaining the user's true intentions in a noisy environment and improves the user's user experience.
Smart Images

Figure CN113763941B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human-computer interaction, and specifically provides a voice recognition method, a voice recognition system, and an electrical appliance device. Background Art
[0002] Electrical appliance devices can be divided into household electrical appliances and commercial electrical appliances. Among them, household electrical appliances mainly include washing machines, refrigerators, air conditioners, etc. With the increasing popularity of electrical appliance devices, the functions of various electrical appliance devices have become more and more powerful, and people's requirements for electrical appliance devices have also become higher and higher.
[0003] Taking a washing machine as an example, in order to improve the user experience, some washing machines have added a voice recognition function. Users can directly perform voice control on the washing machine, which is both convenient and fast. However, in actual applications, we found that the voice recognition function of the washing machine often makes mistakes, resulting in the washing machine being unable to accurately obtain the true intention of the user, which affects the user experience.
[0004] A patent application document with the publication number of CN111045639A discloses a voice input method. The method includes: when receiving a voice input instruction, receiving a voice signal collected by a microphone; obtaining ambient noise from the voice signal; when the sound intensity of the ambient noise is greater than a preset intensity threshold, obtaining a lip image; performing lip language recognition on the lip image to obtain a lip language recognition result; performing voice recognition on the voice signal to obtain a voice input result; when the lip language recognition result and the voice input result match, using the voice input result as the user input information corresponding to the lip image; and displaying the content corresponding to the user input information. That is to say, after obtaining the lip language recognition result in the above patent, only the lip language recognition result is used to judge the accuracy of the voice input recognition result. If the lip language recognition result and the voice input result match, the voice input result can be adopted. However, what if they do not match? The patent does not explain how to handle the situation when the lip language recognition result and the voice input result do not match.
[0005] The patent application document with the publication number CN108319912 discloses a lip-reading recognition method, which includes: obtaining lip-reading information by using a camera, and the lip-reading information includes three categories: mouth shapes, mouth muscle movement information, lip color, skin color information around the mouth, and facial expression information; comparing and analyzing the mouth shapes and mouth muscle movement information with the mouth shape model information in the memory to obtain the first lip-reading information; detecting the pigment distribution of the lip color and the skin color around the mouth according to the lip color and the skin color information around the mouth, judging the mouth movement characteristics by using the intelligent algorithm in the memory, and comparing and analyzing with the mouth shape model information in the memory to obtain the second lip-reading information; using image processing technology to perform expression recognition on the facial expression information, and comparing and analyzing with the expression model information in the memory to obtain the third lip-reading information; performing normalization processing on the first lip-reading information, the second lip-reading information, and the third lip-reading information, and converting the normalized lip-reading information into voice information. That is to say, in this patent, only the lip-reading information is used to obtain the true intention of the user, and no voice information is collected. However, it is also easy to make mistakes when obtaining the true intention of the user only through lip-reading information.
[0006] Therefore, there is a need in the art for a voice recognition method, a voice recognition system, and an electrical device to solve the above problems. Summary of the Invention
[0007] To solve the above problems in the prior art, that is, to solve the problem that the existing voice recognition method is prone to inaccurately obtaining the true intention of the user, the present invention provides a voice recognition method, and the voice recognition method includes: obtaining voice information and lip-reading information, and generating a first sentence according to the voice information; generating a second sentence according to the lip-reading information; generating a final sentence according to the first sentence and the second sentence.
[0008] In a preferred technical solution of the above voice recognition method, the step of "generating a final sentence according to the first sentence and the second sentence" specifically includes: disassembling the first sentence to obtain a plurality of first intention words of different categories; disassembling the second sentence to obtain a plurality of second intention words of different categories; respectively judging whether the semantic similarity between the first intention words and the second intention words in the same category reaches the requirement; according to the judgment result, selectively retaining the first intention word or the second intention word; generating a final sentence according to the finally retained first intention words and second intention words.
[0009] In a preferred technical solution of the above voice recognition method, the step of "selectively retaining the first intention word or the second intention word according to the judgment result" specifically includes: if the semantic similarity between the first intention word and the second intention word does not reach the requirement, then selectively retain the first intention word or the second intention word according to the magnitude of the environmental noise.
[0010] In the preferred technical solution of the above voice recognition method, the step of "selectively retaining the first intended word or the second intended word according to the magnitude of the environmental noise" specifically includes: if the environmental noise is in the low-noise area, retain the first intended word; if the environmental noise is in the medium-noise area and the first intended word and the second intended word belong to the high-stability category, retain the first intended word; if the environmental noise is in the medium-noise area and the first intended word and the second intended word belong to the low-stability category, retain the second intended word; if the environmental noise is in the high-noise area, retain the second intended word, wherein the degree of influence of the words in the high-stability category by the environmental noise is less than the degree of influence of the words in the low-stability category by the environmental noise.
[0011] In the preferred technical solution of the above voice recognition method, the step of "selectively retaining the first intended word or the second intended word according to the judgment result" further includes: if the semantic similarity between the first intended word and the second intended word reaches the requirement, retain the first intended word, or retain the second intended word, or randomly retain one of the first intended word and the second intended word.
[0012] On the other hand, the present invention also provides a voice recognition system, which includes: a sound acquisition device configured to be able to collect voice information; an image acquisition device configured to be able to collect lip language information; and an information processing device configured to be able to generate a first sentence and a second sentence respectively according to the voice information collected by the sound acquisition device and the lip language information collected by the image acquisition device, and generate a final sentence according to the first sentence and the second sentence.
[0013] In the preferred technical solution of the above voice recognition system, the information processing device includes: a sound information processing module configured to be able to generate the first sentence according to the voice information; an image information processing module configured to be able to generate the second sentence according to the lip language information; a sentence analysis and processing module configured to be able to disassemble the first sentence and the second sentence respectively to obtain a plurality of first intended words of different categories and a plurality of second intended words of different categories, and be able to respectively judge whether the semantic similarity between the first intended word and the second intended word in the same category reaches the requirement, and selectively retain the first intended word or the second intended word according to the judgment result, and finally be able to generate a final sentence according to the finally retained first intended word and second intended word.
[0014] In the preferred technical solution of the above voice recognition system, the statement analysis and processing module is further configured to: when the semantic similarity between the first intention word and the second intention word meets the requirement, retain the first intention word, or retain the second intention word, or randomly retain one of the first intention word and the second intention word; when the semantic similarity between the first intention word and the second intention word does not meet the requirement, selectively retain the first intention word or the second intention word according to the magnitude of the ambient noise collected by the sound acquisition device.
[0015] In the preferred technical solution of the above voice recognition system, the statement analysis and processing module is further configured to: in the case where the semantic similarity between the first intention word and the second intention word does not meet the requirement, when the ambient noise is in the low noise area, retain the first intention word; when the ambient noise is in the medium noise area and the first intention word and the second intention word belong to the high stability category, retain the first intention word; when the ambient noise is in the medium noise area and the first intention word and the second intention word belong to the low stability category, retain the second intention word; when the ambient noise is in the high noise area, retain the second intention word, where the degree of influence of the words in the high stability category by the ambient noise is less than the degree of influence of the words in the low stability category by the ambient noise.
[0016] On the other hand, the present invention also provides an electrical device, and the motor device includes the above voice recognition system.
[0017] Those skilled in the art can understand that in the preferred technical solution of the present invention, by simultaneously obtaining speech information and lip language information, generating a first sentence and a second sentence according to the obtained speech information and lip language information respectively, and then generating a final sentence according to the first sentence and the second sentence, the meaning expressed by the final sentence is regarded as the true intention of the user. Compared with the patent with the publication number CN111045639A, after obtaining the lip language recognition result in that patent, only the lip language recognition result is used to judge whether the speech input recognition result is accurate. However, in the present invention, a sentence is generated according to the speech information and the lip language information respectively, and then these two sentences are analyzed and compared to synthesize a new sentence, that is, the final sentence, and the meaning expressed by the final sentence is regarded as the true intention of the user, that is, the true intention of the user is jointly judged by the speech information and the lip language information. By mutual verification and comparison between the speech information and the lip language information, the accuracy of the judgment can be significantly improved, so that the true intention of the user can be obtained more accurately and the user experience can be improved. In addition, compared with the prior art of obtaining the true intention of the user only through speech information and the patent with the publication number CN108319912A of obtaining the true intention of the user only through lip language information, the present invention jointly judges the true intention of the user through speech information and lip language information, which can significantly improve the accuracy of the judgment.
[0018] Further, if the semantic similarity between the first intention word and the second intention word does not meet the requirement, the first intention word or the second intention word is selectively retained according to the magnitude of the ambient noise. Through such a setting, that is, selectively retaining the first intention word or the second intention word according to the magnitude of the ambient noise, the interference of the ambient noise can be effectively excluded, and the accuracy of the judgment can be further improved.
[0019] Further, the step of "selectively retaining the first intention word or the second intention word according to the magnitude of the ambient noise" specifically includes: if the ambient noise is in the low-noise area, the first intention word is retained; if the ambient noise is in the medium-noise area and the first intention word and the second intention word belong to the high-stability category, the first intention word is retained; if the ambient noise is in the medium-noise area and the first intention word and the second intention word belong to the low-stability category, the second intention word is retained; if the ambient noise is in the high-noise area, the second intention word is retained, where the degree of influence of the words in the high-stability category by the ambient noise is less than the degree of influence of the words in the low-stability category by the ambient noise. Through such a setting, that is, when the ambient noise is in the medium-noise area, the selection is made according to the categories of the first intention word and the second intention word, the accuracy of the judgment can be further improved. Description of the Drawings
[0020] Figure 1 is a flowchart of the speech recognition method of the present invention;
[0021] Figure 2 is a flowchart of an embodiment of the speech recognition method of the present invention;
[0022] Figure 3 is a schematic structural diagram of the speech recognition system of the present invention. Specific Embodiments
[0023] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention. For example, although the following embodiments are explained in conjunction with a washing machine, this is not restrictive. The technical solutions of the present invention are equally applicable to other electrical devices, such as household appliances like refrigerators and air conditioners, as well as commercial electrical appliances. Such changes in the application object do not deviate from the principle and scope of the present invention and should all be limited within the protection scope of the present invention.
[0024] Based on the problem pointed out in the background art that the existing speech recognition method is prone to the problem of being unable to accurately obtain the true intention of the user. The present invention provides a speech recognition method, a speech recognition device, and an electrical device, aiming to obtain the true intention of the user according to both voice information and lip movement information.
[0025] The washing machine of the present invention includes a speech recognition system, and the washing machine can accurately obtain the true intention of the user through this speech recognition system.
[0026] First, refer to Figure 3 , Figure 3 which is a schematic structural diagram of the speech recognition system of the present invention. As Figure 3 shown, the speech recognition system of the present invention includes a sound acquisition device, an image acquisition device, and an information processing device. Both the sound acquisition device and the image acquisition device can communicate with the information processing device. Among them, the sound acquisition device is a device such as a microphone that can receive sound information, the image acquisition device is a device such as a camera that can collect image information, and the information processing device is a processor.
[0027] When the user operates the washing machine, the voice acquisition device can receive what the user says, that is, collect voice information, and transmit the collected voice information to the information processing device. After receiving the voice information, the information processing device starts to analyze and process the voice information, converting the voice information into a sentence, which can be recorded as the first sentence; while the voice acquisition device is collecting voice information, the image acquisition device can collect the user's facial image, that is, collect lip language information, and transmit the collected lip language information to the information processing device. After receiving the lip language information, the information processing device starts to analyze and process the lip language information, also converting the lip language information into a sentence, which can be recorded as the second sentence. Finally, the information processing device generates a final sentence based on the first sentence and the second sentence obtained from the previous analysis, and this final sentence is the expression of the user's true intention.
[0028] In the present invention, by simultaneously acquiring voice information and lip language information, generating the first sentence and the second sentence respectively according to the acquired voice information and lip language information, and then generating the final sentence according to the first sentence and the second sentence, the meaning expressed by this final sentence is regarded as the true intention of the user.
[0029] Compared with the patent with the publication number CN111045639A, after obtaining the lip language recognition result in that patent, only the lip language recognition result is used to judge whether the speech input recognition result is accurate. However, in the present invention, a sentence is generated respectively according to the voice information and the lip language information, and then these two sentences are analyzed and compared to synthesize a new sentence, that is, the final sentence. The meaning expressed by this final sentence is regarded as the true intention of the user, that is, the true intention of the user is jointly judged through the voice information and the lip language information. Through the mutual verification and comparison between the voice information and the lip language information, the accuracy of the judgment can be significantly improved, so that the true intention of the user can be obtained more accurately and the user experience can be enhanced.
[0030] In addition, compared with the prior art of obtaining the true intention of the user only through voice information and the patent with the publication number CN108319912A of obtaining the true intention of the user only through lip language information, the present invention jointly judges the true intention of the user through voice information and lip language information, which can significantly improve the accuracy of the judgment.
[0031] Continue to refer to Figure 3 , the information processing device of the present invention includes a voice information processing module, an image information processing module, and a sentence analysis and processing module.
[0032] After the information processing device receives the voice information, the voice information processing module starts to analyze and process the voice information, thereby generating the first sentence. Then, through the sentence analysis and processing module, the first sentence is disassembled into multiple words of different categories, and each word can represent a different intention, which can be recorded as the first intention word.
[0033] Similarly, after the information processing device receives the lip-reading information, the image information processing module starts to analyze and process the lip-reading information to generate a second statement. Then, the statement analysis and processing module disassembles the second statement into multiple words of different categories, and each word can represent a different intention, which can be recorded as the second intention word.
[0034] After the statement analysis and processing module disassembles the first statement and the second statement into multiple first intention words of different categories and multiple second intention words of different categories respectively, it starts to analyze and compare the first intention words and the second intention words in the same category, determines whether the semantic similarity between the first intention words and the second intention words meets the requirements, and then selectively retains one of them according to the judgment result. Finally, only one word is retained in each category, and this word may be the first intention word or the second intention word. Finally, these retained first intention words and second intention words are combined into a complete statement, which is the final statement.
[0035] Exemplarily, the first statement is disassembled into 3 words with different intentions, that is, 3 first intention words are obtained, which are respectively classified into the first category, the second category, and the third category. Similarly, the second statement is also disassembled into 3 words with different intentions, that is, 3 second intention words are obtained, which are also respectively classified into the first category, the second category, and the third category. Each category contains two words, one is the first intention word and the other is the second intention word. Then, the two words in the first category are analyzed and compared to determine whether the semantic similarity between the two words meets the requirements. Finally, only one word is retained, and this word may be the first intention word or the second intention word. Similarly, the two words in the second category and the two words in the third category are also analyzed and compared, and only one word is retained in each of them. Finally, only one word is retained in each category, and a total of three words are obtained. Then, these three words are combined into a complete statement, which is the final statement.
[0036] When the statement analysis and processing module analyzes and compares the first intention words and the second intention words in the same category, the following two situations will occur:
[0037] In the first situation, the semantic similarity between the first intention word and the second intention word meets the requirements, indicating that the intentions expressed by the first intention word and the second intention word are basically the same. In this situation, the first intention word can be directly retained, or the second intention word can be directly retained, or one of the first intention word or the second intention word can be randomly retained.
[0038] In the second case, the semantic similarity between the first intended word and the second intended word does not meet the requirements, indicating that the intentions expressed by the first intended word and the second intended word are quite different. In this case, one of them cannot be randomly selected. Preferably, according to the magnitude of the ambient noise, the first intended word or the second intended word is selectively retained. The specific method is as follows:
[0039] When the ambient noise is in the low-noise area, retain the first intended word;
[0040] When the ambient noise is in the medium-noise area and the first intended word and the second intended word belong to the high-stability category, retain the first intended word;
[0041] When the ambient noise is in the medium-noise area and the first intended word and the second intended word belong to the low-stability category, retain the second intended word;
[0042] When the ambient noise is in the high-noise area, retain the second intended word.
[0043] Among them, the influence of words in the high-stability category by the ambient noise is less than that of words in the low-stability category by the ambient noise.
[0044] That is, when the ambient noise is in the low-noise area, the influence of the ambient noise on the judgment of the voice information is small, so the first intended word is used as the standard; on the contrary, when the ambient noise is in the high-noise area, the influence of the ambient noise on the judgment of the voice information is large, so the second intended word is used as the standard.
[0045] However, when the ambient noise is in the medium-noise area, the specific categories of the first intended word and the second intended word need to be considered. The inventor found through a large number of experimental studies that some categories of words are less affected by the ambient noise, and these categories are recorded as the high-stability category. There are also some categories of words that are more affected by the ambient noise, and these categories are recorded as the low-stability category. Therefore, when the ambient noise is in the medium-noise area and the first intended word and the second intended word belong to the high-stability category, retain the first intended word. On the contrary, when the ambient noise is in the medium-noise area and the first intended word and the second intended word belong to the low-stability category, retain the second intended word.
[0046] It should be noted that the ambient noise can also be collected by a sound acquisition device.
[0047] On the other hand, the present invention also provides a voice recognition method. As Figure 1 shown, the voice recognition method of the present invention includes the following steps:
[0048] S100: Obtain voice information and lip language information;
[0049] S200: Generate a first statement according to the voice information;
[0050] S300: Generate a second statement based on the lip-reading information;
[0051] S400: Generate a final statement based on the first statement and the second statement.
[0052] In the present invention, by simultaneously acquiring voice information and lip-reading information, generating a first statement and a second statement respectively according to the acquired voice information and lip-reading information, and then generating a final statement based on the first statement and the second statement, the meaning expressed by the final statement is regarded as the true intention of the user. Compared with the prior art where only voice information is used to judge the true intention of the user.
[0053] Compared with the patent with the publication number CN111045639A, after obtaining the lip-reading recognition result in that patent, only the lip-reading recognition result is used to judge whether the voice input recognition result is accurate. However, in the present invention, a statement is generated respectively according to the voice information and the lip-reading information, and then these two statements are analyzed and compared to synthesize a new statement, that is, the final statement. The meaning expressed by the final statement is regarded as the true intention of the user, that is, the true intention of the user is jointly judged through the voice information and the lip-reading information. By mutual verification and comparison between the voice information and the lip-reading information, the accuracy of the judgment can be significantly improved, so that the true intention of the user can be obtained more accurately and the user experience can be enhanced.
[0054] In addition, compared with the prior art where only voice information is used to obtain the true intention of the user and the patent with the publication number CN108319912A where only lip-reading information is used to obtain the true intention of the user, the present invention jointly judges the true intention of the user through voice information and lip-reading information, which can significantly improve the accuracy of the judgment.
[0055] It should be noted that step S200 and step S300 can be executed synchronously or in any order successively. Such flexible adjustment and change do not deviate from the principle and scope of the present invention and should all be defined within the protection scope of the present invention.
[0056] Preferably, as Figure 2 shown, step S400 (generating a final statement based on the first statement and the second statement) specifically includes the following steps:
[0057] S401: Decompose the first statement to obtain multiple first intention words of different categories;
[0058] S402: Decompose the second statement to obtain multiple second intention words of different categories;
[0059] S403: Respectively judge whether the semantic similarity between the first intention word and the second intention word in the same category reaches the requirement;
[0060] S404: Selectively retain the first intended word or the second intended word according to the judgment result;
[0061] S405: Generate a final statement according to the finally retained first intended word and second intended word.
[0062] Decompose the first statement into multiple words of different categories, each of which can represent different intentions, and can be recorded as the first intended word; decompose the second statement into multiple words of different categories, each of which can represent different intentions, and can be recorded as the second intended word; then, analyze and compare the first intended word and the second intended word in the same category, judge whether the semantic similarity between the first intended word and the second intended word meets the requirements, and then, according to the judgment result, selectively retain one of them. Finally, only one word is retained in each category, and this word may be the first intended word or the second intended word. Finally, combine these retained first intended words and second intended words into a complete statement, which is the final statement.
[0063] Exemplarily, the first statement is decomposed into 3 words with different intentions, that is, 3 first intended words are obtained, which are respectively classified into the first category, the second category and the third category. Similarly, the second statement is also decomposed into 3 words with different intentions, that is, 3 second intended words are obtained, which are also respectively classified into the first category, the second category and the third category. Each category contains two words, one is the first intended word and the other is the second intended word. Then, analyze and compare the two words in the first category, judge whether the semantic similarity between the two words meets the requirements, and finally only retain one of the words, which may be the first intended word or the second intended word. Similarly, analyze and compare the two words in the second category and the two words in the third category, and also only retain one of the words. Finally, only one word is retained in each category, and a total of three words are obtained. Then, combine these three words into a complete statement, which is the final statement.
[0064] It should be noted that step S401 and step S402 can be executed synchronously or in any order successively. Such flexible adjustment and change do not deviate from the principle and scope of the present invention and should all be limited within the protection scope of the present invention.
[0065] Preferably, the steps of step S404 (selectively retaining the first intended word or the second intended word according to the judgment result) specifically include:
[0066] If the semantic similarity between the first intended word and the second intended word meets the requirements, then retain the first intended word, or retain the second intended word, or randomly retain one of the first intended word or the second intended word;
[0067] If the semantic similarity between the first intended word and the second intended word does not meet the requirement, then depending on the level of environmental noise, either the first intended word or the second intended word is selectively retained.
[0068] That is, when analyzing and comparing the first intended word and the second intended word in the same category, the following two situations will occur:
[0069] In the first situation, the semantic similarity between the first intended word and the second intended word meets the requirement, indicating that the intentions expressed by the first intended word and the second intended word are basically the same. In this situation, the first intended word can be directly retained, or the second intended word can be directly retained, or one of the first intended word or the second intended word can be randomly retained.
[0070] In the second situation, the semantic similarity between the first intended word and the second intended word does not meet the requirement, indicating that the intentions expressed by the first intended word and the second intended word are quite different. In this situation, one of them cannot be randomly selected. Preferably, depending on the level of environmental noise, either the first intended word or the second intended word is selectively retained.
[0071] Preferably, the step of "selectively retaining the first intended word or the second intended word according to the level of environmental noise" specifically includes:
[0072] If the environmental noise is in the low-noise area, then the first intended word is retained;
[0073] If the environmental noise is in the medium-noise area and the first intended word and the second intended word belong to the high-stability category, then the first intended word is retained;
[0074] If the environmental noise is in the medium-noise area and the first intended word and the second intended word belong to the low-stability category, then the second intended word is retained;
[0075] If the environmental noise is in the high-noise area, then the second intended word is retained.
[0076] Among them, the degree to which words in the high-stability category are affected by environmental noise is less than the degree to which words in the low-stability category are affected by environmental noise.
[0077] That is, when the environmental noise is in the low-noise area, the impact on the judgment of voice information is small, so the first intended word is used as the standard; on the contrary, when the environmental noise is in the high-noise area, the impact on the judgment of voice information is large, so the second intended word is used as the standard.
[0078] However, when the ambient noise is in the medium noise area, the specific categories of the first intention words and the second intention words need to be considered. Through a large number of experimental studies, the inventor found that some categories of words are less affected by ambient noise, and these categories are recorded as high-stability categories. There are also some categories of words that are more affected by ambient noise, and these categories are recorded as low-stability categories. Therefore, when the ambient noise is in the medium noise area and the first intention words and the second intention words belong to the high-stability category, the first intention words are retained. On the contrary, when the ambient noise is in the medium noise area and the first intention words and the second intention words belong to the low-stability category, the second intention words are retained.
[0079] The technical solution of the present invention will be described in detail below in conjunction with a specific embodiment.
[0080] For washing machines, we divide the intentions into two major categories: main function intentions and auxiliary function intentions. Among them, the main function intentions can include program type intentions, clothing type intentions, etc. The auxiliary function intentions mainly involve various washing parameters, such as washing time, number of rinses, etc.
[0081] It should be noted that the words of the main function intentions all belong to the high-stability category, and the words of the auxiliary function intentions all belong to the low-stability category.
[0082] When the user operates the washing machine, the first sentence generated according to the obtained voice information is "Wash the woolen sweater, 30 minutes, rinse 2 times", and the second sentence generated according to the obtained lip language information is "Wash the down jacket, washing time 40 minutes, rinse 3 times".
[0083] Then, the first sentence and the second sentence are disassembled respectively to obtain multiple different categories of first intention words and multiple different categories of second intention words. For a clearer representation, a comparison is made in tabular form. The table is as follows:
[0084] Intention category First intention word Second intention word Program type Washing Wash Clothing type Sweater Down jacket Washing time 30 minutes 40 minutes Number of rinses 2 times 2 times
[0085] After analysis and comparison, it can be seen that:
[0086] For the program type, the meanings of "wash" and "clean" are similar, that is, the semantic similarity of the first intention word and the second intention word meets the requirements. In this case, either the first intention word or the second intention word can be retained. Taking the retention of the first intention word as an example, that is, the program type is "wash";
[0087] For the clothing type, the meanings of "woolen sweater" and "down jacket" are significantly different, that is, the semantic similarity of the first intention word and the second intention word does not meet the requirements. In this case, it is necessary to selectively retain the first intention word or the second intention word according to the size of the ambient noise, as follows:
[0088] When the ambient noise belongs to the low-noise area (for example, the low-noise area is less than 70 decibels), retain the first intended word, that is, the clothing type is "woolen sweater";
[0089] When the ambient noise belongs to the medium-noise area (for example, the medium-noise area is 70 - 90 decibels), since the clothing type belongs to the high-stability category, retain the first intended word, that is, the clothing type is "woolen sweater";
[0090] When the ambient noise belongs to the high-noise area (for example, the high-noise area is greater than 90 decibels), then retain the second intended word, that is, the clothing type is "down jacket";
[0091] Regarding the washing time, the meanings of "30 minutes" and "40 minutes" are significantly different, that is, the semantic similarity of the first and second intended words does not meet the requirements. In this case, it is necessary to selectively retain the first or second intended word according to the magnitude of the ambient noise, as follows:
[0092] When the ambient noise belongs to the low-noise area, retain the first intended word, that is, the washing time is "30 minutes";
[0093] When the ambient noise belongs to the medium-noise area, since the washing time belongs to the low-stability category, retain the second intended word, that is, the washing time is "40 minutes";
[0094] When the ambient noise belongs to the high-noise area, then retain the second intended word, that is, the washing time is "40 minutes";
[0095] Regarding the number of rinses, the meanings of "2 times" and "2 times" are the same, that is, the semantic similarity of the first and second intended words meets the requirements. In this case, either the first intended word or the second intended word can be retained. Taking the retention of the first intended word as an example, that is, the number of rinses is "2 times";
[0096] Assume that the ambient noise is 80 decibels, belonging to the medium-noise area, and the finally obtained statement is "Wash the woolen sweater, the washing time is 40 minutes, and rinse 2 times."
[0097] Finally, it should be noted that general algorithm models such as Hidden Markov Model (HMM), Time Delay Neural Network (TDNN), or Convolutional Neural Network (CNN) can be used to analyze the speech information and convert it into a statement.
[0098] In addition, it should be noted that lip reading refers to identifying the content spoken by observing the lip movements of the speaker. In a noisy environment, we can "guess" what the speaker is saying by observing the characteristics of the lip movements, thus making up for the deficiency of the auditory signal. Visual signals can provide more distinguishable information for phonemes sensitive to noise. For example, some pronunciations that are difficult to distinguish in the speech signal channel are easily distinguishable visually.
[0099] To implement a complete lip reading system, it is necessary to complete multiple complex work links such as collecting the video information of the speaker, going through lip detection, feature extraction, recognition, etc. According to different implemented functions, the lip reading system can be divided into the following three main links:
[0100] The first step is lip detection, which is to find the approximate position of the lips from the given image or video. This is a prerequisite for lip reading. The approximate range of the lips can be determined mainly through the following methods: Method 1: Determine according to the physiological structure of the human face. Since the pupil of the eye has a lower gray level compared with the surrounding face and is relatively easy to locate, usually the pupil is located first, and then the approximate position of the lips is determined according to the position of the eyes and the positional relationship between the eyes and the mouth; Method 2: Determine the position of the lips according to the gray level information or color information of the lips; Method 3: Monitor the lips according to the motion information.
[0101] The second step is lip movement localization and feature extraction. In the lip reading system, it is necessary to automatically and real-time localize and track the lip movement. Extracting lip movement features is a prerequisite for further recognition. The quality of localization and feature extraction directly affects the result of lip reading. It can be achieved through methods such as variable template and Snake method, principal component analysis, or optical flow analysis method.
[0102] Third, lip reading. Perform lip reading on the extracted feature quantities. Lip reading and speech recognition both belong to the category of dynamic sequence feature recognition. It is also possible to analyze the lip reading information and convert it into a sentence through general algorithm models such as Hidden Markov Model (HMM), Time Delay Neural Network (TDNN), or Convolutional Neural Network (CNN).
[0103] In addition, it should be noted that it is possible to use sentence analysis and processing such as recurrent neural network or LSTM (Long Short Term Memory network) neural network model.
[0104] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A voice recognition method, characterized in that, The speech recognition method includes: Obtaining speech information and lip-reading information; Generating a first statement according to the speech information; Generating a second statement according to the lip-reading information; Generating a final statement according to the first statement and the second statement; The step of "generating a final statement according to the first statement and the second statement" specifically includes: Decomposing the first statement to obtain a plurality of first intent words of different categories; Decomposing the second statement to obtain a plurality of second intent words of different categories; Respectively determining whether the semantic similarity between the first intent word and the second intent word in the same category reaches the requirement; According to the judgment result, selectively retaining the first intent word or the second intent word; Generating a final statement according to the finally retained first intent word and second intent word.
2. The speech recognition method according to claim 1, wherein The step of "selectively retaining the first intent word or the second intent word according to the judgment result" specifically includes: If the semantic similarity between the first intent word and the second intent word does not reach the requirement, then selectively retain the first intent word or the second intent word according to the magnitude of the environmental noise.
3. The voice recognition method according to claim 2, characterized in that The step of "selectively retaining the first intent word or the second intent word according to the magnitude of the environmental noise" specifically includes: If the environmental noise is in the low-noise area, retain the first intent word; If the environmental noise is in the medium-noise area and the first intent word and the second intent word belong to the high-stability category, retain the first intent word; If the environmental noise is in the medium-noise area and the first intent word and the second intent word belong to the low-stability category, retain the second intent word; If the environmental noise is in the high-noise area, retain the second intent word; Wherein, the influence degree of the words in the high-stability category by the environmental noise is less than the influence degree of the words in the low-stability category by the environmental noise.
4. The speech recognition method according to claim 2 or 3, characterized in that The step of "selectively retaining the first intent word or the second intent word according to the judgment result" further includes: If the semantic similarity between the first intent word and the second intent word reaches the requirement, then retain the first intent word, or retain the second intent word, or randomly retain one of the first intent word and the second intent word.
5. A voice recognition system, characterized in that, The speech recognition system includes: A sound acquisition device configured to be able to collect speech information; An image acquisition device configured to be able to collect lip-reading information; An information processing device configured to be able to generate a first statement and a second statement respectively according to the speech information collected by the sound acquisition device and the lip-reading information collected by the image acquisition device and generate a final statement according to the first statement and the second statement; The information processing device includes: A sound information processing module configured to be able to generate the first statement according to the speech information; An image information processing module configured to be able to generate the second statement according to the lip-reading information; The statement analysis and processing module is configured to disassemble the first statement and the second statement respectively to obtain a plurality of first intention words of different categories and a plurality of second intention words of different categories, and to respectively determine whether the semantic similarity between the first intention words and the second intention words in the same category meets the requirements, and according to the judgment result, selectively retain the first intention words or the second intention words, and finally generate a final statement according to the finally retained first intention words and second intention words.
6. The speech recognition system according to claim 5, wherein The statement analysis and processing module is further configured to: When the semantic similarity between the first intention word and the second intention word meets the requirements, retain the first intention word, or retain the second intention word, or randomly retain one of the first intention word or the second intention word; When the semantic similarity between the first intention word and the second intention word does not meet the requirements, selectively retain the first intention word or the second intention word according to the magnitude of the ambient noise collected by the sound acquisition device.
7. The speech recognition system according to claim 6, wherein, The statement analysis and processing module is further configured to: In the case where the semantic similarity between the first intention word and the second intention word does not meet the requirements, When the ambient noise is in the low noise area, retain the first intention word; When the ambient noise is in the medium noise area and the first intention word and the second intention word belong to the high stability category, retain the first intention word; When the ambient noise is in the medium noise area and the first intention word and the second intention word belong to the low stability category, retain the second intention word; When the ambient noise is in the high noise area, retain the second intention word, wherein, the influence degree of the words in the high stability category by the ambient noise is less than the influence degree of the words in the low stability category by the ambient noise.
8. An electrical device, characterized in that, The electrical equipment includes the voice recognition system according to any one of claims 5 to 7.
Citation Information
Patent Citations
Lip language identification method, device and system, and intelligent glasses
CN108319912A
Voice input method and device, electronic equipment and storage medium
CN111045639A
Method and device for speech recognition
CN106157956A
Speech recognition method, apparatus and user equipment
CN106157957A