Speech semantic understanding method and device

By performing speech preprocessing and fuzzy correction of semantic understanding on the terminal server, combined with home and environmental information, the problem of inaccurate speech semantic understanding in the prior art is solved, and a more accurate user intention understanding and a better user experience are achieved.

CN114203179BActive Publication Date: 2025-06-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111266564.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-06-06
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

In the prior art, the pronunciation and semantic understanding is inaccurate and the user's intention cannot be accurately judged, resulting in poor recognition results experience.

Method used

By obtaining user voice information on the terminal server, preprocessing and sending it to the cloud server, combining family member information, living habits, device scenarios and environment information, the semantic understanding fuzzy correction model is trained, and the semantic understanding results returned by the cloud server are fuzzyly corrected to generate more accurate semantic understanding results.

Benefits of technology

It realizes an accurate understanding of user voice semantics, improves the accuracy of recognition results and user experience, reduces cloud data interaction, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114203179B_ABST
    Figure CN114203179B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for speech semantic understanding, which is used in a terminal server, wherein the method includes: obtaining user speech information and preprocessing it, obtaining the preprocessed speech information and sending it to a cloud server; obtaining a first semantic understanding result returned by the cloud server, wherein the first semantic understanding result is obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding text, and performing semantic understanding on the text; performing semantic understanding fuzzy correction on the first semantic understanding result to obtain a second semantic understanding result; and sending a first instruction to the cloud server and / or sending a second instruction to the local terminal according to the second semantic understanding result. The present invention combines environmental information, user family information, user usage habits and user input corpus through fuzzy processing of semantic understanding, realizes intelligent fuzzy recognition of semantic understanding, truly understands users, and provides customized services for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for understanding speech semantics. Background Art

[0002] With the gradual development of the artificial intelligence industry, semantic understanding based on pure audio-to-text conversion has become the mainstream intelligent semantic understanding application solution, but it has also brought various problems. For example, when a user says "ji rou", there is no way to determine whether the user is talking about "muscle" or "chicken". It does not fall into the category of reference disambiguation and cannot be accurately determined. For example, the user says "remind me to take medicine at 12 o'clock", but affected by the front-end algorithm, the user's corpus is not complete, and the actual input corpus is "take medicine at 2 o'clock", which will lead to unrecognition and even the opposite of the execution result.

[0003] Due to incomplete user corpus, the user's accent cannot be fully and correctly recognized, the user's usage or front-end algorithm causes inaccurate recognition after audio truncation, and the strong ambiguity of user corpus, the recognition result experience is not good. This problem has always been a pain point for users, but it is difficult to solve it uniformly from optimizing the front-end algorithm, optimizing the recognition effect or optimizing the user interaction prompts. This is because the technical solutions for semantic understanding in the existing technology cannot well assist in judging user intentions. Summary of the invention

[0004] The present invention provides a speech semantics understanding method and device, which are used to solve the defect of inaccurate understanding of user semantics in the prior art, and realize the construction of a speech semantics fuzzy understanding model to accurately understand the user's speech semantics.

[0005] In a first aspect, the present invention provides a method for understanding speech semantics, which is used in a terminal server and comprises:

[0006] Get user voice information;

[0007] Preprocessing the user voice information to obtain preprocessed voice information;

[0008] Sending the preprocessed voice information to a cloud server;

[0009] Obtaining the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding speech recognition text and performing semantic understanding on the speech recognition text;

[0010] Performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result;

[0011] Send a first instruction to the cloud server and / or send a second instruction to the local terminal according to the second semantic understanding result.

[0012] Further, according to a speech semantic understanding method provided by the present invention, wherein the step of performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result specifically includes:

[0013] Inputting the speech recognition result and the first semantic understanding result into a semantic understanding fuzzy correction model to obtain a second semantic understanding result;

[0014] The semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results.

[0015] Further, according to a speech semantic understanding method provided by the present invention, after sending a first instruction to a cloud server according to the second semantic understanding result, the method specifically includes:

[0016] The cloud server completes the corresponding instruction action and speech synthesis result according to the first instruction;

[0017] The cloud server feeds back the completion result of the instruction action to the local terminal, and sends the speech synthesis result to the speech broadcast device of the local terminal for speech broadcast.

[0018] In a second aspect, the present invention provides a speech semantic understanding method for a cloud server, comprising:

[0019] Acquiring user voice information, wherein the user voice information is obtained after the terminal server processes the acquired original user voice information;

[0020] Performing speech recognition on the user speech information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information;

[0021] Sending the speech recognition result and the first semantic understanding information to the terminal server;

[0022] Acquire a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to a speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzziness correction on the first semantic understanding information;

[0023] Execute corresponding command actions according to the command information and generate corresponding synthesized voice information;

[0024] The completion result of the instruction action and the synthesized voice information are sent to the local terminal.

[0025] Further, according to a speech semantic understanding method provided by the present invention, the step of obtaining a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server for performing semantic understanding fuzzy correction on the first semantic understanding information according to the speech recognition result to obtain a second semantic understanding result, specifically comprising:

[0026] Inputting the speech recognition result and the first semantic understanding information into the semantic understanding fuzzy correction model of the terminal server to obtain corresponding second semantic understanding information; wherein the semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results;

[0027] The terminal server generates corresponding first instruction information according to the second semantic understanding information and sends the first instruction information to the cloud server.

[0028] Furthermore, a method for understanding speech semantics provided by the present invention further comprises:

[0029] The terminal server generates corresponding second instruction information according to the second semantic understanding information, and sends the fourth instruction information to the corresponding local terminal.

[0030] In a third aspect, the present invention provides a speech semantic understanding device for a terminal server, comprising:

[0031] The first processing module is used to obtain user voice information;

[0032] A second processing module, used for preprocessing the user voice information to obtain preprocessed voice information;

[0033] A third processing module, used to send the pre-processed voice information to a cloud server;

[0034] A fourth processing module is used to obtain the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding speech recognition text and performing semantic understanding on the speech recognition text;

[0035] A fifth processing module, configured to perform semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result;

[0036] The sixth processing module is used to send a first instruction to the cloud server and / or send a second instruction to the local terminal according to the second semantic understanding result.

[0037] In a fourth aspect, the present invention provides a speech semantic understanding device for a cloud server, comprising:

[0038] A seventh processing module, configured to obtain user voice information, wherein the user voice information is obtained after the terminal server processes the obtained original user voice information;

[0039] an eighth processing module, configured to perform voice recognition on the user voice information to obtain a voice recognition result, and perform semantic understanding on the voice recognition result to obtain first semantic understanding information;

[0040] A ninth processing module, configured to send the speech recognition result and the first semantic understanding information to the terminal server;

[0041] a tenth processing module, configured to obtain a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to a speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzziness correction on the first semantic understanding information;

[0042] an eleventh processing module, configured to execute a corresponding instruction action according to the instruction information and generate corresponding synthesized speech information;

[0043] The twelfth processing module is used to send the completion result of the instruction action and the synthesized voice information to the local terminal.

[0044] In a fifth aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described speech semantic understanding methods are implemented.

[0045] In a sixth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described speech semantic understanding methods.

[0046] The present invention provides a method and device for speech semantic understanding, which is used in a terminal server, and obtains user speech information and performs preprocessing to obtain preprocessed speech information; the preprocessed speech information is sent to a cloud server to obtain a first semantic understanding result, wherein the first semantic understanding result is obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding text, and performing semantic understanding on the text; then, the first semantic understanding result is subjected to semantic understanding fuzzy correction to obtain a second semantic understanding result; then, according to the second semantic understanding result, a first instruction is sent to the cloud server and / or a second instruction is sent to the local terminal. The present invention combines environmental information, user family information, user usage habits and user input corpus through fuzzy processing of semantic understanding on the edge side of intelligent edge computing, realizes preprocessing and desensitization of semantic understanding data, and enables intelligent life to realize intelligent fuzzy recognition of data in a high-speed edge network, truly understand users, and provide customized services for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 This is a flow chart of a speech semantic understanding method provided by the present invention, which is used for a terminal server;

[0049] Figure 2 This is a flow chart of a speech semantic understanding method provided by the present invention, which is used for a cloud server;

[0050] Figure 3 This is a schematic diagram of the structure of a speech semantic understanding device provided by the present invention, which is used for a terminal server;

[0051] Figure 4 This is a structural diagram of a speech semantic understanding device provided by the present invention, which is used in a cloud server;

[0052] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Combine the following Figure 1 A speech semantic understanding method provided by the present invention is described, which is used in a terminal server and includes:

[0055] Step 110: Obtain user voice information;

[0056] Specifically, the user initiates human-computer voice interaction and sends a voice command to a voice receiving device, wherein the voice receiving device includes various devices capable of voice collection such as a microphone, which is not limited in the present invention.

[0057] Step 120: preprocessing the user voice information to obtain preprocessed voice information;

[0058] Specifically, in the present invention, after obtaining the voice information of the user, the terminal pre-processes the voice information, specifically including performing various processes such as noise removal and information amplification on the voice information.

[0059] Step 130: Sending the pre-processed voice information to a cloud server;

[0060] Specifically, in the present invention, the pre-processed voice information is sent to a cloud server in a limited or wireless manner, and the cloud server provides a general voice semantic understanding service.

[0061] Step 140: Obtaining the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding speech recognition text and performing semantic understanding on the speech recognition text;

[0062] Specifically, the cloud server in the present invention uses automatic speech recognition to convert the received voice information into text information. The goal of automatic speech recognition (ASR) technology is to enable computers to "dictate" continuous speech spoken by different people, which is commonly known as "speech dictation machine". It is a technology that realizes the conversion of "sound" to "text". Automatic speech recognition is also called speech recognition or computer speech recognition.

[0063] Then, semantic understanding is performed on the obtained text information to obtain the corresponding understanding results. Since the semantic understanding model in the cloud server is a general model that does not take into account the differences in the specific circumstances of each family, the semantic understanding results obtained may be different, or inaccurate or completely different. For example, when the user asks to play the movie I watch most often, there is no corresponding result in the cloud server, and the cloud server cannot execute the corresponding instruction.

[0064] Step 150: performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result;

[0065] Specifically, in the present invention, the result of semantic understanding of the cloud server is transmitted back to the local server. Since the local server is a fuzzy processing of semantic understanding on the edge side of intelligent edge computing, it combines environmental information, user family information, user usage habits and user input corpus. Due to the isolation of local data and cloud data, the user's privacy information will not be leaked in the semantic understanding data. At the same time, the fuzzy semantic understanding of the first semantic understanding result combined with environmental information, family user information, family user habits and other comprehensive judgments can well assist in providing a family model that is closer to the user's true intention, thereby giving the user the most suitable feedback for the user scenario, thereby correcting the first semantic understanding and obtaining a more practical and accurate semantic understanding result. Distance explanation, assuming that in the family, the mother educates the child to brush his teeth after a meal, then, in order to verify whether the child has followed the rules, the parent sends a voice message to the microphone, whether someone has brushed his teeth, etc., but the cloud server cannot recognize who someone is, so it is impossible to perform accurate semantic analysis, let alone execute the corresponding instructions and give an answer.

[0066] That is, the technical solution of the present invention forms a small family model of the terminal in the smart terminal by allowing the edge side device terminal to collect user information, family member habits, smart terminal device conditions connected to the edge side, environmental information and other aspects on the edge side. When there is ambiguity, punctuation, incomplete recognition, and recognition errors caused by accents in the user corpus, the result of semantic understanding is corrected through the small family model, thereby providing users with more intimate feedback. The real-time correction of the small family model does not require cloud operation. All small family models can be completed in real time on the edge side device terminal for changes in the family's recent status. The user's daily fixed habits and family member information are configured and stored on the edge side, the real-time edge records of home devices are on the edge side, the model training is completed on the edge side, and the full edge side of semantic understanding fuzzy correction is realized. This solution reduces family data interaction more efficiently in the cloud, protects the user's privacy data, sinks cloud resource capabilities, and integrates some cloud capabilities and some unrealizable functions on the edge side. Let home intelligence truly enter the era of edge-side intelligence.

[0067] Step 160: Send a first instruction to the cloud server and / or send a second instruction to the local terminal according to the second semantic understanding result.

[0068] Specifically, after obtaining the second semantic understanding result, corresponding instruction information is generated based on the semantic understanding result, some of which need to be completed by the cloud server, such as playing movies that are not stored in the local server. Some instructions are completed by the local terminal device, such as answering questions with private attributes such as whether to take medicine. Therefore, when generating instructions according to the second semantic understanding result in the present invention, different instructions are generated according to different subjects executing the instructions, and the instructions are sent to different execution subjects.

[0069] A speech semantic understanding method provided by the present invention is used for a terminal server, by obtaining user voice information and preprocessing, to obtain preprocessed voice information; the preprocessed voice information is sent to a cloud server to obtain a first semantic understanding result, wherein the first semantic understanding result is obtained by the cloud server performing speech recognition on the preprocessed voice information to obtain the corresponding text, and performing semantic understanding on the text; then, the first semantic understanding result is subjected to semantic understanding fuzzy correction to obtain a second semantic understanding result; then, according to the second semantic understanding result, a first instruction is sent to the cloud server and / or a second instruction is sent to the local terminal. The present invention combines environmental information, user family information, user usage habits and user input corpus through fuzzy processing of semantic understanding on the edge side of intelligent edge computing, to achieve preprocessing and desensitization of semantic understanding data, so that smart life can realize intelligent fuzzy recognition of data in a high-speed edge network, truly understand users, and provide customized services for users.

[0070] Further, according to a speech semantic understanding method provided by the present invention, wherein the step of performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result specifically includes:

[0071] Inputting the speech recognition result and the first semantic understanding result into a semantic understanding fuzzy correction model to obtain a second semantic understanding result;

[0072] The semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results.

[0073] Specifically, since the fuzzy understanding of speech semantics of the present invention can be based on the user's family information, the user's daily living habits data, the user's home device reporting data, indoor and outdoor environment information, the user's scene actively triggers the user terminal set on the edge side to train the edge side small model, namely the speech semantic fuzzy correction model. That is, the neural network model is trained through data such as user family information, living habits information and corresponding predetermined results. Moreover, in the embodiment of the present invention, the living habits data and the like are dynamic, and the speech semantic fuzzy correction model is continuously updated according to the dynamic data, so that it has the characteristics of keeping pace with the times, so that it can more accurately reflect the specific information changes of the user's family.

[0074] Then, the received speech recognition result and the first semantic understanding result are input into the dynamically learned speech semantic fuzzy correction model, and then the corresponding second semantic understanding result is input.

[0075] For example, a family member says to feed Xiao Hu some milk into the microphone. In the cloud server, the voice recognition of this sentence is fed Xiao Hu some milk. The first semantic understanding result is feed someone some milk, but it cannot fully understand what Xiao Hu is. Therefore, it cannot provide corresponding operation instructions. In the voice semantic fuzzy correction model in the terminal, it can be judged that Xiao Hu is a cat by combining family information, and then the voice result and the first semantic understanding result can be corrected to feed the kitten some milk. This is done by sending instructions to feed the cat the corresponding milk, and the amount of milk to be fed this time can also be determined based on the daily feeding amount.

[0076] Further, according to a speech semantic understanding method provided by the present invention, after sending a first instruction to a cloud server according to the second semantic understanding result, the method specifically includes:

[0077] The cloud server completes the corresponding instruction action and speech synthesis result according to the first instruction;

[0078] The cloud server feeds back the completion result of the instruction action to the local terminal, and sends the speech synthesis result to the speech broadcast device of the local terminal for speech broadcast.

[0079] Specifically, in the present invention, if the instruction determined by the voice information sent by the user needs to be completed by the cloud server, then the instruction is sent to the cloud server, and the cloud server completes the corresponding action according to the received instruction. At the same time, the cloud server performs voice synthesis on the relevant information in the process of executing the instruction, and sends the synthesis result to the terminal, which performs voice broadcast. For example, the user sends an instruction to watch a movie, but the movie is not stored in the local terminal, so the cloud server needs to provide corresponding services. In the process of searching for the movie, the cloud server generates corresponding voice information according to the search results, such as the movie is not found, the movie is successfully found, etc., and then transmits the corresponding voice information to the voice broadcast device of the terminal to broadcast the corresponding voice content.

[0080] Combination Figure 2 As shown, the present invention provides a speech semantic understanding method for a cloud server, comprising:

[0081] Step 210: Acquire user voice information, wherein the user voice information is obtained after the terminal processes the acquired original user voice information;

[0082] Specifically, the user initiates human-computer voice interaction and sends original voice information to the voice receiving device, wherein the voice receiving device includes various devices capable of voice collection such as microphones, which are not limited by the present invention. The original voice information is preprocessed, specifically including noise removal, information amplification and other processing of the voice information.

[0083] Step 220: performing speech recognition on the user speech information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information;

[0084] Specifically, the cloud server in the present invention uses automatic speech recognition to convert the received voice information into text information. The goal of automatic speech recognition (ASR) technology is to enable computers to "dictate" continuous speech spoken by different people, which is commonly known as "speech dictation machine". It is a technology that realizes the conversion of "sound" to "text". Automatic speech recognition is also called speech recognition or computer speech recognition.

[0085] Then, semantic understanding is performed on the obtained text information to obtain the corresponding understanding results. Since the semantic understanding model in the cloud server is a general model that does not take into account the differences in the specific circumstances of each family, the semantic understanding results obtained may be different, or inaccurate or completely different. For example, when the user asks to play the movie I watch most often, there is no corresponding result in the cloud server, and the cloud server cannot execute the corresponding instruction.

[0086] Step 230: Sending the speech recognition result and the first semantic understanding information to the terminal;

[0087] Specifically, in the present invention, the result of semantic understanding of the cloud server is transmitted back to the local terminal server.

[0088] Step 240: Acquire a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to the speech recognition result to obtain a second semantic understanding result by performing semantic understanding fuzziness correction on the first semantic understanding information;

[0089] Specifically, since the local server is used as the fuzzy processing of semantic understanding on the edge side of intelligent edge computing, it combines environmental information, user family information, user usage habits and user input corpus. Due to the isolation of local data and cloud data, the user's privacy information will not be leaked in the semantic understanding data. At the same time, the fuzzy semantic understanding of the first semantic understanding result combined with environmental information, family user information, family user habits and other comprehensive judgments can well assist in providing a family model that is closer to the user's true intention, thereby giving the user the most suitable feedback for the user scenario, thereby correcting the first semantic understanding and obtaining a more practical and accurate semantic understanding result. Distance explanation, assuming that in the family, the mother educates the child to brush his teeth after a meal, then, in order to verify whether the child has followed the rules, the parent sends a voice message to the microphone, whether someone has brushed his teeth, etc., but the cloud server cannot recognize who someone is, so it is impossible to perform accurate semantic analysis, let alone execute the corresponding instructions and give an answer.

[0090] That is, the technical solution of the present invention forms a small family model of the terminal in the smart terminal by allowing the edge side device terminal to collect user information, family member habits, smart terminal device conditions connected to the edge side, environmental information and other aspects on the edge side. When there is ambiguity, punctuation, incomplete recognition, and recognition errors caused by accents in the user corpus, the result of semantic understanding is corrected through the small family model, thereby providing users with more intimate feedback. The real-time correction of the small family model does not require cloud operation. All small family models can be completed in real time on the edge side device terminal for changes in the family's recent status. The user's daily fixed habits and family member information are configured and stored on the edge side, the real-time edge records of home devices are on the edge side, the model training is completed on the edge side, and the full edge side of semantic understanding fuzzy correction is realized. This solution reduces family data interaction more efficiently in the cloud, protects the user's privacy data, sinks cloud resource capabilities, and integrates some cloud capabilities and some unrealizable functions on the edge side. Let home intelligence truly enter the era of edge-side intelligence.

[0091] Step 250: executing corresponding instruction actions according to the instruction information, and generating corresponding synthesized speech information;

[0092] Specifically, in the present invention, if the instruction determined by the voice information sent by the user needs to be completed by the cloud server, then the instruction is sent to the cloud server, and the cloud server completes the corresponding action according to the received instruction. At the same time, the cloud server performs voice synthesis on the relevant information in the process of executing the instruction, and sends the synthesis result to the terminal, which performs voice broadcast. For example, the user sends an instruction to watch a movie, but the movie is not stored in the local terminal, so the cloud server needs to provide corresponding services, and in the process of the cloud server searching for the movie, the corresponding voice information is generated according to the search result, such as the movie is not found, the movie is successfully found, etc.

[0093] Step 260: Send the completion result of the instruction action and the synthesized voice information to the local terminal.

[0094] Specifically, as described in step 250, the corresponding voice information is further transmitted to the voice broadcasting device of the terminal to broadcast the corresponding voice content.

[0095] Further, according to a speech semantic understanding method provided by the present invention, the step of obtaining a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server for performing semantic understanding fuzzy correction on the first semantic understanding information according to the speech recognition result to obtain a second semantic understanding result, specifically comprising:

[0096] Inputting the speech recognition result and the first semantic understanding information into the semantic understanding fuzzy correction model of the terminal server to obtain corresponding second semantic understanding information; wherein the semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results;

[0097] The terminal server generates corresponding first instruction information according to the second semantic understanding information and sends the first instruction information to the cloud server.

[0098] Specifically, since the fuzzy understanding of speech semantics of the present invention can be based on the user's family information, the user's daily living habits data, the user's home device reporting data, indoor and outdoor environment information, the user's scene actively triggers the user terminal set on the edge side to train the edge side small model, namely the speech semantic fuzzy correction model. That is, the neural network model is trained through data such as user family information, living habits information and corresponding predetermined results. Moreover, in the embodiment of the present invention, the living habits data and the like are dynamic, and the speech semantic fuzzy correction model is continuously updated according to the dynamic data, so that it has the characteristics of keeping pace with the times, so that it can more accurately reflect the specific information changes of the user's family.

[0099] Then, the received speech recognition result and the first semantic understanding result are input into the dynamically learned speech semantic fuzzy correction model, and then the corresponding first semantic understanding result is input. Based on the first semantic understanding result, a corresponding instruction is generated and sent to the cloud server.

[0100] Furthermore, a method for understanding speech semantics provided by the present invention further comprises:

[0101] The terminal server generates corresponding second instruction information according to the second semantic understanding information, and sends the fourth instruction information to the corresponding local terminal.

[0102] Specifically, if the instruction generated by the second semantic understanding information is executed in the local terminal service, such instruction information is called the fourth instruction information. It does not need to be uploaded to the cloud server, but only needs to send the corresponding instruction directly to the corresponding local terminal, for example, playing downloaded songs, etc.

[0103] Combination Figure 3 As shown, the present invention provides a speech semantic understanding device for a terminal server, comprising:

[0104] The first processing module 31 is used to obtain user voice information;

[0105] The second processing module 32 is used to pre-process the user voice information to obtain pre-processed voice information;

[0106] The third processing module 33 is used to send the pre-processed voice information to the cloud server;

[0107] The fourth processing module 34 is used to obtain a first semantic understanding result returned by the cloud server, wherein the first semantic understanding result is obtained by the cloud server performing speech recognition on the pre-processed speech information to obtain a corresponding text, and performing semantic understanding on the text;

[0108] A fifth processing module 35, configured to perform semantic understanding fuzziness correction on the first semantic understanding result to obtain a second semantic understanding result;

[0109] The sixth processing module 36 is used to send a first instruction to the cloud server and / or send a second instruction to the local terminal according to the second semantic understanding result.

[0110] Since the device provided in the embodiment of the present invention can be used to execute the method described in the above embodiment, its working principle and beneficial effects are similar, so they are not described in detail here. For specific contents, please refer to the introduction of the above embodiment.

[0111] The present invention provides a speech semantic understanding device, which is used for a terminal server, and obtains user voice information and performs preprocessing to obtain preprocessed voice information; the preprocessed voice information is sent to a cloud server to obtain a first semantic understanding result, wherein the first semantic understanding result is obtained by the cloud server performing speech recognition on the preprocessed voice information to obtain the corresponding text, and performing semantic understanding on the text; then, the first semantic understanding result is subjected to semantic understanding fuzzy correction to obtain a second semantic understanding result; then, according to the second semantic understanding result, a first instruction is sent to the cloud server and / or a second instruction is sent to the local terminal. The present invention combines environmental information, user family information, user usage habits and user input corpus through fuzzy processing of semantic understanding on the edge side of intelligent edge computing, realizes preprocessing and desensitization of semantic understanding data, and enables intelligent life to realize intelligent fuzzy recognition of data in a high-speed edge network, truly understand users, and provide customized services for users.

[0112] Further, according to a speech semantic understanding device provided by the present invention, the fifth processing module 35 is specifically used for:

[0113] Inputting the speech recognition result and the first semantic understanding result into a semantic understanding fuzzy correction model to obtain a second semantic understanding result;

[0114] The semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results.

[0115] Further, according to a speech semantic understanding device provided by the present invention, the sixth processing module 36 is specifically used for:

[0116] The cloud server completes the corresponding instruction action and speech synthesis result according to the first instruction;

[0117] The cloud server feeds back the completion result of the instruction action to the local terminal, and sends the speech synthesis result to the speech broadcast device of the local terminal for speech broadcast.

[0118] Combination Figure 4 As shown, the present invention provides a speech semantic understanding device for a cloud server, comprising:

[0119] The seventh processing module 41 is used to obtain user voice information, wherein the user voice information is obtained after the terminal processes the obtained original user voice information;

[0120] An eighth processing module 42, configured to perform voice recognition on the user voice information to obtain a voice recognition result, and perform semantic understanding on the voice recognition result to obtain first semantic understanding information;

[0121] A ninth processing module 43, configured to send the speech recognition result and the first semantic understanding information to the terminal server;

[0122] A tenth processing module 44 is used to obtain a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to the speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzziness correction on the first semantic understanding information;

[0123] An eleventh processing module 45 is used to execute a corresponding instruction action according to the instruction information and generate corresponding synthesized speech information;

[0124] The twelfth processing module 46 is used to send the completion result of the instruction action and the synthesized voice information to the local terminal.

[0125] Since the device provided in the embodiment of the present invention can be used to execute the method described in the above embodiment, its working principle and beneficial effects are similar, so they are not described in detail here. For specific contents, please refer to the introduction of the above embodiment.

[0126] Further, according to a speech semantic understanding device provided by the present invention, the tenth processing module 44 is specifically used for:

[0127] Inputting the speech recognition result and the first semantic understanding information into the semantic understanding fuzzy correction model of the terminal server to obtain corresponding second semantic understanding information; wherein the semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results;

[0128] The terminal server generates corresponding first instruction information according to the second semantic understanding information and sends the first instruction information to the cloud server.

[0129] Furthermore, according to a speech semantic understanding device provided by the present invention, a seventh processing module is further included, wherein the seventh processing module is specifically used to:

[0130] The terminal server generates corresponding second instruction information according to the second semantic understanding information, and sends the fourth instruction information to the corresponding local terminal.

[0131] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a speech semantic understanding method for a terminal server, the method comprising: obtaining user speech information; preprocessing the user speech information to obtain preprocessed speech information; sending the preprocessed speech information to a cloud server; obtaining the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information and performing semantic understanding on the speech recognition text; performing semantic understanding fuzzy correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result; sending a first instruction to the cloud server and / or sending a second instruction to the local terminal according to the second semantic understanding result.

[0132] Alternatively, a method for speech semantic understanding is used in a cloud server, the method comprising: obtaining user voice information, wherein the user voice information is obtained by a terminal server after processing the original user voice information obtained; performing speech recognition on the user voice information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information; sending the speech recognition result and the first semantic understanding information to the terminal server; obtaining a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to the speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzzy correction on the first semantic understanding information to obtain a second semantic understanding result; executing a corresponding instruction action according to the instruction information, and generating corresponding synthesized speech information; and sending the completion result of the instruction action and the synthesized speech information to a local terminal.

[0133] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0134] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a speech semantic understanding method provided by the above methods, which is used for a terminal server, and the method includes: obtaining user voice information; preprocessing the user voice information to obtain preprocessed voice information; sending the preprocessed voice information to a cloud server; obtaining a speech recognition text and a first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the corresponding speech recognition text obtained after the cloud server performs speech recognition on the preprocessed voice information, and a first semantic understanding result obtained by performing semantic understanding on the speech recognition text; performing semantic understanding fuzzy correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result; and sending a first instruction to the cloud server and / or sending a second instruction to the local terminal according to the second semantic understanding result.

[0135] Alternatively, a method for speech semantic understanding is used in a cloud server, the method comprising: obtaining user voice information, wherein the user voice information is obtained by a terminal server after processing the original user voice information obtained; performing speech recognition on the user voice information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information; sending the speech recognition result and the first semantic understanding information to the terminal server; obtaining a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to the speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzzy correction on the first semantic understanding information to obtain a second semantic understanding result; executing a corresponding instruction action according to the instruction information, and generating corresponding synthesized speech information; and sending the completion result of the instruction action and the synthesized speech information to a local terminal.

[0136] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a speech semantic understanding method provided above, which is used for a terminal server, and the method includes: obtaining user voice information; preprocessing the user voice information to obtain preprocessed voice information; sending the preprocessed voice information to a cloud server; obtaining a speech recognition text and a first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the corresponding speech recognition text obtained after the cloud server performs speech recognition on the preprocessed voice information, and a first semantic understanding result obtained by performing semantic understanding on the speech recognition text; performing semantic understanding fuzzy correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result; and sending a first instruction to the cloud server and / or a second instruction to the local terminal according to the second semantic understanding result.

[0137] Alternatively, a method for speech semantic understanding is used in a cloud server, the method comprising: obtaining user voice information, wherein the user voice information is obtained by a terminal server after processing the original user voice information obtained; performing speech recognition on the user voice information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information; sending the speech recognition result and the first semantic understanding information to the terminal server; obtaining a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to the speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzzy correction on the first semantic understanding information to obtain a second semantic understanding result; executing a corresponding instruction action according to the instruction information, and generating corresponding synthesized speech information; and sending the completion result of the instruction action and the synthesized speech information to a local terminal.

[0138] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0139] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for understanding speech semantics, It is characterized in that For terminal servers, including: Get user voice information; Preprocessing the user voice information to obtain preprocessed voice information; Sending the preprocessed voice information to a cloud server; Obtaining the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding speech recognition text and performing semantic understanding on the speech recognition text; Performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result; Sending a first instruction to a cloud server and / or sending a second instruction to a local terminal according to the second semantic understanding result; The performing semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain the second semantic understanding result specifically includes: Inputting the speech recognition result and the first semantic understanding result into a semantic understanding fuzzy correction model to obtain a second semantic understanding result; The semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results.

2. The method for speech semantic understanding according to claim 1, It is characterized in that After sending the first instruction to the cloud server according to the second semantic understanding result, the method specifically includes: The cloud server completes the corresponding instruction action and speech synthesis result according to the first instruction; The cloud server feeds back the completion result of the instruction action to the local terminal, and sends the speech synthesis result to the voice broadcast device of the local terminal for voice broadcast.

3. A method for understanding speech semantics, It is characterized in that For cloud servers, including: Acquiring user voice information, wherein the user voice information is obtained after the terminal server processes the acquired original user voice information; Performing speech recognition on the user speech information to obtain a speech recognition result, and performing semantic understanding on the speech recognition result to obtain first semantic understanding information; Sending the speech recognition result and the first semantic understanding information to the terminal server; Acquire a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to a speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzziness correction on the first semantic understanding information; Execute corresponding command actions according to the command information and generate corresponding synthesized voice information; Sending the completion result of the command action and the synthesized voice information to the local terminal; The obtaining of the first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server for performing semantic understanding fuzziness correction on the first semantic understanding information according to the speech recognition result to obtain a second semantic understanding result, specifically comprising: Inputting the speech recognition result and the first semantic understanding information into the semantic understanding fuzzy correction model of the terminal server to obtain corresponding second semantic understanding information; wherein the semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results; The terminal server generates corresponding first instruction information according to the second semantic understanding information and sends the first instruction information to the cloud server.

4. The method for speech semantic understanding according to claim 3, It is characterized in that Also includes: The terminal server generates corresponding second instruction information according to the second semantic understanding information, and sends the second instruction information to the corresponding local terminal.

5. A speech semantic understanding device, It is characterized in that For terminal servers, including: The first processing module is used to obtain user voice information; A second processing module, used for preprocessing the user voice information to obtain preprocessed voice information; A third processing module, used to send the pre-processed voice information to a cloud server; A fourth processing module is used to obtain the speech recognition text and the first semantic understanding result returned by the cloud server, wherein the speech recognition text and the first semantic understanding result are the first semantic understanding result obtained by the cloud server performing speech recognition on the preprocessed speech information to obtain the corresponding speech recognition text and performing semantic understanding on the speech recognition text; A fifth processing module, configured to perform semantic understanding fuzziness correction on the first semantic understanding result according to the speech recognition result to obtain a second semantic understanding result; A sixth processing module, configured to send a first instruction to a cloud server and / or send a second instruction to a local terminal according to the second semantic understanding result; The fifth processing module is specifically used for: Inputting the speech recognition result and the first semantic understanding result into a semantic understanding fuzzy correction model to obtain a second semantic understanding result; The semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results.

6. A speech semantic understanding device, It is characterized in that For cloud servers, including: A seventh processing module, configured to obtain user voice information, wherein the user voice information is obtained after the terminal server processes the obtained original user voice information; an eighth processing module, configured to perform voice recognition on the user voice information to obtain a voice recognition result, and perform semantic understanding on the voice recognition result to obtain first semantic understanding information; A ninth processing module, configured to send the speech recognition result and the first semantic understanding information to the terminal server; a tenth processing module, configured to obtain a first instruction sent by the terminal server, wherein the first instruction is generated by the terminal server according to a speech recognition result to obtain a second semantic understanding result obtained by performing semantic understanding fuzziness correction on the first semantic understanding information; an eleventh processing module, configured to execute a corresponding instruction action according to the instruction information and generate corresponding synthesized speech information; A twelfth processing module, used for sending the completion result of the instruction action and the synthesized voice information to the local terminal; The tenth processing module is specifically used for: Inputting the speech recognition result and the first semantic understanding information into the semantic understanding fuzzy correction model of the terminal server to obtain corresponding second semantic understanding information; wherein the semantic understanding fuzzy correction model is trained based on family member information, family member living habit information, device scene information, indoor and outdoor environment information, and predetermined output results; The terminal server generates corresponding first instruction information according to the second semantic understanding information and sends the first instruction information to the cloud server.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the speech semantic understanding method as described in any one of claims 1 to 4 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the speech semantic understanding method as claimed in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Man-machine interaction method based on artificial intelligence voice

    CN111145746A

  • Voice recognition method, device, equipment, system and storage medium

    CN113436614A