Semantic Recognition Method and Device

By combining the context, device status and user behavior habits, the user's intentions are identified, and the user's voice signal intent is solved, and the intent hit rate and voice interaction efficiency are achieved.

CN113806470BActive Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010543520.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-15
Publication Date
2025-05-27
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

When existing voice assistants understand the user's voice signal's intention, they are prone to misidentification due to mismatch of whitelist databases, and cannot guarantee that the intention of each corpus can be accurately identified.

Method used

By combining the context, the status of the electronic device and the user's behavioral habits, the intention of the voice signal sent by the user is recognized to improve the accuracy of intention recognition. Specific methods include converting historical dialogue records, device status information and user behavior habits into coded vectors, and fusing these vectors to determine the target intention of the speech signal.

Benefits of technology

It significantly improves the probability of intention hitting, improves the efficiency of voice interaction, and reduces the situation of misidentification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113806470B_ABST
    Figure CN113806470B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a semantic recognition method and device, which relate to the field of artificial intelligence technology, and particularly to the field of semantic recognition technology. In this method, in the nth round of conversation, the electronic device sends the voice signal sent by the user in this round of conversation to the server, and the server converts the voice signal into a text to be processed. Based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits, and the text to be processed, the target intention of the text to be processed is determined from the whitelist database, the instruction corresponding to the target intention is executed, and the execution result is fed back to the user. In this process, for the multi-round conversation scenario, since the expression of the user in the current round of conversation, that is, the text to be processed, is no longer the only basis for determining the intention, the server's simultaneous consideration of the text to be processed, the historical conversation record, the device status information, and the user behavior habits can significantly improve the hit probability of the intention and improve the voice interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a semantic recognition method and apparatus. Background Art

[0002] Currently, with the rapid development of artificial intelligence (AI), more and more electronic devices are installed with voice assistants (VAs). The voice assistant can help users complete many functions such as opening applications, playing music, hailing a taxi, booking tickets, etc.

[0003] Generally, a voice assistant includes three parts: a speech engine, a semantic engine, and a dialogue management system. Among them, the speech engine is used to convert the speech uttered by the user into text, the semantic engine is used to understand the user's intention according to the text, and the dialogue management system is used to make the communication between the user and the voice assistant smoother. Among them, the core of the semantic engine is natural language understanding (NLU). Since Chinese characters can express a large number of contents, the intention of the same sentence may be different in different scenarios. Therefore, NLU cannot guarantee accurate recognition of the intention of each piece of corpus. To avoid misrecognition, a common practice is to establish a whitelist database for special corpus, and the corresponding relationship between the special corpus and the intention is stored in the whitelist database. For example, the special corpus is "Beijing Beijing", and the intention of this special corpus is a TV drama. In this way, after the user emits a voice signal of "Beijing Beijing", the voice assistant can query the whitelist database to determine that the user's intention is to search for the TV drama "Beijing Beijing".

[0004] However, the above whitelist database is prone to incorrect matching. For example, when the real intention of the user's voice signal of "Beijing Beijing" is to search for the weather in Beijing, the voice assistant understands that the user wants to search for the song "Beijing Beijing". Therefore, how to correctly understand the intention of the user's voice signal is regarded as an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of this application provide a semantic recognition method and apparatus, which recognize the intention of the voice signal emitted by the user by combining the context, the state of the application on the electronic device, and the user's behavior habits, and improve the accuracy of intention recognition.

[0006] In a first aspect, an embodiment of the present application provides a semantic recognition method. This method is applied to a server and can also be applied to a chip on the server. Hereinafter, taking the application to the server as an example, this method will be described: In the nth round of conversation between the user and the electronic device, the electronic device sends the voice signal sent by the user in this round of conversation to the server. The server converts the voice signal into a text to be processed, and based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits, and the text to be processed, determines the target intention of the text to be processed from the whitelist database, executes the instruction corresponding to the target intention, and feeds back the execution result to the user. In this process, for the multi-round conversation scenario, since the expression of the user in the current round of conversation, that is, the text to be processed, is no longer the only basis for determining the intention, the server's simultaneous consideration of the text to be processed, the historical conversation record, the device status information, and the user behavior habits can significantly improve the hit probability of the intention and improve the voice interaction efficiency.

[0007] In a feasible design, when the server determines the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, it first determines the first coding vector according to the historical conversation record and the text to be processed, determines the second coding vector according to the device status information, and determines the third coding vector according to the user behavior habits of the user. Then, the server fuses the first coding vector, the second coding vector, and the third coding vector to obtain a fusion vector, and determines the target intention of the text to be processed from the whitelist database according to the fusion vector. By adopting this solution, the purpose of converting the historical conversation record, the device status information, or the user behavior habits into coding vectors is achieved.

[0008] In a feasible design, when the server determines the first coding vector according to the historical conversation record and the text to be processed, it determines the semantic vector of each word in the user's statement in the (n - 1)th round of conversation according to the historical conversation record. The user's statement is the text corresponding to the voice signal sent by the user in the (n - 1)th round of conversation. It determines the contribution distribution of each word in the user's statement to the semantics of the user's statement, and determines the semantic vector of the user's statement in the (n - 1)th round of conversation according to the semantic vector of each word in the user's statement and the contribution distribution of each word in the user's statement to the semantics of the user's statement. Then, it determines the first coding vector according to the semantic vector of the user's statement in the (n - 1)th round of conversation and the word vector corresponding to the text to be processed. By adopting this solution, the historical conversation record is used as a basis in the process of determining the target intention, making the finally determined target intention more accurate.

[0009] In a feasible design, when the server determines the semantic vector of each word in the user's statement in the (n - 1)-th round of conversation based on the historical conversation record, it first determines the word vector of the user's statement in the (n - 1)-th round of conversation. Then, using a long short-term memory network (LSTM), it determines the past semantics and future semantics of each word in the word vector. Based on the past semantics and future semantics of each word in the user's statement in the (n - 1)-th round of conversation, it determines the semantic vector of the corresponding word in the user's statement in the (n - 1)-th round of conversation. By adopting this solution, the historical conversation record is used as a basis in the process of determining the target intent, making the finally determined target intent more accurate.

[0010] In a feasible design, when the server determines the second encoding vector based on the device status information, it performs one-hot encoding on the software and hardware on the electronic device according to the running status of the software and hardware on the electronic device to obtain the second encoding vector. The dimension of the second encoding vector is the same as the number of software and hardware on the electronic device, and the running status includes on or off. By adopting this solution, the device status information is used as a basis in the process of determining the target intent, improving the accuracy of the target intent.

[0011] In a feasible design, the dimension of the above-mentioned third encoding vector is the same as the number of user behavior habits.

[0012] In a feasible design, when the server fuses the first encoding vector, the second encoding vector, and the third encoding vector to obtain a fusion vector, it inputs the first encoding vector into the first autoencoder to obtain a first intermediate layer vector, inputs the second encoding vector into the second autoencoder to obtain a second intermediate layer vector, inputs the third encoding vector into the third autoencoder to obtain a third intermediate layer vector, and fuses the first intermediate layer vector, the second intermediate layer vector, and the third intermediate layer vector to obtain a concatenated vector. By adopting this solution, the historical conversation record, the device status information, and the user behavior habits are considered simultaneously in the process of determining the target intent, improving the accuracy of the target intent.

[0013] In a feasible design, when the server determines the target intent of the text to be processed from the whitelist database based on the fusion vector, it uses the concatenated vector to determine the matching probability of each intent among the multiple intents corresponding to the text to be processed, determines whether the matching probability of each intent exceeds the threshold of the corresponding intent, and takes the intent whose matching probability exceeds the threshold of the corresponding intent as the target intent of the text to be processed. By adopting this solution, the target intent is accurately determined by determining the probabilities of the multiple intents corresponding to the text to be processed.

[0014] In a feasible design, before determining the target intent of the text to be processed from the whitelist database based on the historical conversation records between the user and the electronic device, the device status information of the electronic device, and the user behavior habits of the user, the server also receives an indication message sent by the electronic device, and the indication message carries the device status information and / or the user behavior habits. By adopting this solution, the purpose of accurately determining the user's intent by using multi-source data is achieved by obtaining the device status and / or the user behavior habits.

[0015] In a feasible design, before determining the target intent of the text to be processed from the whitelist database based on the historical conversation records between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, the server also performs data cleaning on at least one of the historical conversation records, the device status information, or the user behavior habits. By adopting this solution, through data cleaning, the invalid data in the historical conversation records, the device status information, or the user behavior habits is filtered, reducing the data processing volume.

[0016] In a feasible design, after executing the instruction corresponding to the target intent, the execution result corresponding to the instruction is also determined and sent to the electronic device.

[0017] In a second aspect, an embodiment of the present application provides a semantic recognition device, including:

[0018] A transceiver unit, configured to receive a voice signal from the electronic device, where the voice signal is the voice signal sent by the user in the nth round of conversation between the user and the electronic device, n≥2 and is an integer.

[0019] A processing unit, configured to convert the voice signal into a text to be processed, and determine the target intent of the text to be processed from a whitelist database based on the historical conversation records between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed. The device status information is used to indicate the status of the software and hardware on the electronic device, and the user behavior habits are used to indicate the habits of the user using the software and hardware on the electronic device. The whitelist database stores the corresponding relationship between at least one candidate text and the intent of each candidate text in the at least one candidate text. Each candidate text in the at least one candidate text corresponds to at least one intent, and the at least one candidate text includes the text to be processed, and execute the instruction corresponding to the target intent.

[0020] In a feasible design, when determining the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, the processing unit is configured to determine a first encoding vector according to the historical conversation record and the text to be processed, determine a second encoding vector according to the device status information, determine a third encoding vector according to the user behavior habits of the user, fuse the first encoding vector, the second encoding vector, and the third encoding vector to obtain a fused vector, and determine the target intention of the text to be processed from the whitelist database according to the fused vector.

[0021] In a feasible design, when determining the first encoding vector according to the historical conversation record and the text to be processed, the processing unit is configured to determine the semantic vector of each word in the user's statement in the (n - 1)-th round of conversation according to the historical conversation record, where the user's statement is the text corresponding to the voice signal sent by the user in the (n - 1)-th round of conversation, determine the contribution distribution of each word in the user's statement to the semantics of the user's statement, determine the semantic vector of the user's statement in the (n - 1)-th round of conversation according to the semantic vector of each word in the user's statement and the contribution distribution of each word in the user's statement to the semantics of the user's statement, and determine the first encoding vector according to the semantic vector of the user's statement in the (n - 1)-th round of conversation and the word vector corresponding to the text to be processed.

[0022] In a feasible design, when determining the semantic vector of each word in the user's statement in the (n - 1)-th round of conversation according to the historical conversation record, the processing unit is configured to determine the word vector of the user's statement in the (n - 1)-th round of conversation, use a long short-term memory network (LSTM) to determine the past semantics and future semantics of each word in the word vector, and determine the semantic vector of the corresponding word in the user's statement in the (n - 1)-th round of conversation according to the past semantics and the future semantics of each word in the user's statement in the (n - 1)-th round of conversation.

[0023] In a feasible design, when determining the second encoding vector according to the device status information, the processing unit is configured to perform one-hot encoding on the software and hardware on the electronic device according to the running status of the software and hardware on the electronic device to obtain the second encoding vector, where the dimension of the second encoding vector is the same as the number of software and hardware on the electronic device, and the running status includes on or off.

[0024] In a feasible design, the dimension of the third encoding vector is the same as the number of user behavior habits

[0025] In a feasible design, when the processing unit fuses the first coding vector, the second coding vector, and the third coding vector to obtain a fused vector, it is configured to input the first coding vector into a first autoencoder to obtain a first intermediate layer vector, input the second coding vector into a second autoencoder to obtain a second intermediate layer vector, input the third coding vector into a third autoencoder to obtain a third intermediate layer vector, and fuse the first intermediate layer vector, the second intermediate layer vector, and the third intermediate layer vector to obtain the concatenated vector.

[0026] In a feasible design, when the processing unit determines the target intent of the text to be processed from the whitelist database according to the fused vector, it is configured to use the concatenated vector to determine the matching probability of each intent among the multiple intents corresponding to the text to be processed, determine whether the matching probability of each intent exceeds the threshold of the corresponding intent, and use the intent whose matching probability exceeds the threshold of the corresponding intent as the target intent of the text to be processed.

[0027] In a feasible design, before the processing unit determines the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, and the user behavior habits of the user, the transceiver unit is further configured to receive indication information sent by the electronic device, where the indication information carries the device status information and / or the user behavior habits.

[0028] In a feasible design, before the processing unit determines the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, it is further configured to perform data cleaning on at least one of the historical conversation record, the device status information, or the user behavior habits.

[0029] In a feasible design, after the processing unit executes the instruction corresponding to the target intent, it is further configured to determine the execution result corresponding to the instruction; the transceiver unit is further configured to send the execution result to the electronic device.

[0030] In a fifth aspect, an embodiment of the present application provides a semantic recognition device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the semantic recognition device implements the method in any possible implementation manner of the first aspect or the first aspect above.

[0031] Sixth aspect, an embodiment of the present application provides a chip, including: a logic circuit and an input interface, wherein the input interface is used to obtain data to be processed, and the logic circuit is used to execute the method described in any item of the first aspect on the data to be processed to obtain processed data.

[0032] In a feasible design, the chip further includes: an output interface, and the output interface is used to output the processed data.

[0033] Seventh aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium is used to store a program, and the program is used to execute the method described in any item of the first aspect when executed by a processor.

[0034] Eighth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a semantic recognition device, the semantic recognition device is enabled to execute the method described in any item of the first aspect.

[0035] The semantic recognition method and device provided by the embodiments of the present application relate to the field of artificial intelligence technology, and particularly to the field of semantic recognition technology. In the semantic recognition method, in the nth round of conversation between the user and the electronic device, the electronic device sends the voice signal sent by the user in this round of conversation to the server, and the server converts the voice signal into text to be processed. Based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits, and the text to be processed, the target intention of the text to be processed is determined from the whitelist database, and the instruction corresponding to the target intention is executed and the execution result is fed back to the user. In this process, for the multi-round conversation scenario, since the expression of the user in the current round of conversation, that is, the text to be processed, is no longer the only basis for determining the intention, the server considering the text to be processed, the historical conversation record, the device status information, and the user behavior habits at the same time can significantly improve the hit probability of the intention and improve the voice interaction efficiency. Description of the Drawings

[0036] Figure 1 is a network schematic diagram of the semantic recognition method provided by the embodiment of the present application;

[0037] Figure 2 is a flowchart of the semantic recognition method provided by the embodiment of the present application;

[0038] Figure 3 is another flowchart of the semantic recognition method provided by the embodiment of the present application;

[0039] Figure 4 is a schematic diagram of the process of multi-source feature encoding and multi-source feature fusion in the semantic recognition method provided by the embodiment of the present application;

[0040] Figure 5Schematic structural diagram of a semantic recognition device provided by an embodiment of the present application;

[0041] Figure 6 Schematic structural diagram of another semantic recognition device provided by an embodiment of the present application. Detailed implementation manners

[0042] The voice assistant on an electronic device aims to help users complete functions such as playing music, hailing a taxi, booking tickets, etc. through human-machine voice interaction, thus simplifying the original multi-round manual operations into a single voice command, significantly optimizing the user experience, reducing the difficulty of human-machine interaction, and facilitating the user to use the electronic device quickly and conveniently.

[0043] Generally, a voice assistant includes three parts: a voice engine, a semantic engine, and a dialogue management system. Among them, the core of the semantic engine is the NLU. Since Chinese characters can express a wide variety of contents, the intention of the same sentence may be different in different scenarios. Therefore, the NLU cannot guarantee accurate recognition of the intention of each piece of corpus. Such corpus that may lead to misrecognition is called special corpus. For example, the most common intention of the corpus "Beijing Beijing" is a city, but when the TV drama "Beijing Beijing" is very popular, the intention of this corpus is a TV drama. To ensure that the voice assistant can recognize special corpus, a common approach is to establish a whitelist database for storing the corresponding relationship between special corpus and intentions. For example, the special corpus is "Beijing Beijing", and the intention of this special corpus is a TV drama. In this way, after the user sends the voice signal of "Beijing Beijing", the voice assistant can first query the whitelist database to determine that the user's intention is to search for the TV drama "Beijing Beijing".

[0044] However, the above whitelist database matching method is prone to incorrect matching. Taking "Beijing Beijing" as an example, when the TV drama "Beijing Beijing" is very popular, in order to prioritize the development of this intention, the whitelist database is usually used for intention matching, which will trigger the following error scenarios:

[0045] Scenario 1:

[0046] In the first round of conversation, the user says: "What's the weather like tomorrow?" True intention

Weather

[0047] The machine says: "It will be sunny in Tianjin tomorrow." Reasoning graph

Weather

[0048] In the second round of conversation, the user says: "What about Beijing Beijing?" True intention

Weather

[0049] The machine says: "Found the TV drama 'Beijing Beijing' for you." Whitelist intention

Video

[0050] Scenario 2:

[0051] In the first round of conversation, the user said, "Help me translate this." True intention: [Translation]

[0052] The machine said, "What do you want to translate?" True intention: [Translation]

[0053] In the second round of conversation, the user said, "Beijing, Beijing." True intention: [Translation]

[0054] The machine said, "Found the TV drama 'Beijing, Beijing' for you." Whitelist intention: [Video]

[0055] According to the above, it can be seen that the traditional whitelist database matching method is prone to incorrect matching, and the intentions in the whitelist database are single. Obviously, when a special corpus has multiple intentions, it cannot be correctly matched.

[0056] In view of this, the embodiment of the present application provides a semantic recognition method, which recognizes the intention of the voice signal sent by the user by combining the context, the state of the applications on the electronic device, and the user's behavior habits, so as to improve the accuracy of intention recognition.

[0057] Figure 1 It is a network schematic diagram of the semantic recognition method provided by the embodiment of the present application. Please refer to Figure 1 , this network architecture includes an electronic device 1 and a server 2, and a network connection is established between the electronic device 1 and the server 2. Among them, a client of the voice assistant is installed on the electronic device 1, and the electronic device 1 uses the client to send the voice signal sent by the user to the server 2.

[0058] The server 2 has components such as a voice engine, a semantic engine, and a dialogue management system of the voice assistant. The server can obtain the state of the software and / or hardware on the electronic device from the electronic device 1, as well as the user's behavior habits of using the software and / or hardware on the electronic device 1. Among them, the software on the electronic device refers to various application programs (applications, APPs) installed on the electronic device, and the state of the software is, for example, APP is on, APP is off, etc.; the hardware on the electronic device includes, but is not limited to, the left channel, right channel, lens, light, etc. on the electronic device, and the state of the hardware includes on, off, the intensity of the light, etc.

[0059] The server 2 also pre-stores a whitelist database, which stores the corresponding relationship between at least one candidate text and the intention of each candidate text, and each candidate text corresponds to at least one intention. For example, a candidate text is "Beijing, Beijing", and the intentions corresponding to this candidate text include Figure 1 , city; Figure 2 , One , a song; Figure 3 ,One a TV drama

[0060] During the conversation between the user and the electronic device, for each round of conversation, the server stores the user's statement and the machine's statement in that round of conversation. For example, in the first round of conversation, the user's statement is "What's the weather like tomorrow?", and the machine says "It will be sunny in Tianjin tomorrow"; in the second round of conversation, the user's statement is "How about Beijing?", then the user's statement and the machine's statement in the first round are stored in the historical conversation record. For the user's statement in the current round, that is, the text to be processed, the server identifies the intention of the text to be processed based on the historical conversation record, the device status information of the electronic device, and the user's behavior habits.

[0061] After the server 2 identifies the target intention for the text to be processed, it executes the instruction according to the target intention and feeds back the execution result to the electronic device 1. Taking the text to be processed "How about Beijing?" as an example, when the target intention is a TV drama, the server searches for relevant videos of the TV drama and sends the link of the video to the electronic device 1.

[0062] Figure 1 In this case, the electronic device 1 is a desktop terminal or a mobile terminal, etc. The desktop terminal is, for example, a computer or a self-service device, and the mobile terminal is a mobile phone, a tablet computer, a laptop computer, etc.; the server 2 is an independently deployed server or a server cluster composed of multiple servers, etc.

[0063] Figure 2 It is a flowchart of the semantic recognition method provided by the embodiment of the present application. This embodiment is described from the perspective of the server. This embodiment includes:

[0064] 101. Receive a voice signal from the electronic device.

[0065] Among them, the voice signal is the voice signal sent by the user in the nth round of conversation between the user and the electronic device, where n≥2 and n is an integer.

[0066] The embodiment of the present application is applicable to any round of conversation after the first round of conversation. This is because in the first round of conversation, no historical conversation record has been collected on the server. The user sends a voice signal within the voice perception range of the electronic device 1, and the electronic device collects the voice signal and sends it to the server.

[0067] 102. Convert the voice signal into text to be processed.

[0068] In the nth round of conversation, after the server receives the voice signal from the electronic device, it uses a voice recognition model, etc. to convert the voice signal into text to be processed.

[0069] 103. Determine the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user's behavior habits, and the text to be processed.

[0070] In the embodiments of the present application, the historical conversation record is the conversation record of one or more rounds of conversations before the nth round of conversation. For example, the historical conversation record is the conversation record of the (n - 1)th round of conversation; for another example, the historical conversation record is the conversation record from the 1st round to the (n - 1)th round of conversation; for yet another example, the historical conversation record is the conversation record from the (n - 3)th round to the (n - 1)th round of conversation.

[0071] In the dialogue management stage, the semantic engine on the server uses the valid data in the current round and the previous rounds of conversations as the context, and obtains the device status information of the electronic device and the user's behavior habits. Among them, the device status information is used to indicate the status of the software and hardware on the electronic device. The software includes, for example, video playback APP, music playback APP, novel reading APP, etc. The status of the software refers to whether the APP is turned on or off, etc.; the hardware includes the left channel, right channel, etc. of the electronic device, and the status of the hardware includes, but is not limited to, on, off, light intensity, volume level, etc. The user's behavior habits include the user's operating system for the software and / or hardware. Taking the video playback APP as an example of the software, the user's behavior habit is to open the video playback APP -> play a certain TV series -> exit the video playback APP. Taking the left channel as an example of the hardware, the user's behavior habit is that when the user uses the music playback APP, the left channel is turned on and the right channel is turned off.

[0072] Then, the server determines the target intent of the text to be processed from the whitelist database based on the context, the device status information of the electronic device, and the user's behavior habits.

[0073] In this process, based on the context, device status information, and user behavior habits, a multi-dimensional matching operation is performed on the user's statement, that is, the text to be processed, which solves the problem of mis-matching caused by the method of only relying on text for matching, and the matching method is more flexible and efficient.

[0074] 104. Execute the instruction corresponding to the target intent.

[0075] After the server identifies the current round, that is, the target intent corresponding to the text to be processed corresponding to the user's statement in the nth round of conversation, it executes an instruction to respond to the user. For example, in the first round of conversation, the user says, "What's the weather like tomorrow?" and the machine says, "It will be sunny in Tianjin tomorrow." In the second round of conversation, the user says, "How about Beijing?" The machine says, "It will be sunny in Beijing tomorrow." Another example, in the first round of conversation, the user says, "I want to watch TV." The machine says, "What kind of TV do you want to watch?" In the second round of conversation, the user says, "How about Beijing?" The machine says, "The TV drama 'Beijing' is currently a hit. I've found the TV drama 'Beijing' for you."

[0076] In the semantic recognition method provided by the embodiments of the present application, in the nth round of conversation between the user and the electronic device, the electronic device sends the voice signal sent by the user in this round of conversation to the server. The server converts the voice signal into text to be processed, and determines the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits, and the text to be processed, executes the instruction corresponding to the target intent, and feeds back the execution result to the user. In this process, for the multi-round conversation scenario, since the expression of the user in the current round of conversation, that is, the text to be processed is no longer the only basis for determining the intent, the server's simultaneous consideration of the text to be processed, the historical conversation record, the device status information, and the user behavior habits can significantly improve the hit probability of the intent and improve the voice interaction efficiency.

[0077] For the detailed process of step 103 in the above embodiment, reference can be made to Figure 3 。 Figure 3 is another flowchart of the semantic recognition method provided by the embodiments of the present application. Please refer to Figure 3 ,this embodiment includes:

[0078] 201. Obtain the context, the device status information of the electronic device, and the user behavior habits.

[0079] Among them, the context refers to the effective data obtained by the server based on the user's statement in the nth round of conversation and the previous rounds of conversation. This effective data, for example, the inferred user intent, such as weather, etc.

[0080] 202. Data cleaning.

[0081] In this step, the server performs data cleaning on the context, the device status information of the electronic device, and the user behavior habits.

[0082] For the context, the data cleaning includes at least one of the following processes:

[0083] 1) Merging and processing of consecutive numbers. This is because the user's statement may contain consecutive number strings, such as phone numbers, measurement numbers, amounts, etc. This type of data needs to be considered as a whole and cannot be separated.

[0084] 2) Merge and process consecutive English characters. This is because consecutive English characters appearing in Chinese usually represent an English word or a string with special meaning, and such characters need to be treated as a whole.

[0085] 3) Filter out stop words. This is because the user's statements in each round of conversation may contain meaningless stop words such as "de", "de", "di", etc. Although these words are used in many situations, they cannot be used to distinguish the user's intention.

[0086] 4) Slot synonym filtering. This is because in the user's statements, there may be both system-defined slots and related synonyms, as well as user-defined slots and synonyms. Among them, system-defined slots and related synonyms are, for example, country, city, time, etc. User-defined slots and synonyms are, for example, the slot is train ticket, and the slot values are high-speed train ticket, bullet train ticket, hard sleeper, hard seat, etc. These words can be randomly replaced in the user's statements without changing the user's intention. For example, replace all synonyms of train ticket with train ticket.

[0087] For device status information, data cleaning includes: encoding the status of the hardware or software of the electronic device using one-hot code. For example, for software or hardware in the on state, encode it as 1, and for software or hardware in the off state, encode it as 0.

[0088] For user behavior habits, define n common user behavior habits, and initialize each user behavior habit with a random n-dimensional vector. For example, define 100 user behavior habits, and initialize each of the 100 user behavior habits with a 100-dimensional random vector.

[0089] Through data cleaning, invalid data in historical conversation records, device status information, or user behavior habits is filtered out, reducing the data processing volume.

[0090] 203. Multi-source feature encoding.

[0091] In this step, the server performs multi-source feature encoding on the context, device status information, and user behavior habits that have undergone data cleaning.

[0092] Figure 4 It is a schematic diagram of the process of multi-source feature encoding and multi-source feature fusion in the semantic recognition method provided by the embodiment of the present application.

[0093] Please refer to Figure 4 , in the multi-source data encoding stage, the server determines the first encoding vector according to the historical conversation record and the text to be processed, such as Figure 4 the x inn ; Determine the second encoding vector according to the device status information, such as Figure 4 c in Figure 4 ; Determine the third encoding vector according to the user behavior habits of the user line, such as

[0094] q in

[0095] In the embodiments of the present application, the server can obtain the first encoding vector corresponding to the context according to the text to be processed and the historical conversation record. Taking the user's statement in the (n-1)-th round of conversation as the historical conversation record as an example, in the process of determining the first encoding vector, the server uses the glove model to vectorize the user's statement in the (n-1)-th round of conversation to obtain the word vector of the user's statement in the (n-1)-th round of conversation, denoted as w = [w 1 , w 2 ,... w t ∈ R d×t , where d represents the dimension of the word vector, for example, 300, etc., and t represents the number of words after segmenting the user's statement. Then, the server determines the semantic vector of each word in the user's statement in the (n-1)-th round of conversation.

[0096] When the server determines the semantic vector of each word in the user's statement in the (n-1)-th round of conversation, the server uses the Long Short-Term Memory (LSTM) network to determine the past semantics and future semantics of each word in the word vector, and determines the semantic vector of the corresponding word in the user's statement in the (n-1)-th round of conversation according to the past semantics and the future semantics of each word in the user's statement in the (n-1)-th round of conversation. Among them, the past semantics and future semantics of the i-th word in the word vector are shown in formulas (2-1) and (2-2).

[0097]

[0098]

[0099] Among them, i represents the i-th word in the word vector, w represents the random parameter value corresponding to the i-th word in the LSTM, n represents the n-th round of conversation, n-1 represents the (n-1)-th round of conversation, represents the vector output by the bidirectional LSTM for the i-th word in the (n-1)-th round; q takes values from 1 to n. Different directions of the vectors in Formula (2-1) and Formula (2-2) represent LSTM models in different directions.

[0100] For the i-th word in the word vector corresponding to the user's statement in the (n - 1)-th round of conversation, after the server obtains the past semantics and future semantics of the i-th word, the past semantics and future semantics are concatenated to obtain the semantic vector of the i-th word in the word vector corresponding to the user's statement in the (n - 1)-th round of conversation. The semantic vector of the i-th word is shown in Formula (2-3):

[0101]

[0102] In Formula (2-3), ⊕ represents the concatenation of two vectors.

[0103] After the server obtains the semantic vector of the i-th word, based on determining the contribution distribution of the i-th word to the semantics of the user's statement, as shown in Formula (2-4):

[0104]

[0105] In Formula (2-4), represents the contribution distribution of the i-th word to the user's statement in the (n - 1)-th round of conversation, sigmod represents a function, t represents the number of words after segmenting the user's statement, represents the random parameter corresponding to t in the (n - 1)-th round of conversation. Since the value of t is different in each round of conversation, so in each round of conversation the values taken may be different, represents the value of the random parameter corresponding to the i-th word in the (n - 1)-th round in the LSTM, and tanh represents the hyperbolic function.

[0106] The server determines the semantic vector of the user's statement in the (n - 1)-th round of conversation according to the semantic vector of each word in the user's statement and the contribution distribution of each word in the user's statement to the semantics of the user's statement. The semantic vector of the user's statement in the (n - 1)-th round of conversation is shown in Formula (2-5):

[0107]

[0108] The server obtains the semantic vector x of the user's statement in the (n - 1)-th round of conversation n-1 and the user's statement (i.e., the text to be processed) in the n-th round of conversation. After that, according to the text to be processed, the server obtains the word vector corresponding to the user's statement in the n-th round of conversation, and concatenates the word vector corresponding to the user's statement in the n-th round of conversation and the semantic vector x of the user's statement in the (n - 1)-th round of conversation n-1 to obtain the first encoding vector x corresponding to the text to be processed in the n-th round of conversation by using the calculation methods of the above Formulas (2-1) to (2-5). nObviously, the first encoding vector x n incorporates the context.

[0109] Secondly, the server determines a second encoding vector according to the device status information.

[0110] In the embodiments of the present application, the device status information refers to the running status of software and / or hardware on the electronic device, etc. The server obtains the device status information from the electronic device and obtains the second encoding vector according to the device status information. For example, the server represents the running status of 30 commonly used APPs on the electronic device with a 30-dimensional random vector. Again, the server represents the running status of 50 commonly used APPs on the electronic device with a 50-dimensional random vector.

[0111] Finally, the server determines a third encoding vector according to the user's behavior habits.

[0112] In the embodiments of the present application, the user behavior habits are used to indicate the habits of the user using the software and hardware on the electronic device. Taking a software as an APP as an example, the user behavior habits refer to the operation sequence of the user for the APP. After pre-defining the commonly used APPs of the user, the server randomly initializes the elements in the operation sequence, such as opening a certain application, playing the next song, etc. Then, the server uses the same calculation method as the context to calculate the third encoding vector q corresponding to the user behavior habits. The dimension of the third encoding vector is the same as the number of the user behavior habits

[0113] 204. Multi-source feature fusion.

[0114] The server inputs the first encoding vector, the second encoding vector, and the third encoding vector into the autoencoder 1, the autoencoder 2, and the autoencoder 3 respectively. Exemplarily, reference can be made to formula (2-6), formula (2-7), and formula (2-8).

[0115] y = rule(Wx + b) (2-6)

[0116] z = rule(W T y + b′) (2-7)

[0117]

[0118] In the above formula (2-6), formula (2-7), and formula (2-8), rule is a non-linear algorithm in the industry, W, W TLet \(A\) and \(B\) represent two different matrices of randomly initialized parameters, and \(x\) represent the input vector of the autoencoder. \(B\) represents the randomly initialized parameter values, \(y\) represents the corresponding vector in the middle layer, \(b'\) represents the randomly initialized parameter values, \(z\) represents the vector value corresponding to the last layer of the autoencoder, \(l\) represents the cross-entropy formula, and \(d\) represents the dimension.

[0119] When the above formulas (2-6), (2-7), and (2-8) are used for Autoencoder 1, then \(x\) is the first encoded vector \(x\) n , and \(d\) represents the first encoded vector \(x\) n with dimension. \(i = 1, l\) 1 represents the calculation formula for the cross-entropy of Autoencoder 1. When the above formulas (2-6), (2-7), and (2-8) are used for Autoencoder 2, then \(x\) is the second encoded vector \(c\), and \(d\) represents the dimension of the second encoded vector \(c\), for example, 30. \(i = 2, l\) 2 represents the calculation formula for the cross-entropy of Autoencoder 2. When the above formulas (2-6), (2-7), and (2-8) are used for Autoencoder 3, then \(x\) is the third encoded vector \(q\), and \(d\) represents the dimension of the third encoded vector \(q\), for example, 100. \(i = 3, l\) 3 represents the calculation formula for the cross-entropy of Autoencoder 3.

[0120] From the above, it can be known that for any one of Autoencoders 1 to 3, according to the above formula (2-6), the middle layer vector \(y\) of the autoencoder can be obtained. That is, inputting the first encoded vector into Autoencoder 1 to obtain the first middle layer vector, inputting the second encoded vector into Autoencoder 2 to obtain the second middle layer vector, and inputting the third encoded vector into Autoencoder 3 to obtain the third middle layer vector. Then, substituting the middle layer vector into formula (2-7) can determine the vector \(z\) corresponding to the last layer of the autoencoder. Finally, based on the vector \(z\) corresponding to the last layer of the autoencoder calculated according to formulas (2-8) and (2-7) and the input vector of the autoencoder, the cross-entropy between the input and output of the autoencoder can be determined. The cross-entropy of Autoencoder 1 is \(L\) 1 , the cross-entropy of Autoencoder 2 is \(L\) 2 , and the cross-entropy of Autoencoder 3 is \(L\) 3 . These three cross-entropies are the least square error loss functions.

[0121] 205. Result prediction.

[0122] After the server obtains the middle layer vectors \(y\) of the three autoencoders respectively, it concatenates these three middle layer vectors \(y\) to obtain a concatenated vector. Input this concatenated vector into a multi-classifier (as shown by the dashed arrow in Figure 4 ), and use the multi-classifier to calculate the loss function \(L\) of the multi-classifier 0 .

[0123] Further, calculate L 0 After that, the backpropagation algorithm function can be determined using the following formula (9)

[0124]

[0125] In formula (2-9), the values of each parameter are determined according to the actual situation. For example, α 0 = 0.4, α 1 = α 2 = α 3 = 0.2

[0126] Obtain the backpropagation algorithm After that, take this backpropagation algorithm as a constraint condition. Under the constraint of this constraint condition, input the concatenated vector into a multi-classifier for multi-classification

[0127] During the classification process, use the concatenated vector to determine the matching probability of each intention among the multiple intentions corresponding to the text to be processed, determine whether the matching probability of each intention exceeds the threshold corresponding to the intention, and take the intention whose matching probability exceeds the threshold corresponding to the intention as the target intention of the text to be processed. Taking the text to be processed as "Beijing Beijing" and the intentions as "video", "city", and "song" as an example, the multi-classification results are shown in Table 1

[0128] Table 1

[0129]

[0130] Please refer to Table 1. The function of the multi-classifier is, for example, formula (2-4), where m represents the number of intentions. In this example, m = 3. In the nth round of conversation, when the user's statement (i.e., the text to be processed) is "Beijing Beijing", the probability P 1 = 0.8 that the target intention is "video", the probability P2 = 0.55 that the target intention is "location", and the probability P3 = 0.4 that the target intention is "song". Among them, P 1 is greater than the threshold 0.7. The electronic device considers that the target intention of the text to be processed is "video". At this time, the electronic device outputs "1", indicating that "video" is the intention matched by the text to be processed. If P 2 is less than the threshold 0.6, the electronic device considers that the target intention of the text to be processed does not match the intention of "location". Similarly, if P 3 is less than the threshold 0.5, the electronic device considers that the target intention of the text to be processed does not match the intention of "song".

[0131] Figure 5The figure is a schematic structural diagram of a semantic recognition device provided by an embodiment of the present application. Optionally, the semantic recognition device involved in this embodiment is the first network device in the foregoing embodiments, or a chip applied to the first network device. The semantic recognition device is used to execute the functions of the first network device in the foregoing embodiments. Optionally, as Figure 5 shown, the semantic recognition device 100 includes: a transceiver unit 11 and a processing unit 12.

[0132] The transceiver unit 11 is configured to receive a voice signal from an electronic device, where the voice signal is a voice signal emitted by the user in the nth round of conversation between the user and the electronic device, and n≥2 and is an integer.

[0133] The processing unit 12 is configured to convert the voice signal into a text to be processed, and determine a target intention of the text to be processed from a whitelist database based on a historical conversation record between the user and the electronic device, device status information of the electronic device, user behavior habits of the user, and the text to be processed. The device status information is used to indicate the status of software and hardware on the electronic device, and the user behavior habits are used to indicate the habits of the user using the software and hardware on the electronic device. The whitelist database stores the correspondence between at least one candidate text and the intention of each candidate text in the at least one candidate text. Each candidate text in the at least one candidate text corresponds to at least one intention, and the at least one candidate text includes the text to be processed, and execute an instruction corresponding to the target intention.

[0134] In a feasible design, when determining the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, the processing unit 12 is configured to determine a first coding vector according to the historical conversation record and the text to be processed, determine a second coding vector according to the device status information, determine a third coding vector according to the user behavior habits of the user, fuse the first coding vector, the second coding vector, and the third coding vector to obtain a fusion vector, and determine the target intention of the text to be processed from the whitelist database according to the fusion vector.

[0135] In a feasible design, when determining the first encoded vector according to the historical conversation record and the text to be processed, the processing unit 12 is configured to determine, according to the historical conversation record, the semantic vectors of each word in the user's utterance in the (n-1)-th round of conversation, where the user's utterance is the text corresponding to the voice signal sent by the user in the (n-1)-th round of conversation, determine the contribution distribution of each word in the user's utterance to the semantics of the user's utterance, determine the semantic vector of the user's utterance in the (n-1)-th round of conversation according to the semantic vectors of each word in the user's utterance and the contribution distribution of each word in the user's utterance to the semantics of the user's utterance, and determine the first encoded vector according to the semantic vector of the user's utterance in the (n-1)-th round of conversation and the word vector corresponding to the text to be processed.

[0136] In a feasible design, when determining the semantic vectors of each word in the user's utterance in the (n-1)-th round of conversation according to the historical conversation record, the processing unit 12 is configured to determine the word vectors of the user's utterance in the (n-1)-th round of conversation, use a long short-term memory network (LSTM) to determine the past semantics and future semantics of each word in the word vectors, and determine the semantic vectors of the corresponding words in the user's utterance in the (n-1)-th round of conversation according to the past semantics and the future semantics of each word in the user's utterance in the (n-1)-th round of conversation.

[0137] In a feasible design, when determining the second encoded vector according to the device status information, the processing unit 12 is configured to perform one-hot encoding on the software and hardware on the electronic device according to the running status of the software and hardware on the electronic device to obtain the second encoded vector, where the dimension of the second encoded vector is the same as the number of software and hardware on the electronic device, and the running status includes on or off.

[0138] In a feasible design, the dimension of the third encoded vector is the same as the number of the user's behavior habits.

[0139] In a feasible design, when fusing the first encoded vector, the second encoded vector, and the third encoded vector to obtain a fused vector, the processing unit 12 is configured to input the first encoded vector into a first autoencoder to obtain a first intermediate layer vector, input the second encoded vector into a second autoencoder to obtain a second intermediate layer vector, input the third encoded vector into a third autoencoder to obtain a third intermediate layer vector, and fuse the first intermediate layer vector, the second intermediate layer vector, and the third intermediate layer vector to obtain the concatenated vector.

[0140] In a feasible design, when determining the target intent of the text to be processed from the whitelist database according to the fusion vector, the processing unit 12 is used to determine the matching probability of each intent among the multiple intents corresponding to the text to be processed by using the splicing vector, determine whether the matching probability of each intent exceeds the threshold of the corresponding intent, and use the intent whose matching probability exceeds the threshold of the corresponding intent as the target intent of the text to be processed.

[0141] In a feasible design, before the processing unit 12 determines the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, and the user behavior habit of the user, the transceiver unit 11 is further used to receive the indication information sent by the electronic device, where the indication information carries the device status information and / or the user behavior habit.

[0142] In a feasible design, before the processing unit 12 determines the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habit of the user, and the text to be processed, the processing unit 12 is further used to perform data cleaning on at least one of the historical conversation record, the device status information, or the user behavior habit.

[0143] In a feasible design, after executing the instruction corresponding to the target intent, the processing unit 12 is further used to determine the execution result corresponding to the instruction, and the transceiver unit 11 is further used to send the execution result to the electronic device.

[0144] The semantic recognition device provided by the embodiments of the present application can perform the actions of the server in the above embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0145] Figure 6 It is a schematic structural diagram of another semantic recognition device provided by the embodiments of the present application. As Figure 6 shown, the semantic recognition device 200 includes a processor 21 and a memory 22.

[0146] The memory 22 stores computer execution instructions.

[0147] The processor 21 executes the computer execution instructions stored in the memory 22, so that the semantic recognition device executes the semantic recognition method executed by the server as above.

[0148] The specific implementation process of the processor 21 can be referred to in the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0149] Optionally, the semantic recognition device 200 further includes a communication interface 23. Among them, the processor 21, the memory 22, and the communication interface 23 can be connected through a bus 24.

[0150] In the implementation of the above semantic recognition device, the memory and the processor are directly or indirectly electrically connected to achieve data transmission or interaction, that is, the memory and the processor are connected or integrated through an interface. For example, these components are electrically connected to each other through one or more communication buses or signal lines, such as being connected through a bus. The memory stores computer execution instructions for implementing the data access control method, including software function modules stored in the memory in the form of at least one software or firmware. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory.

[0151] The memory includes, but is not limited to, random access memory (Random Access Memory, abbreviated as RAM), read-only memory (Read Only Memory, abbreviated as ROM), programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, abbreviated as EEPROM), etc. Among them, the memory is used to store programs, and the processor executes the programs after receiving the execution instructions. Further, the software programs and modules in the above memory may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide a running environment for other software components.

[0152] The processor is an integrated circuit chip with signal processing capabilities. The above-mentioned processor is, for example, a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor is, for example, a microprocessor or any conventional processor, etc.

[0153] On the above basis, the present application further provides an electronic device, including the semantic recognition device implemented in any of the above implementation manners.

[0154] On the basis described above, the present application further provides a chip, including: a logic circuit and an input interface, where: the input interface is used to obtain data to be processed, such as voice signals, historical conversation records, device status information, user behavior habits, etc.; the logic circuit is used to execute the technical solutions on the server side in the foregoing method embodiments for the data to be processed, and obtain the processed data. The data to be processed includes a target intention, an instruction corresponding to the target intention, etc.

[0155] Optionally, the chip further includes: an output interface, and the output interface is used to output the processed data.

[0156] The present application further provides a computer-readable storage medium, and the computer-readable storage medium is used to store a program, and the program is used to execute the technical solutions on the server side in the foregoing embodiments when being executed by a processor.

[0157] The embodiments of the present application further provide a computer program product, and when the computer program product runs on a semantic recognition device, the semantic recognition device is enabled to execute the technical solutions on the server side in the foregoing embodiments.

[0158] Those of ordinary skill in the art should understand that all or part of the steps for implementing the foregoing method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the foregoing method embodiments; and the foregoing storage medium includes: various media for storing program codes such as ROM, RAM, magnetic disks, or optical discs, and the specific type of the medium is not limited in the present application.

Claims

1. A semantic recognition method, characterized in that, it includes: Receiving a voice signal from an electronic device, where the voice signal is the voice signal emitted by the user in the nth round of conversation between the user and the electronic device, n≥2 and n is an integer; Converting the voice signal into a text to be processed; Based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, determining the target intention of the text to be processed from a whitelist database, where the device status information is used to indicate the status of the software and hardware on the electronic device, the user behavior habits are used to indicate the habits of the user using the software and hardware on the electronic device, the whitelist database stores the corresponding relationship between at least one candidate text and the intention of each candidate text in the at least one candidate text, each candidate text in the at least one candidate text corresponds to at least one intention, and the at least one candidate text includes the text to be processed; Executing the instruction corresponding to the target intention.

2. The method according to claim 1, characterized in that, The determining the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed includes: Determining a first coding vector according to the historical conversation record and the text to be processed; Determining a second coding vector according to the device status information; Determining a third coding vector according to the user behavior habits of the user; Fusing the first coding vector, the second coding vector, and the third coding vector to obtain a fused vector; Determining the target intention of the text to be processed from the whitelist database according to the fused vector.

3. The method according to claim 2, characterized in that, The determining the first coding vector according to the historical conversation record and the text to be processed includes: Determining the semantic vector of each word in the user's statement in the (n - 1)th round of conversation according to the historical conversation record, where the user's statement is the text corresponding to the voice signal emitted by the user in the (n - 1)th round of conversation; Determining the contribution distribution of each word in the user's statement to the semantics of the user's statement; Determining the semantic vector of the user's statement in the (n - 1)th round of conversation according to the semantic vector of each word in the user's statement and the contribution distribution of each word in the user's statement to the semantics of the user's statement; Determining the first coding vector according to the semantic vector of the user's statement in the (n - 1)th round of conversation and the word vector corresponding to the text to be processed.

4. The method according to claim 3, characterized in that, The determining the semantic vector of each word in the user's statement in the (n - 1)th round of conversation according to the historical conversation record includes: Determining the word vector of the user's statement in the (n - 1)th round of conversation; Using a long short-term memory network LSTM to determine the past semantics and future semantics of each word in the word vector. Determine the semantic vector of the corresponding word in the user's statement in the (n - 1)-th round of conversation according to the past semantics and the future semantics of each word in the user's statement in the (n - 1)-th round of conversation.

5. The method according to any one of claims 2 - 4, wherein, the determining the second encoding vector according to the device state information includes: performing one-hot encoding on the software and hardware on the electronic device according to the operating states of the software and hardware on the electronic device to obtain the second encoding vector, the dimension of the second encoding vector being the same as the number of software and hardware on the electronic device, and the operating states including on or off.

6. The method according to any one of claims 2 - 4, wherein, the fusing the first encoding vector, the second encoding vector, and the third encoding vector to obtain a fused vector includes: inputting the first encoding vector into a first autoencoder to obtain a first intermediate layer vector; inputting the second encoding vector into a second autoencoder to obtain a second intermediate layer vector; inputting the third encoding vector into a third autoencoder to obtain a third intermediate layer vector; fusing the first intermediate layer vector, the second intermediate layer vector, and the third intermediate layer vector to obtain a concatenated vector.

7. The method according to claim 6, wherein, the determining the target intention of the text to be processed from the whitelist database according to the fused vector includes: using the concatenated vector to determine the matching probability of each intention among the multiple intentions corresponding to the text to be processed; determining whether the matching probability of each intention exceeds the threshold of the corresponding intention, and taking the intention whose matching probability exceeds the threshold of the corresponding intention as the target intention of the text to be processed.

8. The method according to claim 1, wherein, before determining the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device state information of the electronic device, the user's user behavior habit, and the text to be processed, further includes: receiving the indication information sent by the electronic device, the indication information carrying the device state information and / or the user behavior habit.

9. The method according to claim 1, wherein, before determining the target intention of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device state information of the electronic device, the user's user behavior habit, and the text to be processed, further includes: performing data cleaning on at least one of the historical conversation record, the device state information, or the user behavior habit.

10. A semantic recognition device, wherein, comprises: a transceiver unit, configured to receive a voice signal from an electronic device, the voice signal being the voice signal sent by the user in the n-th round of conversation between the user and the electronic device, n≥2 and being an integer; A processing unit, configured to convert the voice signal into a text to be processed, and determine a target intent of the text to be processed from a whitelist database based on a historical conversation record between the user and the electronic device, device status information of the electronic device, user behavior habits of the user, and the text to be processed. The device status information is used to indicate the status of software and hardware on the electronic device, and the user behavior habits are used to indicate the habits of the user using the software and hardware on the electronic device. The whitelist database stores the correspondence between at least one candidate text and the intent of each candidate text in the at least one candidate text. Each candidate text in the at least one candidate text corresponds to at least one intent, and the at least one candidate text includes the text to be processed, and execute an instruction corresponding to the target intent.

11. The apparatus according to claim 10, wherein, when determining the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, the processing unit is configured to determine a first coding vector according to the historical conversation record and the text to be processed, determine a second coding vector according to the device status information, determine a third coding vector according to the user behavior habits of the user, fuse the first coding vector, the second coding vector, and the third coding vector to obtain a fused vector, and determine the target intent of the text to be processed from the whitelist database according to the fused vector.

12. The apparatus according to claim 11, wherein, when determining the first coding vector according to the historical conversation record and the text to be processed, the processing unit is configured to determine a semantic vector of each word in the user's statement in the (n-1)-th round of conversation according to the historical conversation record, where the user's statement is the text corresponding to the voice signal sent by the user in the (n-1)-th round of conversation, determine the contribution distribution of each word in the user's statement to the semantics of the user's statement, determine the semantic vector of the user's statement in the (n-1)-th round of conversation according to the semantic vector of each word in the user's statement and the contribution distribution of each word in the user's statement to the semantics of the user's statement, and determine the first coding vector according to the semantic vector of the user's statement in the (n-1)-th round of conversation and the word vector corresponding to the text to be processed.

13. The apparatus according to claim 12, wherein, When determining the semantic vector of each word in the user's statement in the (n - 1)-th round of conversation according to the historical conversation record, the processing unit is used to determine the word vector of the user's statement in the (n - 1)-th round of conversation, use a long short-term memory network (LSTM) to determine the past semantics and future semantics of each word in the word vector, and determine the semantic vector of the corresponding word in the user's statement in the (n - 1)-th round of conversation according to the past semantics and the future semantics of each word in the user's statement in the (n - 1)-th round of conversation.

14. The apparatus according to any one of claims 11 - 13, wherein, when determining the second encoding vector according to the device status information, the processing unit is used to perform one-hot encoding on the software and hardware on the electronic device according to the running status of the software and hardware on the electronic device to obtain the second encoding vector, the dimension of the second encoding vector is the same as the number of software and hardware on the electronic device, and the running status includes on or off.

15. The apparatus according to any one of claims 11 - 13, wherein, when fusing the first encoding vector, the second encoding vector and the third encoding vector to obtain a fused vector, the processing unit is used to input the first encoding vector into a first autoencoder to obtain a first intermediate layer vector, input the second encoding vector into a second autoencoder to obtain a second intermediate layer vector, input the third encoding vector into a third autoencoder to obtain a third intermediate layer vector, and fuse the first intermediate layer vector, the second intermediate layer vector and the third intermediate layer vector to obtain a concatenated vector.

16. The apparatus according to claim 15, wherein, when determining the target intention of the text to be processed from the white list database according to the fused vector, the processing unit is used to use the concatenated vector to determine the matching probability of each intention in the multiple intentions corresponding to the text to be processed, determine whether the matching probability of each intention exceeds the threshold of the corresponding intention, and use the intention whose matching probability exceeds the threshold of the corresponding intention as the target intention of the text to be processed.

17. The apparatus according to claim 10, wherein, before the processing unit determines the target intention of the text to be processed from the white list database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user's user behavior habits, and the text to be processed, the transceiver unit is further used to receive indication information sent by the electronic device, and the indication information carries the device status information and / or the user behavior habits.

18. The apparatus according to claim 10, wherein, Before determining the target intent of the text to be processed from the whitelist database based on the historical conversation record between the user and the electronic device, the device status information of the electronic device, the user behavior habits of the user, and the text to be processed, the processing unit is further configured to perform data cleaning on at least one of the historical conversation record, the device status information, or the user behavior habits.

19. A semantic recognition device, characterized in that it includes: a processor and a memory, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that it is used to store a computer program or instructions, and when the computer program or instructions run on the semantic recognition device, the semantic recognition device is caused to execute the method according to any one of claims 1 to 9.

21. A chip, characterized in that the chip includes a programmable logic circuit and an input interface, the input interface is used to obtain data to be processed, and the logic circuit is used to execute the method according to any one of claims 1 to 9 on the data to be processed.

22. A computer program product, characterized in that the computer program product includes computer program code, and when the computer program code runs on a computer, the computer is caused to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Operation execution method and device and computer readable storage medium

    CN109885652A