Multi-intention semantic understanding model training and using method and related products

By transforming the attention matrix and encoding mask of the multi-intention semantic understanding model, and adding special characters at the end of the entire sentence data, training the multi-intention model solves the problems of high cost and low precision, and achieving efficient streaming semantic analysis and instant response.

CN120356468APending Publication Date: 2025-07-22AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510559900.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the multi-intention task that is performed while speaking, the prior art, the training cost of semantic understanding models is high and the deployment is difficult, and the analytical accuracy decreases when the user speaks quickly.

Method used

By transforming the attention matrix and encoding mask of the multi-intention semantic understanding model, and adding k special characters with posterior window value at the end of the entire sentence data, the multi-intention understanding model is trained and the semantic analysis of streaming input is supported.

Benefits of technology

It reduces the cost of data generation and model training, and improves the analytical accuracy and instant response capabilities of the model under streaming input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356468A_ABST
    Figure CN120356468A_ABST
Patent Text Reader

Abstract

The invention discloses a training and using method of a multi-intention semantic understanding model, electronic equipment and a storage medium, and the training method of the multi-intention semantic understanding model is characterized in that the multi-intention semantic understanding model comprises a clause model and a plurality of distributed cascaded slot position models. Transforming the clause model into a streaming clause model. The training method comprises the following steps: acquiring whole sentence data for training; k special characters are added to the tail of each whole sentence of the whole sentence data, training data are obtained, and k is a posterior window value; and training the multi-intention understanding model based on the training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of speech recognition technology, and in particular to a method for training and using a multi-intent semantic understanding model and related products. Background Art

[0002] In the related art, the specific task of performing while speaking is defined as follows: when a user expects to input a task instruction with multiple intents, before the complete instruction is finished speaking, for the instruction of a complete single intent that has been clearly expressed, an immediate response and interaction can be performed to achieve an extremely fast understanding effect. For example, assuming that the user's complete instruction is "Open the window, call mom, then play the song of XXX, and then navigate to Window of the World", when the user says "Open the window", "Call mom", "Play the song of XXX", and "Navigate to Window of the World", immediate semantic understanding and dialogue interaction can be performed. The effect experienced by the user is that when saying "Call", the instruction of "Open the window" before has been executed.

[0003] The inventors found that in the related art, generally, semantic determination and semantic parsing are performed on the user's streaming input. Instead of analyzing from the perspective of natural language processing, the required training cost is extremely high and the deployment difficulty is great. Or use a natural language processing model for sentence segmentation, but at the same time introduce time dimension features. When the user's speech rate is relatively fast, the time features will interfere with the progress of streaming segmentation, resulting in a decrease in the model parsing accuracy. Summary of the Invention

[0004] The embodiments of the present invention provide a method for training and using a multi-intent semantic understanding model and related products, which are used to solve at least one of the above technical problems.

[0005] In a first aspect, the embodiments of the present invention provide a method for training a multi-intent semantic understanding model, including: obtaining whole-sentence data for training; adding k special characters at the end of each whole sentence of the whole-sentence data to obtain training data, where k is a posterior window value; training the multi-intent understanding model based on the training data.

[0006] In a second aspect, the embodiments of the present invention provide a method for using a multi-intent semantic understanding model, including: sending the user's current streaming input into the streaming clause splitting model to split the complete intent clauses that can be determined; sending the complete intent clauses into the slot model for semantic parsing to obtain at least one semantic result; performing multi-intent fusion on the at least one semantic result to obtain a final multi-intent semantic result and execute it.

[0007] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training and usage methods of any one of the above-mentioned multi-intent semantic understanding models of the present invention.

[0008] In a fourth aspect, an embodiment of the present invention provides a storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) for executing the training and usage methods of any one of the above-mentioned multi-intent semantic understanding models of the present invention.

[0009] In a fifth aspect, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the training and usage methods of any one of the above-mentioned multi-intent semantic understanding models.

[0010] The method of the present application obtains training data and trains a multi-intent semantic understanding model by modifying the attention matrix and encoding mask of the model, and adding k special characters at the end of each whole sentence of the whole sentence data, thereby reducing the data generation cost and the computing power cost of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a flowchart of a training method for a multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 2 It is a flowchart of another usage method for a multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of a single-intent semantic understanding model of a specific example of the training and usage methods of a multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of a clause sequence annotation model of a specific example of the training and usage methods of a multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 5Schematic diagram of a conventional multi-intent model in a distributed cascade, which is a specific example of the training and usage method of the multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 6 Schematic diagram of the decoding effect of the streaming mode at k = 3, which is a specific example of the training and usage method of the multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 7 Engineering flowchart of speaking and executing simultaneously, which is a specific example of the training and usage method of the multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 8 Flowchart of the training and usage of the multi-intent semantic understanding model, which is a specific example of the training and usage method of the multi-intent semantic understanding model provided by an embodiment of the present invention; Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0014] Please refer to Figure 1 , which shows a flowchart of a training method of a multi-intent semantic understanding model provided by an embodiment of the present invention. Among them, the multi-intent understanding model includes a sentence splitting model and a plurality of distributed cascade slot models, and the sentence splitting model is transformed into a streaming sentence splitting model.

[0015] As Figure 1 shown, in step 101, the whole sentence data for training is obtained; In step 102, k special characters are added to the end of each whole sentence of the whole sentence data to obtain training data, where k is the posterior window value; In step 103, the multi-intent understanding model is trained based on the training data.

[0016] In this embodiment, for step 101, the training device of the multi-intent semantic understanding model obtains the multi-intent whole sentence data for training. In a specific embodiment, in order to support the user input of single intent at the same time, in the overall training data, a part of the single-intent whole sentence data is also added, and some rejection data and noise data are additionally introduced to improve the generalization ability of the model.

[0017] Then, for step 102, the training device of the multi-intent semantic understanding model adds k special characters at the end of each whole sentence of the whole-sentence data to obtain training data, where k is the posterior window value. For example, due to the innate nature of streaming multi-intent, the label calculation of tokens at fixed positions only performs parameter calculation through the information of the previous tokens (lemmas / tokens) and the subsequent k tokens. Therefore, it is not necessary to traverse and generate all streaming input data. Taking "open the window", "open the window and", "open the window and close", "open the window and turn off", "open the window and turn off the electricity", "open the window and turn off the phone" as examples, only one complete multi-intent sentence, such as "open the window and turn off the phone", is needed to enable the model to learn the results output by each state of the whole sentence under streaming input. At the same time, in order to enable the model to support the parsing of complete sentences, that is, to determine the labels of the last k input tokens, k special characters are added at the end of each piece of training data as the whole sentence for training. For example, the special character is "&", k = 3, and the whole sentence is "open the window and turn off the phone", then the training data is "open the window and turn off the phone&&&".

[0018] Finally, for step 103, the training device of the multi-intent semantic understanding model trains the multi-intent understanding model based on the training data.

[0019] The method of this embodiment obtains training data and trains the multi-intent semantic understanding model by adding k special characters at the end of each whole sentence of the whole-sentence data, thereby reducing the data generation cost and the computing power cost of model training.

[0020] In some optional embodiments, the clause model includes an encoding layer, a bidirectional attention matrix, and an encoding mask. Transforming the clause model into a streaming clause model includes: When transforming the bidirectional attention matrix, the bidirectional attention matrix is transformed into an attention matrix that only focuses on the forward sequence, and then the attention matrix is shifted to the right so that when the encoding layer of the clause model calculates the attention at the i-th position, it only focuses on the subsequent k tokens; When transforming the encoding mask, the word vectors at the first n - k positions are inserted into the word vectors of each layer, and the token information at the subsequent k positions is not passed to the next layer through the mask, and is remapped back to the encoding layer of the clause model through a linear mapping.

[0021] In some alternative embodiments, the whole-sentence data includes single-intention whole sentences, multi-intention whole sentences, rejection recognition data, and / or noise data. For example, using multi-intention whole sentences can enable the model to support multi-intention user inputs. Further, in order to support single-intention user inputs simultaneously, a part of single-intention whole-sentence data is added to the overall training data, and some rejection recognition data and noise data are also introduced additionally to improve the generalization ability of the model.

[0022] In some alternative embodiments, k is a posterior window value, and k supports customization. When k is set to a lower value, it is used for real-time response and execution while speaking. When k is set to a higher value, it is used to reduce the mis-triggering of single-intention sentences for execution while speaking. For example, users who tend to have a real-time response and execution while speaking can set a smaller k value. For customers who do not want single-intention sentences to be mis-triggered for execution while speaking, a higher window k value can be set. Thus, the model can have additional diversified customization capabilities.

[0023] Please refer to Figure 2 , which shows a flowchart of a method for using a multi-intention semantic understanding model provided by an embodiment of the present invention. As Figure 2 shown, in step 201, the current streaming input of the user is sent into the streaming clause-splitting model to split the complete intention clauses that can be determined. In step 202, the complete intention clauses are sent into the slot model for semantic parsing to obtain at least one semantic result. In step 203, multi-intention fusion is performed on the at least one semantic result to obtain the final multi-intention semantic result and execute it.

[0024] In this embodiment, for step 201, the device for using the multi-intention semantic understanding model sends the current streaming input of the user into the streaming clause-splitting model to split the complete intention clauses that can be determined. Taking "Open the window and then navigate to XXX" as an example, it can be split into clause 1 "Open the window" and clause 2 "Navigate to XXX".

[0025] Then, for step 202, the device for using the multi-intention semantic understanding model sends the complete intention clauses into the slot model for semantic parsing to obtain at least one semantic result. Taking "Open the window and then navigate to XXX" as an example, first, clause 1 "Open the window" is sent into the slot model for semantic parsing to obtain the task-based semantic result of the first intention, and then clause 2 "Navigate to XXX" is sent into the slot model for parsing to obtain the task-based semantic result of the second intention.

[0026] Finally, for step 203, the device for using the multi-intention semantic understanding model performs multi-intention fusion on the at least one semantic result to obtain the final multi-intention semantic result and execute it.

[0027] The method of this embodiment sends the user's current streaming input into a streaming sentence model to split the complete intent clauses that can be determined, and then sends the complete intent clauses into a slot model for semantic analysis to obtain at least one semantic result, and finally performs multi-intent fusion on the at least one semantic result to obtain the final multi-intent semantic result and execute it, thereby completing the task of speaking and executing.

[0028] In some optional embodiments, the step of feeding the user's current streaming input into the streaming sentence model and splitting the complete intent clauses that can be determined includes: The device using the multi-intention semantic understanding model performs a rule match based on the obtained streaming request to determine whether the match is successful; if the match is successful, delete the last clause of the streaming request; then send the other clauses of the streaming request to multiple slot models for semantic analysis to obtain the final multi-intention task-type instructions. For example, in rule matching, if a half-sentence matches the rule that originally matched the whole sentence, it means that the last clause is a half-sentence with low confidence, where a half-sentence refers to a sentence that is truncated or incomplete, and only a part of the rule is matched, because the rule originally requires a complete context to determine the intent, but the lack of the second half may lead to misjudgment. Taking "Navigate to XX Park" as an example, the complete sentence "Navigate to XX Park" has a high rule matching confidence. If it is a half-sentence "Navigate to...", "XX Park" is missing, and the system cannot confirm the intent, and the confidence is low. Many rules require the second half to determine the semantics. In streaming input, the sentence may be truncated midway, resulting in an incomplete last clause. When the sentence only partially matches the rule, the system will consider the reliability of the matching result to be low, because the missing information may lead to misjudgment, so the last clause needs to be deleted. This can solve the problem of incomplete input processing.

[0029] In some optional embodiments, the streaming request is obtained based on a rule match to determine whether the match is successful, including: if the match fails, the query of the streaming request is sent to the streaming sentence model for sentence segmentation to determine whether the sentence segmentation is successful; if the sentence segmentation fails, no response is given; if the sentence segmentation is successful, the segmented sentences are sent to multiple slot models respectively to obtain the final multi-intention task-type instructions.

[0030] Please refer to Figure 3 , which shows a schematic diagram of a single-intent semantic understanding model as a specific example of a method for training and using a multi-intent semantic understanding model provided by an embodiment of the present invention.

[0031] like Figure 3 As shown in the figure, for a single-intent semantic understanding model, a sequence labeling model is usually used to perform token-level part-of-speech tagging.

[0032] Please refer to Figure 4 , which shows a schematic diagram of a clause sequence annotation model for a specific example of the training and usage method of a multi-intent semantic understanding model provided by an embodiment of the present invention.

[0033] As Figure 4 shown, for multi-intent tasks, a clause model needs to be added before the single-intent semantic understanding model. Therefore, traditional semantic understanding models usually use a distributed cascade method for multi-intent semantic understanding. For example, for a complete clause of a single intent, the model classifies the tokens with "B, I, E, O" labels. If the labels can segment a chunk composed of "BII..E", the clause can be segmented, and then the clause is sent to the single-intent semantic understanding model for slot annotation, and finally a multi-intent result is obtained.

[0034] Please refer to Figure 5 , which shows a schematic diagram of a conventional multi-intent model with distributed cascade for a specific example of the training and usage method of a multi-intent semantic understanding model provided by an embodiment of the present invention.

[0035] As Figure 5 shown, in the distributed cascade model under the conventional multi-intent architecture, a sentence will be divided into (turn on the air conditioner) and (navigate to XXX), and then they are respectively sent to the slot model for semantic parsing to obtain the task-based semantic results of the first intent and the task-based semantic results of the second intent.

[0036] Please refer to Figure 6 , which shows a schematic diagram of the decoding effect of the streaming mode at k = 3 for a specific example of the training and usage method of a multi-intent semantic understanding model provided by an embodiment of the present invention.

[0037] As Figure 6 shown, for the task of performing while speaking, the clause model needs to be structurally modified and adjusted. There are two ways to make this adjustment, namely, modifying the attention matrix (attention mask) and modifying the encoding mask (embedding). First, in order to make the model interpretable, a posterior window value k needs to be set so that the model can calculate the posterior probability of the annotation category for the previous n - k tokens based on the last k tokens (assuming that n tokens are currently input).

[0038] When modifying the attention matrix, the original bidirectional attention mask needs to be changed to a mask that only focuses on the forward sequence, and then it is shifted to the right. In this way, when the encoding layer of the model calculates the attention at the i-th position, it only focuses on the last k tokens. ​​

[0039] When modifying the encoding mask, the word vectors at the first n-k positions are inserted into the word vectors of each layer, and the token information at the last k positions is not passed to the next layer through masking, and is remapped back to 1 through linear mapping and enters the encoder layer of the model.

[0040] Due to the innate nature of streaming multi-intentions, the label calculation of tokens at fixed positions only performs parameter calculation through the information of the previous tokens and the last k tokens. Therefore, it is not necessary to traverse and generate all streaming input data, such as "open the window", "open the window and", "open the window and close", "open the window and turn off", "open the window and turn off the electricity", "open the window and turn off the phone". Only one complete multi-intention sentence needs to be prepared, such as "open the window and turn off the phone", so that the model can learn the results output by the whole sentence in each state under streaming input. At the same time, in order to enable the model to support the parsing of complete sentences, that is, to determine the labels of the last k input tokens, k special characters will be added at the end of each training data as the whole sentence for training. Assuming that the special character is "&" and k = 3, and the whole sentence is "open the window and turn off the phone", the training data is "open the window and turn off the phone&&&".

[0041] Meanwhile, in order to support single-intention user inputs at the same time, in the overall training data, some single-intention complete sentences can also be added, and some rejection recognition data and noise data can be additionally introduced to improve the generalization ability of the model.

[0042] The posterior window value K is extremely important in the design of the model. It gives the model the ability to have a fixed number of posterior tokens, and based on the design of the posterior window, the training data only needs a single sentence to complete the training of the entire streaming ability of the sentence. In addition, due to the adjustability of the posterior window value, it also provides additional diversified customization capabilities. Users who are more inclined to immediate response and execute while speaking can set a smaller k value; while for customers who do not want single-intention sentences to be accidentally triggered and executed while speaking, a higher window k value can be set.

[0043] Please refer to Figure 7 , which shows a specific example of the engineering flowchart of speaking and executing for the training and use method of the multi-intention semantic understanding model provided by an embodiment of the present invention.

[0044] As Figure 7As shown, during the process of engineering while speaking and executing, the user's current query (streaming input) will first be sent to a streaming clause splitting model, and then the complete intent clauses that can be determined will be split. Then, the clauses will be sent to a slot model for semantic parsing, and finally, multi-intent fusion (merge) will be performed to obtain the final (Final) multi-intent semantic result and execute it.

[0045] Please refer to Figure 8 , which shows a flowchart of the training and use of a multi-intent semantic understanding model provided by an embodiment of the present invention, which is a specific example of the training and use method of the multi-intent semantic understanding model.

[0046] As Figure 8 shown, in the training process, it is necessary to train the streaming clause splitting model and the slot model. For the streaming clause splitting model, it is necessary to collect single-intent whole sentences, multi-intent whole sentences, and rejection and noise data. For the collected whole sentence data, k (the set posterior window value) special characters need to be added, and then the streaming clause splitting model is trained. For the slot model, it is only necessary to collect single-intent slot data and train it.

[0047] In the inference process, the trained streaming clause splitting model and slot model will be directly applied to each module of the inference. First, the streaming request will perform a rule matching. If the matching is successful, the last clause will be deleted, and the other clauses will be sent to the slot model for semantic parsing and obtain the final multi-intent task-based instruction. If the rule matching fails, the streaming input query will be sent to the streaming clause splitting model for clause splitting. If the clause splitting fails, there will be no response; if the clause splitting is successful, the split clauses will be sent to the slot model and the final multi-intent task-based instruction will be obtained.

[0048] In some other embodiments, the embodiments of the present invention also provide a non-volatile computer storage medium, and the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the training and use method of the multi-intent semantic understanding model in any of the above method embodiments; As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as: Obtain the whole sentence data for training; Add k special characters at the end of each whole sentence in the whole sentence data to obtain training data, where k is the posterior window value; Train the multi-intent understanding model based on the training data.

[0049] As another implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as: Send the user's current streaming input into the streaming clause splitting model to split the complete intent clauses that can be determined; Send the complete intent clauses into the slot model for semantic parsing to obtain at least one semantic result; Perform multi-intent fusion on the at least one semantic result to obtain the final multi-intent semantic result and execute it.

[0050] The non-volatile computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area can store an operating system and application programs required for at least one function; the storage data area can store data created according to the training and use of the multi-intent semantic understanding model and the like. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the training and use device of the multi-intent semantic understanding model through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0051] The embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the training and use method of the multi-intent semantic understanding model described in any one of the above.

[0052] Figure 9 is a schematic structural diagram of the electronic device provided by the embodiment of the present invention, as Figure 9 shown, the device includes: one or more processors 910 and a memory 920, Figure 9 Taking one processor 910 as an example. The device for training and using the multi-intent semantic understanding model may further include: an input device 930 and an output device 940. The processor 910, the memory 920, the input device 930, and the output device 940 can be connected through a bus or other means, Figure 9 Taking the connection through the bus as an example. The memory 920 is the above-mentioned non-volatile computer-readable storage medium. The processor 910 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 920, that is, implements the training and use method of the multi-intent semantic understanding model in the above method embodiment. The input device 930 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the training and use device of the multi-intent semantic understanding model. The output device 540 may include a display device such as a display screen.

[0053] The above products can execute the method provided by the embodiments of the present invention, and have the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present invention.

[0054] As an implementation manner, the above electronic device is applied to a training and use device of a multi-intent semantic understanding model, and includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtain the whole-sentence data for training; Add k special characters at the end of each whole sentence of the whole-sentence data to obtain training data, where k is the posterior window value; Train the multi-intent understanding model based on the training data.

[0055] As another implementation manner, the above electronic device is applied to a training and use device of a multi-intent semantic understanding model, and includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Send the current streaming input of the user to the streaming clause-splitting model to split the complete intent clauses that can be determined; Send the complete intent clauses to the slot model for semantic parsing to obtain at least one semantic result; Perform multi-intent fusion on the at least one semantic result to obtain the final multi-intent semantic result and execute it.

[0056] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.

[0057] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc., such as iPad.

[0058] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players (such as iPod), handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0059] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but due to the need to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.

[0060] (5) Other electronic devices with data interaction functions.

[0061] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0062] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

Claims

1. A training method for a multi-intent semantic understanding model, wherein, The multi-intent understanding model includes a sentence splitting model and multiple distributed cascaded slot models. It is characterized in that the sentence splitting model is transformed into a streaming sentence splitting model. The training method includes: Obtain the whole sentence data for training; Add k special characters at the end of each whole sentence in the whole sentence data to obtain the training data, where k is the posterior window value; Train the multi-intent understanding model based on the training data.

2. The method according to claim 1, wherein The sentence splitting model includes an encoding layer, a bidirectional attention matrix, and an encoding mask. Transforming the sentence splitting model into a streaming sentence splitting model includes: When transforming the bidirectional attention matrix, transform the bidirectional attention matrix into an attention matrix that only focuses on the forward sequence, and then shift the attention matrix to the right, so that when the encoding layer of the sentence splitting model calculates the attention at the i-th position, it only focuses on the last k tokens; When transforming the encoding mask, insert the word vectors at the first n-k positions into the word vectors of each layer, and do not transmit the token information at the last k positions to the next layer through masking, and remap it back to the encoding layer of the sentence splitting model through linear mapping.

3. The method according to claim 1, wherein, The whole sentence data includes single-intent whole sentences, multi-intent whole sentences, rejection recognition data, and / or noise data.

4. The method according to claim 1, wherein, The k is the posterior window value, and the k supports customization. When the k is set to a lower value, it is used for instant response and execution while speaking; when the k is set to a higher value, it is used to reduce the mis-triggering of single-intent sentences for execution while speaking.

5. A method for using a multi-intent semantic understanding model, which is used for the multi-intent semantic understanding model according to any one of claims 1-4, including: Send the current streaming input of the user into the streaming sentence splitting model to split the complete intent clauses that can be determined; Send the complete intent clauses into the slot model for semantic parsing to obtain at least one semantic result; Perform multi-intent fusion on the at least one semantic result to obtain the final multi-intent semantic result and execute it.

6. The method according to claim 5, wherein The step of sending the current streaming input of the user into the streaming sentence splitting model to split the complete intent clauses that can be determined includes: Perform a rule matching on the obtained streaming request once to determine whether the matching is successful; If the matching is successful, delete the last clause of the streaming request; Send the other clauses of the streaming request into multiple slot models for semantic parsing respectively to obtain the final multi-intent task-based instruction.

7. The method according to claim 6, wherein, The step of performing a rule matching on the obtained streaming request once to determine whether the matching is successful includes: If the matching fails, send the query of the streaming request into the streaming sentence splitting model for sentence splitting to determine whether the sentence splitting is successful; If the sentence splitting is successful, send the segmented sentences into multiple slot models respectively to obtain the final multi-intent task-based instruction.

8. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.