Semantic understanding method and device

By using a combination of encoder and decoder in the semantic understanding method, semantic feature data is extracted and intention and entity recognition is performed, the problem of low accuracy of semantic understanding in the prior art is solved, and higher accuracy and applicability of semantic understanding are achieved.

CN115062620BActive Publication Date: 2025-07-18HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110221290.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-27
Publication Date
2025-07-18
Estimated Expiration
2041-02-27

AI Technical Summary

Technical Problem

The semantic understanding of semantic understanding methods in the prior art is not very accurate, and it is difficult to effectively identify intentions and entities in user speech.

Method used

By obtaining the entered speech converted into text data, the encoder is used to extract semantic feature data, and combined with the semantic intent decoder and entity decoder, the encoder is trained using a decoupled pre-training method to reduce the conflict between intent information and entity information.

Benefits of technology

It improves the accuracy and applicability of semantic understanding, can more accurately identify intentions and related entities in user voice, and achieve accurate control of vehicle functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062620B_ABST
    Figure CN115062620B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a semantic understanding method and apparatus in the field of artificial intelligence. The method includes: obtaining text data converted from input speech. Obtaining first semantic feature data corresponding to the text data through an encoder, and determining the semantic intent corresponding to the first semantic feature data through a semantic intent decoder to implement intent recognition of the text data. Generating fusion data based on the text data and the semantic intent, and obtaining second semantic feature data corresponding to the fusion data based on the encoder. Determining the entity corresponding to the second semantic feature data based on a semantic entity decoder to implement entity recognition associated with the semantic intent in the text data. By using the method provided in the present application, the accuracy of semantic understanding can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a semantic understanding method and apparatus. Background Art

[0002] With the development of artificial intelligence and vehicle networking technologies, a safer and more intelligent travel era is coming. Currently, the traditional touch-screen interaction method (i.e., the user clicks on each function button displayed on the display screen of the vehicle terminal to control functions such as navigation, air conditioning, and music in the vehicle) no longer meets the user's needs. How to maximize the user's interaction experience in the vehicle cockpit has become one of the hottest topics of concern. Among them, natural language understanding (NLU), as a key technology in the field of artificial intelligence, is crucial in the voice interaction experience of future intelligent vehicles.

[0003] When performing NLU processing, first, it is necessary to obtain the word vector corresponding to each character in the multiple characters that make up the sentence to be understood. Then, the input matrix composed of the word vectors corresponding to each character is respectively input into the intent classification model and the entity recognition model for calculation to respectively obtain the intent classification result and the entity recognition result of the sentence to be understood. However, the semantic understanding accuracy of the semantic understanding method in the related technology is not high. Summary of the Invention

[0004] This application provides a semantic understanding method and apparatus, which can improve the accuracy of semantic understanding and has strong applicability.

[0005] It should be understood that the method provided in the embodiments of this application can be executed by a semantic understanding apparatus. The semantic understanding apparatus can be a terminal device, or a part of the components in the terminal device, such as a chip applied in the terminal device, etc., or a server, such as a local server or a cloud server, etc., or a part of the components in the server, such as a chip applied in the server, etc. There is no limitation here.

[0006] In a first aspect, the present application provides a semantic understanding method, which includes: obtaining text data converted from input speech. Among them, the input speech can be collected based on a microphone, and the input speech is converted into corresponding text data through a speech recognition system. Then, the encoder is used to obtain the first semantic feature data corresponding to the above text data, and the semantic intention decoder is used to determine the semantic intention corresponding to the first semantic feature data, so as to realize the intention recognition of the above text data. Fusion data is generated based on the above text data and the above semantic intention, and the second semantic feature data corresponding to the fusion data is obtained based on the above encoder. Based on the semantic entity decoder, the entity corresponding to the second semantic feature data is determined, so as to realize the entity recognition associated with the semantic intention in the above text data.

[0007] In the present application, first, the text data is input into the encoder to obtain the first semantic feature data. Then, after the first semantic feature data passes through the semantic intention decoder, a semantic intention output can be obtained. Further, by fusing the semantic intention and the original text data, the fusion result can be input into the above encoder again to obtain the second semantic feature data. Finally, by passing the second semantic feature data through the semantic entity decoder, the output entity can be obtained. It is not difficult to understand that in the present application, the encoder and the semantic intention decoder obtain the semantic intention corresponding to the text data, and use the semantic intention as prior information to fuse with the text data. Then, the fused data obtained by fusion is encoded by the same encoder, and then the entity is determined based on the semantic entity decoder, which can reduce the conflict between the semantic intention information and the entity information and is beneficial to improving the semantic understanding accuracy in voice interaction.

[0008] Combined with the first aspect, in a first possible implementation manner, before obtaining the first semantic feature data corresponding to the above text data through the encoder, the above method further includes: determining the first word vector matrix corresponding to the above text data. The above process of encoding the above text data through the encoder to obtain the first semantic feature data corresponding to the above text data includes: inputting the first word vector matrix into the encoder, and obtaining the semantic feature vector output by the encoder as the first semantic feature data corresponding to the above text data.

[0009] In the present application, by determining the first word vector matrix corresponding to the text data and inputting the first word vector matrix into the encoder for encoding, and using the semantic feature vector output by the encoder as the first semantic feature data corresponding to the text data, it is easy to operate and has high applicability.

[0010] Combined with the first possible implementation manner of the first aspect, in the second possible implementation manner, the determination of the first word vector matrix corresponding to the above text data includes: performing a character splitting process on the above text data to obtain a plurality of characters included in the above text data. Obtaining a character word vector, a position word vector, and a character type word vector corresponding to each character in the above plurality of characters, where the above position word vector is used to represent the position of the character in the above text data. Summing the character word vector, the position word vector, and the character type word vector corresponding to each character to obtain the word vector corresponding to each character. Generating the first word vector matrix corresponding to the above text data according to the plurality of word vectors corresponding to the above plurality of characters.

[0011] In this application, by determining the vector sum of the character word vector, the position word vector, and the character type word vector corresponding to each character as the word vector corresponding to each character, and then generating the first word vector matrix for processing, it fully considers the position information of each character in the text, that is, the order information of the text data, and the character type information of each character, making the encoding effect better and conducive to improving the accuracy of semantic understanding.

[0012] Combined with the second possible implementation manner of the first aspect, in the third possible implementation manner, the obtaining of the character word vector, the position word vector, and the character type word vector corresponding to each character in the above plurality of characters includes: obtaining a character word vector query table, a position word vector query table, and a character type word vector query table. Wherein, the above character word vector query table includes n character word vectors corresponding to n characters. The above position word vector query table includes m position word vectors corresponding to m positions. The above character type word vector query table includes character type word vectors corresponding to non-padding characters and character type word vectors corresponding to padding characters. Both n and m are integers greater than 0. Obtaining the character word vector corresponding to each character in the above plurality of characters from the above character word vector query table, obtaining the position word vector corresponding to the position of each character in the above text data from the above position word vector query table, and obtaining the character type word vector corresponding to the above non-padding character as the character type word vector corresponding to each character from the above character type word vector query table.

[0013] In this application, by querying the pre-set character word vector query table, position word vector query table, and character type word vector query table to obtain the character word vector, position word vector, and character type word vector corresponding to each character, the operation is simple, the data processing efficiency is improved, and the applicability is strong.

[0014] Combined with any one of the first possible implementation manner to the third possible implementation manner of the first aspect, in the fourth possible implementation manner, the encoder includes j encoding units. The input of the first encoding unit in the j encoding units is the first word vector input matrix. The input of any encoding unit after the first encoding unit is the output of the upper encoding unit of the any encoding unit. The output of the jth encoding unit in the j encoding units is the first semantic feature data, where j is an integer greater than 0.

[0015] Combined with any one of the first aspect to the fourth possible implementation manner of the first aspect, in the fifth possible implementation manner, generating the fusion data according to the above text data and the above semantic intention includes: splicing the above text data and the above semantic intention to obtain the fusion data corresponding to the above text data and the above semantic intention.

[0016] In this application, by splicing the determined semantic intention as prior information with the text data to obtain the fusion data, and then processing the fusion data, it is beneficial to improve the recognition accuracy of entities associated with the semantic intention and improve the accuracy of semantic understanding.

[0017] Combined with the fifth possible implementation manner of the first aspect, in the sixth possible implementation manner, before obtaining the second semantic feature data corresponding to the fusion data based on the above encoder, the method further includes: obtaining the word vector corresponding to each character in the multiple characters constituting the fusion data. Generating a second word vector matrix according to the multiple word vectors corresponding to the multiple characters constituting the fusion data. Obtaining the second semantic feature data corresponding to the fusion data based on the above encoder includes: inputting the second word vector matrix into the encoder and obtaining the semantic feature vector output by the encoder as the second semantic feature data corresponding to the fusion data.

[0018] Combined with any one of the first aspect to the sixth possible implementation manner of the first aspect, in the seventh possible implementation manner, the above text data is obtained by converting the input voice, and the input voice is the vehicle control voice. The method further includes: controlling the vehicle to execute the target function according to the above semantic intention and the above entity.

[0019] Combined with any one of the first to seventh possible implementation manners of the first aspect, in the eighth possible implementation manner, the above encoder is trained according to the first training sample and the second training sample, the above semantic intention decoder is trained according to the above first training sample, the above semantic entity decoder is trained according to the above second training sample, the above first training sample includes sample text data and the intention category corresponding to the above sample text data pre-annotated, the above second training sample includes sample fusion data and the entity corresponding to the above sample fusion data pre-annotated, and the above sample fusion data includes the above sample text data and the intention category corresponding to the above sample text data.

[0020] Combined with the eighth possible implementation manner of the first aspect, in the ninth possible implementation manner, the above encoder, the above semantic intention decoder, and the above semantic entity decoder are trained through the following steps:

[0021] Obtain the third semantic feature data corresponding to the above sample text data through the initial encoder, and predict the semantic intention corresponding to the above third semantic feature data through the initial semantic intention decoder;

[0022] Adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the above sample text data pre-annotated, so as to train the above initial encoder and the above initial semantic intention decoder to obtain the first encoder and the above semantic intention decoder;

[0023] Obtain the fourth semantic feature data corresponding to the above sample fusion data through the above first encoder, and predict the entity corresponding to the above fourth semantic feature data based on the initial semantic entity decoder;

[0024] Adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the predicted entity and the entity corresponding to the above sample fusion data pre-annotated, so as to train the above first encoder and the above initial semantic entity decoder to obtain the above encoder and the above semantic entity decoder.

[0025] In this application, the encoder trained by the decoupled pre-training method is beneficial to improving the encoding effect of the encoder, and thus helps to improve the accuracy of semantic understanding when used subsequently. Among them, the decoupled pre-training method can be understood as follows: training a first encoder and a semantic intent decoder based on the sample text data and the intent categories corresponding to the pre-annotated sample text data. Training the first encoder and the initial semantic entity decoder according to the sample fusion data and the entities corresponding to the pre-annotated sample fusion data to obtain the final encoder and semantic entity decoder. In other words, if there are 100,000 pieces of sample text data, then based on the training method of decoupled pre-training, the encoder will learn the features of 200,000 pieces of data (the 200,000 pieces of data include 100,000 pieces of sample text data and 100,000 pieces of sample fusion data), so that the learning effect of the encoder is better.

[0026] In a second aspect, this application provides a semantic understanding method, which includes: obtaining a first training sample and a second training sample, where the first training sample includes sample text data and the intent categories corresponding to the sample text data, and the second training sample includes sample fusion data and the entities corresponding to the pre-annotated sample fusion data, and the sample fusion data includes the sample text data and the intent categories corresponding to the pre-annotated sample text data. Obtaining the third semantic feature data corresponding to the sample text data through the initial encoder, and predicting the semantic intent corresponding to the third semantic feature data through the initial semantic intent decoder. Adjusting the weight parameters of the initial encoder and the initial semantic intent decoder according to the predicted semantic intent and the intent categories corresponding to the pre-annotated sample text data to train the initial encoder and the initial semantic intent decoder to obtain a first encoder and a semantic intent decoder. Obtaining the fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predicting the entity corresponding to the fourth semantic feature data based on the initial semantic entity decoder. Adjusting the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entities corresponding to the pre-annotated sample fusion data to train the first encoder and the initial semantic entity decoder to obtain an encoder and a semantic entity decoder.

[0027] In this application, the encoder trained by the decoupled pre-training method is beneficial to improving the encoding effect of the encoder, and thus helps to improve the accuracy of semantic understanding when used subsequently. Among them, the decoupled pre-training method can be understood as follows: training a first encoder and a semantic intent decoder based on the sample text data and the intent categories corresponding to the pre-annotated sample text data. Training the first encoder and the initial semantic entity decoder according to the sample fusion data and the entities corresponding to the pre-annotated sample fusion data to obtain the final encoder and semantic entity decoder.

[0028] In combination with the second aspect, in the first possible implementation manner, adjusting the weight parameters of the above-mentioned initial encoder and the above-mentioned initial semantic intention decoder according to the above-mentioned semantic intention obtained by prediction and the intention category corresponding to the above-mentioned sample text data pre-annotated includes: determining a first loss according to the above-mentioned semantic intention obtained by prediction and the intention category corresponding to the above-mentioned sample text data pre-annotated. Adjusting the weight parameters of the above-mentioned initial encoder and the above-mentioned initial semantic intention decoder according to the above-mentioned first loss.

[0029] In combination with any one of the second aspect to the first possible implementation manner of the second aspect, in the second possible implementation manner, adjusting the weight parameters of the above-mentioned first encoder and the above-mentioned initial semantic entity decoder according to the above-mentioned entity obtained by prediction and the entity corresponding to the above-mentioned sample fusion data pre-annotated to train the above-mentioned first encoder and the above-mentioned initial semantic entity decoder includes: determining a second loss according to the above-mentioned entity obtained by prediction and the entity corresponding to the above-mentioned sample fusion data pre-annotated. Adjusting the weight parameters of the above-mentioned first encoder and the above-mentioned initial semantic entity decoder according to the above-mentioned second loss.

[0030] In combination with any one of the second aspect to the second possible implementation manner of the second aspect, in the third possible implementation manner, the above-mentioned method includes: obtaining text data converted from the input voice. Wherein, the input voice can be collected based on a microphone, and the input voice is converted into corresponding text data through a speech recognition system. Then, obtaining first semantic feature data corresponding to the above-mentioned text data through an encoder, and determining the semantic intention corresponding to the above-mentioned first semantic feature data through a semantic intention decoder to realize the intention recognition of the above-mentioned text data. Generating fusion data according to the above-mentioned text data and the above-mentioned semantic intention, and obtaining second semantic feature data corresponding to the above-mentioned fusion data based on the above-mentioned encoder. Determining the entity corresponding to the above-mentioned second semantic feature data based on a semantic entity decoder to realize the entity recognition associated with the above-mentioned semantic intention in the above-mentioned text data.

[0031] After the present application obtains the semantic intention corresponding to the text data according to the encoder and the semantic intention decoder, by fusing the semantic intention as prior information with the text data, and then encoding the fused fusion data through the same encoder, and further determining the entity based on the semantic entity decoder, the conflict between the semantic intention information and the entity information can be reduced, which is beneficial to improving the semantic understanding accuracy in voice interaction.

[0032] In a third aspect, the present application provides a semantic understanding device, which includes: a transceiver unit for acquiring text data. A processing unit for obtaining first semantic feature data corresponding to the above-mentioned text data through an encoder, and determining a semantic intention corresponding to the first semantic feature data through a semantic intention decoder, so as to realize intention recognition of the above-mentioned text data. The above-mentioned processing unit is further configured to generate fusion data according to the above-mentioned text data and the above-mentioned semantic intention, and obtain second semantic feature data corresponding to the fusion data based on the above-mentioned encoder. The above-mentioned processing unit is further configured to determine an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to realize entity recognition associated with the semantic intention in the above-mentioned text data.

[0033] In combination with the third aspect, in a first possible implementation manner, the above-mentioned processing unit is further configured to: determine a first word vector matrix corresponding to the above-mentioned text data. Input the first word vector matrix into the encoder, and obtain a semantic feature vector output by the encoder as the first semantic feature data corresponding to the above-mentioned text data.

[0034] In combination with the first possible implementation manner of the third aspect, in a second possible implementation manner, the above-mentioned processing unit is further configured to: perform character splitting on the above-mentioned text data to obtain a plurality of characters included in the above-mentioned text data. Obtain a character word vector, a position word vector, and a character type word vector corresponding to each character in the above-mentioned plurality of characters, where the above-mentioned position word vector is used to represent the position of the character in the above-mentioned text data. Sum the character word vector, the position word vector, and the character type word vector corresponding to each character to obtain a word vector corresponding to each character. Generate a first word vector matrix corresponding to the above-mentioned text data according to the plurality of word vectors corresponding to the above-mentioned plurality of characters.

[0035] In combination with the second possible implementation manner of the third aspect, in a third possible implementation manner, the above-mentioned processing unit is further configured to: obtain a character word vector query table, a position word vector query table, and a character type word vector query table. Wherein, the above-mentioned character word vector query table includes n character word vectors corresponding to n characters. The above-mentioned position word vector query table includes m position word vectors corresponding to m positions. The above-mentioned character type word vector query table includes character type word vectors corresponding to non-padding characters and character type word vectors corresponding to padding characters. Both n and m are integers greater than 0. Obtain the character word vector corresponding to each character in the above-mentioned plurality of characters from the above-mentioned character word vector query table, obtain the position word vector corresponding to the position of each character in the above-mentioned text data from the above-mentioned position word vector query table, and obtain the character type word vector corresponding to the non-padding character as the character type word vector corresponding to each character from the above-mentioned character type word vector query table.

[0036] Combined with any one of the first to third possible implementation manners of the third aspect, in the fourth possible implementation manner, the above-mentioned encoder includes j encoding units. The input of the first encoding unit in the j encoding units is the above-mentioned first word vector input matrix. The input of any encoding unit after the first encoding unit is the output of the upper encoding unit of any encoding unit. The output of the j-th encoding unit in the j encoding units is the above-mentioned first semantic feature data, where j is an integer greater than 0.

[0037] Combined with any one of the third aspect to the fourth possible implementation manner of the third aspect, in the fifth possible implementation manner, the above-mentioned processing unit is further configured to: splice the above-mentioned text data and the above-mentioned semantic intention to obtain the fusion data corresponding to the above-mentioned text data and the above-mentioned semantic intention.

[0038] Combined with the fifth possible implementation manner of the third aspect, in the sixth possible implementation manner, the above-mentioned processing unit is further configured to: obtain the word vector corresponding to each character among the multiple characters constituting the above-mentioned fusion data. Generate a second word vector matrix according to the multiple word vectors corresponding to the multiple characters constituting the above-mentioned fusion data. Input the above-mentioned second word vector matrix into the above-mentioned encoder, and obtain the semantic feature vector output by the above-mentioned encoder as the second semantic feature data corresponding to the above-mentioned fusion data.

[0039] Combined with any one of the third aspect to the sixth possible implementation manner of the third aspect, in the seventh possible implementation manner, the above-mentioned input voice is a vehicle control voice; the above-mentioned processing unit is further configured to: control the above-mentioned vehicle to execute a target function according to the above-mentioned semantic intention and the above-mentioned entity.

[0040] Combined with any one of the third aspect to the seventh possible implementation manner of the third aspect, in the eighth possible implementation manner, the above-mentioned encoder is trained according to a first training sample and a second training sample. The above-mentioned semantic intention decoder is trained according to the above-mentioned first training sample. The above-mentioned semantic entity decoder is trained according to the above-mentioned second training sample. The above-mentioned first training sample includes sample text data and the intention category corresponding to the pre-annotated above-mentioned sample text data. The above-mentioned second training sample includes sample fusion data and the entity corresponding to the pre-annotated above-mentioned sample fusion data. The above-mentioned sample fusion data includes the above-mentioned sample text data and the intention category corresponding to the above-mentioned sample text data.

[0041] Combined with the eighth possible implementation manner of the third aspect, in the ninth possible implementation manner, the above-mentioned processing unit is further configured to train the above-mentioned encoder, the above-mentioned semantic intention decoder, and the above-mentioned semantic entity decoder through the following steps:

[0042] Obtain the third semantic feature data corresponding to the above sample text data through the initial encoder, and predict the semantic intention corresponding to the above third semantic feature data through the initial semantic intention decoder;

[0043] Adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated above sample text data, so as to train the above initial encoder and the above initial semantic intention decoder to obtain the first encoder and the above semantic intention decoder;

[0044] Obtain the fourth semantic feature data corresponding to the above sample fusion data through the above first encoder, and predict the entity corresponding to the above fourth semantic feature data based on the initial semantic entity decoder;

[0045] Adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated above sample fusion data, so as to train the above first encoder and the above initial semantic entity decoder to obtain the above encoder and the above semantic entity decoder.

[0046] Fourthly, the present application provides a semantic understanding device, which includes: a transceiver unit, configured to obtain a first training sample and a second training sample, wherein the above first training sample includes sample text data and the intention category corresponding to the above sample text data, and the above second training sample includes sample fusion data and the entity corresponding to the pre-annotated above sample fusion data, and the above sample fusion data includes the above sample text data and the intention category corresponding to the pre-annotated above sample text data. A processing unit, configured to obtain the third semantic feature data corresponding to the above sample text data through the initial encoder, and predict the semantic intention corresponding to the above third semantic feature data through the initial semantic intention decoder. The above processing unit is further configured to adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated above sample text data, so as to train the above initial encoder and the above initial semantic intention decoder to obtain the first encoder and the semantic intention decoder. The above processing unit is further configured to obtain the fourth semantic feature data corresponding to the above sample fusion data through the above first encoder, and predict the entity corresponding to the above fourth semantic feature data based on the initial semantic entity decoder. The above processing unit is further configured to adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated above sample fusion data, so as to train the above first encoder and the above initial semantic entity decoder to obtain the encoder and the semantic entity decoder.

[0047] In combination with the fourth aspect, in a first possible implementation manner, the above processing unit is specifically configured to:

[0048] Determine a first loss according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data;

[0049] Adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the above first loss.

[0050] In combination with any one of the fourth aspect to the first possible implementation manner of the fourth aspect, in a second possible implementation manner, the above processing unit is specifically configured to:

[0051] Determine a second loss according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data;

[0052] Adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the above second loss.

[0053] In combination with any one of the fourth aspect to the second possible implementation manner of the fourth aspect, in a third possible implementation manner, the above transceiver unit is further configured to obtain text data; the above processing unit is further configured to obtain first semantic feature data corresponding to the text data through an encoder, and determine a semantic intention corresponding to the first semantic feature data through a semantic intention decoder, so as to implement intention recognition of the text data; the above processing unit is further configured to generate fusion data according to the text data and the semantic intention, and obtain second semantic feature data corresponding to the fusion data based on the encoder; the above processing unit is further configured to determine an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to implement entity recognition associated with the semantic intention in the text data.

[0054] Fifth aspect, an embodiment of the present application provides a terminal device. The terminal device includes a memory, a transceiver, and a processor; wherein, the memory, the transceiver, and the processor are connected through a communication bus, or the processor and the transceiver are used to be coupled with the memory. The memory is used to store a set of program codes, and the processor is used to call the program codes stored in the memory to execute the semantic understanding method provided by the first aspect and / or any possible implementation manner in the first aspect, so it can also achieve the beneficial effects of the method provided by the first aspect, or the processor is used to call the program codes stored in the memory to execute the semantic understanding method provided by the second aspect and / or any possible implementation manner in the second aspect, so it can also achieve the beneficial effects of the method provided by the second aspect.

[0055] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a terminal, the terminal is caused to execute the semantic understanding method provided in the first aspect and / or any possible implementation manner in the first aspect, and the beneficial effects achieved by the method provided in the first aspect can also be realized. Or, the terminal is caused to execute the semantic understanding method provided in the second aspect and / or any possible implementation manner in the second aspect, and the beneficial effects achieved by the method provided in the second aspect can also be realized.

[0056] Seventh aspect, an embodiment of the present application provides a communication device. The communication device may be a chip or multiple chips working together. The communication device includes an input device coupled to the communication device (such as a chip) for executing the technical solution provided in the first aspect or the second aspect of the embodiment of the present application. It should be understood that "coupled" here means that two components are directly or indirectly combined with each other. This combination may be fixed or movable, and this combination allows the flow of liquid, electricity, electrical signals or other types of signals to communicate between the two components.

[0057] Eighth aspect, an embodiment of the present application provides a computer program product containing instructions. When the computer program product is run on a terminal, the terminal is caused to execute the semantic understanding method provided in the first aspect or the second aspect, and the beneficial effects achieved by the method provided in the first aspect or the second aspect can also be realized.

[0058] In the semantic understanding method provided in the present application, text data is acquired. First semantic feature data corresponding to the text data is acquired through an encoder, and the semantic intention corresponding to the first semantic feature data is determined through a semantic intention decoder to implement the intention recognition of the text data. Fusion data is generated based on the text data and the semantic intention, and second semantic feature data corresponding to the fusion data is acquired based on the encoder. Entities corresponding to the second semantic feature data are determined based on a semantic entity decoder to implement the entity recognition associated with the semantic intention in the text data. By using the method provided in the present application, the accuracy of semantic understanding can be improved. Description of the Drawings

[0059] Figure 1 is a flowchart of the semantic understanding method;

[0060] Figure 2 is a schematic flowchart of the semantic understanding method provided in an embodiment of the present application;

[0061] Figure 3 is another schematic flowchart of the semantic understanding method provided in an embodiment of the present application;

[0062] Figure 4 is a schematic diagram of an application scenario of semantic intention recognition provided in an embodiment of the present application;

[0063] Figure 5 It is a schematic structural diagram of an encoder provided by an embodiment of the present application;

[0064] Figure 6 It is a schematic structural diagram of a multi-head attention mechanism layer provided by an embodiment of the present application;

[0065] Figure 7 It is another schematic structural diagram of an encoder provided by an embodiment of the present application;

[0066] Figure 8 It is a schematic diagram of an application scenario for entity recognition provided by an embodiment of the present application;

[0067] Figure 9 It is another schematic flowchart of a semantic understanding method provided by an embodiment of the present application;

[0068] Figure 10 It is a schematic structural diagram of a semantic understanding device provided by an embodiment of the present application;

[0069] Figure 11 It is another schematic structural diagram of a semantic understanding device provided by an embodiment of the present application. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0071] From the feature phone era to the smartphone era, the way of interaction between humans and machines has been changing. Especially in recent years, voice interaction has been a research hotspot. The semantic understanding method provided in this application can be widely applied to industries and scenarios such as smart home, automotive, and intelligent customer service. Among them, especially with the rise of the Internet of Vehicles and intelligent vehicles, more and more functions are carried on in-vehicle units, making voice interaction in the automotive intelligent cockpit particularly important. For example, in aspects such as vehicle control, intelligent navigation, and multimedia entertainment, users can control the vehicle to perform corresponding functions through voice commands. For instance, in terms of vehicle control, users can adjust the windows or the in-vehicle temperature through voice commands such as "Open the front right window" and "It's too hot. Please help me turn on the maximum cooling mode", or users can also adjust the vehicle's rearview mirror or shift gears through voice commands, which are specifically determined according to the actual application scenario and are not limited here. In terms of intelligent navigation, users can control the vehicle navigation service through voice commands such as "I want to go to Pudong Avenue". In terms of multimedia entertainment, users can control the playback, pause, and song switching of music through voice commands such as "I want to listen to 'Chili Fragrance'", "I don't want to listen to music anymore", or "I don't like this song. Change it to 'Chili Fragrance'", which are not limited here. Obviously, in the process of voice interaction, if accurate voice control is to be achieved, the accurate understanding of the true meaning in the user's voice by the terminal device is the key. For the convenience of description, hereinafter, this application will be described by taking in-vehicle semantic understanding as an example, that is, taking the voice interaction in the automotive intelligent cockpit as an example.

[0072] Among them, the semantic understanding method provided by the embodiments of the present application can be adapted to various terminal devices. For example, the above terminal device can be an intelligent in-vehicle computer (or called an in-vehicle terminal) installed on the vehicle dashboard, or the above terminal device can also be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., which are not limited herein. Optionally, the semantic understanding method provided by the embodiments of the present application can also be executed by a chip, such as an in-vehicle chip, etc., which are not limited herein. Optionally, the semantic understanding method provided by the embodiments of the present application can also be executed by a server or a chip in the server. Among them, the server can be a local server or a cloud server, etc., which are not limited herein. For the convenience of description, the following will take the terminal device as an example for illustration. Specifically, the semantic understanding method proposed by the present application can perform intent recognition on the user's speech through the terminal device and entity recognition on the entities included in the user's speech that are associated with the user's intent, and then can accurately control the corresponding functions of the vehicle according to the recognized semantic intent and entity. Among them, the entities in the present application include personal names, place names, organization names, dates, currencies, percentages, and other custom entities. For example, other custom entities can be role names, dish names, etc., which are specifically determined according to the actual application scenario and are not limited herein. For example, for the text data 1 "I want to listen to Qi Li Xiang", the semantic intent is to play music, and the entity associated with the semantic intent is "Qi Li Xiang". Another example is that for the text data 2 "Navigate to ZhuHai Bridge", the semantic intent is to start navigation, and the entity associated with the semantic intent is "ZhuHai Bridge".

[0073] Specifically, when the present application performs intent recognition and entity recognition on the user's speech, the user's speech can be converted into corresponding text data, and then the converted text data is processed. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the semantic understanding method provided by the embodiments of the present application. As Figure 1 shown, after obtaining the text data converted from the user's speech, the text data can be input into the encoder to extract the semantic features of the text data through the encoder. Then, the semantic feature data corresponding to the text data output by the encoder is input into the semantic intent decoder to obtain the semantic intent output by the semantic intent decoder. Among them, to improve the recognition accuracy of the entities associated with the semantic intent, the semantic intent output by the above semantic intent decoder can be used as prior information to be fused with the original text data (i.e., the text data corresponding to the user's speech), and then the semantic features of the fused data are extracted based on the same encoder. Finally, the semantic feature data corresponding to the fused data output by the encoder is input into the semantic entity decoder to output the entity recognition result through the semantic entity decoder. Finally, corresponding function control is performed according to the obtained semantic intent and entity.

[0074] Among them, the current NLU processing usually completes intent recognition and entity recognition as two independent tasks. For example, please refer to Figure 2 , Figure 2 which is a flowchart of a semantic understanding method. As Figure 2 shown, after obtaining the text data corresponding to the user's speech, the text data can be input into the first encoder and the second encoder respectively for encoding to obtain the semantic feature data corresponding to the text data output by the first encoder and the semantic feature data corresponding to the text data output by the second encoder. Then, the semantic feature data output by the first encoder is input into the semantic intent decoder to obtain the semantic intent output after passing through the semantic intent decoder. At the same time, the semantic feature data output by the second encoder is input into the semantic entity decoder to obtain the entity output after passing through the semantic entity decoder. Finally, corresponding function control is performed according to the obtained semantic intent and entity.

[0075] Obviously, compared with the prior art that uses the first encoder and the second encoder to extract the semantic feature data of the text data respectively, the present application extracts the semantic feature data of the text data by using an encoder, and extracts the semantic feature data of the fusion data including the text data and the semantic intent based on the same encoder, so that more semantic features can be learned based on this encoder, which is beneficial to improving the accuracy of semantic understanding.

[0076] Next, the technical solutions in the present application will be described in detail in conjunction with Figures 3 to 10 . Please refer to Figure 3 , Figure 3 which is another schematic flowchart of the semantic understanding method provided by the embodiment of the present application. It can be understood that the semantic understanding method provided by the present application can be executed by a terminal device, or a chip in the terminal device, or a server, or a chip in the server, etc., which is not limited herein. For the convenience of description, the following embodiments of the present application will be described by taking the terminal device as an example. As Figure 3 shown, the above semantic understanding method may include the following steps:

[0077] S101. Obtain text data.

[0078] In some feasible embodiments, the voice signal of the user can be collected through the microphone or microphone array (i.e., multiple arranged microphones) of the terminal device. Furthermore, by inputting the voice signal into an automated speech recognition (ASR) system, the voice signal can be converted into text data. For example, during the startup or driving of the vehicle, the voice signal T of the user can be collected through the microphone of the terminal device (i.e., in-vehicle terminal) installed on the vehicle. Then, the collected voice signal T is transmitted to the in-vehicle voice recognition system. When the in-vehicle voice recognition system receives the voice signal T, the voice signal T can be converted into text data X. Therefore, by parsing the text data, the semantic understanding of the voice input by the user can be achieved.

[0079] In some feasible embodiments, after the terminal device collects the voice signal, it can also perform noise reduction processing such as echo cancellation and crosstalk cancellation on the collected voice signal, and then input the voice signal after noise reduction processing into the ASR system for text conversion.

[0080] Optionally, in some feasible embodiments, if the above terminal device does not have the function of voice collection or voice recognition, the terminal device can also receive the text data corresponding to the user's voice from another terminal device with the functions of voice collection and voice recognition through wired or wireless communication methods.

[0081] Optionally, in some feasible embodiments, the terminal device can also obtain text data from the local memory or cloud storage, or can also receive the pre-stored text data from another terminal device through wired or wireless communication methods for semantic understanding, or can also obtain the text data input by the user on the display interface of the terminal device for semantic understanding, etc., which is not limited here.

[0082] S102. Obtain the first semantic feature data corresponding to the text data through the encoder, and determine the semantic intention corresponding to the first semantic feature data through the semantic intention decoder.

[0083] In some feasible embodiments, semantic feature data corresponding to text data, i.e., the first semantic feature data, can be obtained through an encoder, and then the semantic intent corresponding to the first semantic feature data can be determined through a semantic intent decoder. That is to say, the above-mentioned encoder can be used to extract semantic features such as the morphology and syntax of the input text (i.e., text data), and then provide information input for the semantic intent decoder. Specifically, when performing intent recognition on text data, the first word vector matrix corresponding to the above-mentioned text data can be determined first. Then, the first word vector matrix is input into the encoder to obtain the semantic feature vector output by the encoder as the above-mentioned first semantic feature data. Among them, the first word vector matrix includes word vectors corresponding to each of the multiple characters that make up the text data. It should be understood that the semantic feature data involved in this application can be features such as the morphology and syntax in text information.

[0084] Among them, by performing character splitting on text data, multiple characters included in the text data can be obtained. Among them, the above-mentioned character splitting of text data can be understood as splitting the text data by character. For example, for text data 1 "I want to listen to "JASMINE", by performing character splitting on this text data, 6 characters "I", "want", "to", "listen", "to", "JASMINE" can be obtained. Another example, for text data 2 "Navigate to Zhuhai Bridge", by performing character splitting on this text data, 7 characters "navigate", "to", "Zhuhai", "Bridge" can be obtained. Further, by obtaining the character word vector, position word vector, and character type word vector corresponding to each of the multiple characters that make up the text data, the obtained character word vector, position word vector, and character type word vector corresponding to each character can be summed to obtain the word vector corresponding to each character. Therefore, according to the multiple word vectors corresponding to multiple characters, the first word vector matrix corresponding to the text data can be generated. It should be understood that the character word vector corresponding to each character is the vector representation of each character. The position word vector corresponding to each character is used to represent the position of the character in the text data. The character type word vector is used to represent the character type to which the character belongs. Among them, the vector dimensions of the character word vector, position word vector, and character type word vector corresponding to each character are the same. For example, the vector dimensions of the character word vector, position word vector, and character type word vector in the embodiments of this application can be 768 or the vector dimensions are all 312, etc., which are specifically determined according to the actual application scenario and are not limited here.

[0085] Generally speaking, in order to improve processing efficiency, a character word vector query table, a position word vector query table and a character type word vector query table can be pre-set. Among them, the character word vector query table includes n character word vectors corresponding to n characters. The position word vector query table includes m position word vectors corresponding to m positions. Among them, n and m are both integers greater than 0. The character type word vector query table includes character type word vectors corresponding to non-fill characters and character type word vectors corresponding to fill characters, that is, the present application may include 2 types of characters, namely fill characters and non-fill characters, wherein each character constituting the text data belongs to the type of non-fill characters. Therefore, when the multiple characters constituting the text data are determined, the character word vector query table, the position word vector query table and the character type word vector query table can be obtained, and then the character word vector corresponding to each character in the multiple characters constituting the text data can be obtained from the character word vector query table, the position word vector corresponding to the position of each character in the text data can be obtained from the position word vector query table, and the character type word vector corresponding to the non-fill character can be obtained from the character type word vector query table as the character type word vector corresponding to each character. Furthermore, the character word vector, position word vector and character type word vector corresponding to each character are summed to obtain the word vector corresponding to each character.

[0086] For example, taking the character "I" in the text data 1 "I want to listen to Qilixiang" as an example, assuming that the character word vector corresponding to "I" obtained from the character word vector query table, the position word vector query table and the character type word vector query table are [1,2,…,6], the position word vector is [3,4,…,1], and the character type word vector is [1,1,…,1] respectively, then by summing the above three vectors (i.e., the character word vector, the position word vector and the character type word vector), the word vector corresponding to the character "I" can be obtained as [5, 7,…, 8]. For another example, taking the character "想" in the text data 1 "我要听七里香" as an example, assuming that the character word vector corresponding to "想" obtained from the character word vector query table, the position word vector query table and the character type word vector query table are [0,1,…,3], the position word vector is [2,7,…,9], and the character type word vector is [1,1,…,1] respectively, then by summing the above three vectors (i.e., the character word vector, the position word vector and the character type word vector), the word vector corresponding to the character "想" can be obtained as [3, 9,…, 13]. Similarly, the word vectors corresponding to each character in the characters "听", "七", "里" and "香" in the text data 1 can be obtained respectively.

[0087] Among them, after obtaining multiple word vectors corresponding to multiple characters that make up the text data, the first word vector matrix corresponding to the text data can be generated according to the multiple word vectors corresponding to the multiple characters. Generally speaking, the length of the character sequence input to the encoder is fixed, that is, the matrix size of the word vector matrix of the input encoder is fixed. Therefore, when the length of the character data (that is, the number of characters that make up the character data) is less than the length of the character sequence, the character length can be padded to the specified length of the character sequence. Among them, the length of the character sequence of the input encoder in the embodiments of the present application can be 512, 256, etc., which is specifically determined according to the actual application scenario and is not limited here.

[0088] For example, please refer to Figure 4 , Figure 4 is a schematic diagram of the application scenario of semantic intent recognition provided by the embodiments of the present application. As Figure 4 shown, assume that the length of the character sequence that the preset encoder can process is 10. For the text data 1 "I want to listen to Qi Li Xiang" corresponding to the voice input by the user, by adding the character [CLS] at the beginning of the text data 1 "I want to listen to Qi Li Xiang" to indicate the start of the text data, and adding the character [SEP] at the end of the text data 1 "I want to listen to Qi Li Xiang" to indicate the end of the text data, then according to the length 6 of the text data 1 itself (that is, the 6 characters included in the text data 1), as well as the character [CLS] and the character [SEP], it can still be determined that the input length of the text data (here the input length is 8 characters) is less than the set processing length 10 of the encoder. Therefore, the text data can be padded in length based on the padding character [PAD]. That is to say, when the input length of the text data is less than the length of the character sequence preset by the encoder, the text data can be padded in length by the padding character [PAD] so that the input of the text data meets the settings of the encoder. Among them, as Figure 4 shown, the [CLS] character in front of the first character "I" of the text data is used to indicate the start of the text data, and the [SEP] character after the last character "Xiang" of the text data is used to indicate the end of the text data. As Figure 4 shown in the scenario, 2 padding characters [PAD] can be added after the [SEP] character to pad the length of the text data to 10 characters. Further, by obtaining the character word vector, position word vector, and character type word vector corresponding to each of the above 10 characters, the character word vector, position word vector, and character type word vector corresponding to each character can be summed, and the vector obtained by the summation can be determined as the word vector corresponding to the character. It should be understood that since the [CLS] character, [SEP] character, and each character that makes up the text data are all non-padding characters, their character type word vectors are all character type word vectors E corresponding to non-padding characters 11Since the character [PAD] is a padding character, its character type word vectors are all the character type word vectors E corresponding to the padding character. 00 Also, assume that the vector dimension of the word vector corresponding to each character is 768. Then, the character word vectors corresponding to the above 10 characters can generate a first word vector matrix of 10×768, where each row in the first word vector matrix represents the word vector corresponding to a character. It should be understood that by inputting the above first word vector matrix into a pre-trained encoder, the semantic feature vector output after passing through the encoder can be obtained as the first semantic feature data corresponding to the text data. Further, by inputting the first semantic feature data into the semantic intent decoder, the semantic intent output after passing through the semantic intent decoder can be obtained. As Figure 4 shown, the semantic intent of text data 1 "I want to listen to Qi Li Xiang" can be obtained as "Music", that is, to play music.

[0089] Among them, the encoder in the embodiment of the present application can be composed of j layers of encoding units. Here, j is an integer greater than 0. For example, j can be equal to 6. It should be understood that the j layers of encoding units are serially connected. That is to say, the input of the first layer of encoding units in the j layers of encoding units is the first word vector matrix corresponding to the text data, and the input of any layer of encoding units after the first layer of encoding units is the output of the upper layer of encoding units of this layer of encoding units. For example, please refer to Figure 5 , Figure 5 is a schematic structural diagram of the encoder provided by the embodiment of the present application. As Figure 5 shown, assume that the encoder includes a total of 4 layers of encoding units (i.e., j = 4), which are the first layer of encoding units, the second layer of encoding units, the third layer of encoding units, and the fourth layer of encoding units respectively. Among them, the input of the first layer of encoding units is the first word vector matrix corresponding to the text data, the input of the second layer of encoding units is the output of the first layer of encoding units, the input of the third layer of encoding units is the output of the second layer of encoding units, and so on. Finally, the output of the fourth layer of encoding units can be used as the output of the entire encoder.

[0090] It should be understood that in each of the j layers of encoding units included in the encoder, each layer of encoding units can be composed of a multi-head attention mechanism layer and a feed-forward layer, and the multi-head attention mechanism layer and the feed-forward layer are serially connected. For example, please refer to Figure 5, taking the first - layer encoding unit as an example, the feed - forward layer included in the first - layer encoding unit is connected after the multi - head attention mechanism layer. Therefore, for the first - layer encoding unit, by inputting the first word - vector matrix into the multi - head attention mechanism layer of the first - layer encoding unit, the output data obtained through the multi - head attention mechanism layer of the first - layer encoding unit can be input into the feed - forward layer of the first - layer encoding unit. Furthermore, the output data obtained after passing through the feed - forward layer of the first - layer encoding unit can be used as the output data of the first - layer encoding unit and input into the next - layer encoding unit of the first - layer encoding unit (i.e., the second - layer encoding unit), and so on, until the output of the feed - forward layer of the fourth - layer encoding unit is obtained as the output of the encoder. That is to say, the semantic feature vector output by the fourth - layer encoding unit can be used as the first semantic feature data corresponding to the text data. Finally, the semantic intention decoder can be used to determine the semantic intention corresponding to the first semantic feature data to achieve the intention recognition of the text data.

[0091] It should be understood that the multi - head attention mechanism layer is the most core layer in each layer of the encoding unit. Among them, the multi - head attention mechanism layer can learn a weight for each character in the input word - vector matrix. Please refer to Figure 6 , Figure 6 which is the structural schematic diagram of the multi - head attention mechanism layer provided by the embodiments of the present application. Among them, for each word - vector corresponding to each character included in the input word - vector matrix (i.e., the first word - vector matrix), a corresponding query Query vector (for convenience of description, hereinafter referred to as the Q vector), key Key vector (for convenience of description, hereinafter referred to as the K vector), and value Value vector (for convenience of description, hereinafter referred to as the V vector) can be generated. That is to say, for the first word - vector matrix, the first word - vector matrix can be multiplied by three different weight matrices W Q , W K , W V respectively, to obtain the Q - vector matrix, K - vector matrix, and V - vector matrix corresponding to the first word - vector matrix. Furthermore, by performing Attention calculation on the Q - vector matrix, K - vector matrix, and V - vector matrix, the output result of each head among multiple heads can be obtained. As Figure 6 shown, it is assumed that there are h heads in the multi - head attention mechanism layer, and h is an integer greater than 1. Among them, the Attention calculation method for each of the h heads is the same, but the weight matrices W Q , W K , W V for linear transformation in each head are different. For convenience of description, hereinafter, the embodiments of the present application take the Attention calculation of any one head i among the h heads as an example for illustration. Specifically, it is assumed that the parameters for linear transformation of any one head i are W i Q , Wi K and W i V If so, after the first word vector matrix X undergoes a linear transformation, the corresponding Q i vector matrix, K i vector matrix, and V i vector matrix can be obtained. Furthermore, by performing Attention calculation on the Q i vector matrix, K i vector matrix, and V i vector matrix, the scaled dot-product Attention result head i output by any head i can be obtained. Among them:

[0092] Q i = W i Q X;

[0093] K i = W i K X;

[0094] V i = W i V X;

[0095]

[0096] head i = Attention(Q i , K i , V i )

[0097] Among them, d k is the hidden neuron dimension of the multi-head attention mechanism layer. Generally speaking, d k is 64.

[0098] It should be understood that when each of the h heads undergoes the above calculation process, h scaled dot-product Attention results can be obtained. Suppose the h scaled dot-product Attention results are head1, head2,..., head h respectively. Therefore, the h calculation results can be concatenated, and the concatenated result can be linearly transformed again to obtain the output MultiHead(Q, K, V) of the multi-head attention mechanism layer. Among them:

[0099] MultiHead(Q, K, V) = Concat(head1,..., head h )W O

[0100] Among them, W O is the weight matrix used for the linear transformation.

[0101] Optionally, in some feasible embodiments, in addition to the attention mechanism layer and the forward propagation layer, each encoding unit in the above j-layer encoding unit may further include a first vector normalization layer and a second vector normalization layer. It should be understood that the vector normalization layer can be used to normalize the output vector and simplify the learning difficulty. For example, please refer to Figure 7 , Figure 7 which is another structural schematic diagram of the encoder provided in the embodiment of the present application. Assume j = 4, that is to say, this encoder includes a total of 4 encoding units, namely the first-layer encoding unit, the second-layer encoding unit, the third-layer encoding unit, and the fourth-layer encoding unit. As Figure 7 shown, taking the first-layer encoding unit as an example, the first-layer encoding unit may include a multi-head attention mechanism layer, a first vector normalization layer, a forward propagation layer, and a second vector normalization layer. Among them, the multi-head attention mechanism layer is connected to the forward propagation layer through the first vector normalization layer, and the output of the forward propagation layer is connected to the second vector normalization layer. That is to say, the multi-head attention mechanism layer is connected to the first vector normalization layer, the first vector normalization layer is connected to the forward propagation layer, and the forward propagation layer is connected to the second normalization layer. Therefore, for the first-layer encoding unit, by inputting the first word vector matrix into the multi-head attention mechanism layer of the first-layer encoding unit, the output data obtained through the multi-head attention mechanism layer of the first-layer encoding unit can be input into the first vector normalization layer of the first-layer encoding unit for normalization or normalization processing. Then, the output data after the normalization processing is input into the forward propagation layer of the first-layer encoding unit, and after passing through the forward propagation layer and the second vector normalization layer in sequence, the output data after the normalization processing by the second normalization layer can be used as the output data of the first-layer encoding unit, and then input into the next-layer encoding unit (i.e., the second-layer encoding unit) of the first-layer encoding unit, and so on, until the semantic feature vector output by the second vector normalization layer of the fourth-layer encoding unit is used as the first semantic feature data corresponding to the text data. Finally, by inputting the semantic feature vector output by the above fourth-layer encoding unit into the semantic intent decoder, the semantic intent determined based on the semantic intent decoder can be obtained to realize the intent recognition of the text data. That is to say, for a certain layer of encoding unit, when the output MultiHead(Q, K, V) of the multi-head attention mechanism layer included in this layer of encoding unit is obtained, the output of the multi-head attention mechanism layer can be input into the first vector normalization layer, where the first vector normalization layer can satisfy:

[0102] x = LayerNorm(MultiHead(Q, K, V) + Sublayer(MultiHead(Q, K, V)))

[0103] Among them, x is the output of the first vector normalization layer, LayerNorm represents the normalization calculation operation, MultiHead(Q, K, V) is the output of the multi-head attention mechanism layer, and Sublayer represents the residual calculation operation.

[0104] After obtaining the output result x after normalization by the first vector normalization layer, the high-dimensional mapping of the vector space can be realized through the feed-forward layer to extract abstract high-dimensional lexical, syntactic and other semantic information. Specifically, the feed-forward layer can satisfy:

[0105] FFN(x) = max(0, xW1 + b1)W2 + b2

[0106] Among them, FFN(x) represents the output of the feed-forward layer, max represents the maximum value calculation operation, W1 and W2 represent the weight matrices in the feed-forward layer, b1 and b2 are the bias parameters of the weight matrices, and x is the output of the first vector normalization layer.

[0107] Furthermore, after obtaining the output result FFN(x) after the feed-forward layer, the output result of the feed-forward layer can be further input into the second vector normalization layer, where the second vector normalization layer satisfies:

[0108] V = LayerNorm(FFN(x) + Sublayer(FFN(x)))

[0109] Among them, V represents the output result of the second vector normalization layer, LayerNorm represents the normalization calculation operation, FFN(x) is the output of the feed-forward layer, and Sublayer represents the residual calculation operation.

[0110] It should be understood that after obtaining the output result after normalization by the second normalization layer of this layer's encoding unit, this output result can be used as the output result of this layer's encoding unit to further input this output result into the next encoding unit of this layer's encoding unit, and so on, until obtaining the output result of the second vector normalization layer of the j-th layer's encoding unit as the first semantic feature data corresponding to the text data.

[0111] In some feasible implementation manners, after obtaining the output result of the encoder, that is, the output result of the above-mentioned j-th layer's encoding unit, the semantic intention corresponding to this first semantic feature data can be determined according to the output result of the encoder (i.e., the first semantic feature data) and the semantic intention decoder, so as to realize the intention recognition of the text data. Among them, the semantic intention decoder can satisfy:

[0112] y1 = F(W3 * v1 + b3)

[0113] Among them, y1 represents the semantic intention, W3 represents the weight matrix of the semantic intention decoder, b3 represents the bias parameter, and v1 represents the first row vector of the output matrix V of the encoder.

[0114] S103. Generate fusion data according to the text data and the semantic intention, and obtain second semantic feature data corresponding to the fusion data based on the encoder.

[0115] In some feasible embodiments, after determining the semantic intention based on the above-mentioned semantic intention decoder, fusion data can be generated according to the text data and the obtained semantic intention, and second semantic feature data corresponding to the fusion data can be obtained based on the encoder. It should be understood that by splicing the text data and the semantic intention, the fusion data corresponding to the text data and the semantic intention can be obtained. Further, word vectors corresponding to each character in the multiple characters constituting the fusion data are obtained, and finally, a second word vector matrix is generated according to the multiple word vectors corresponding to the multiple characters constituting the fusion data. Among them, the method for obtaining the word vector corresponding to each character in the multiple characters constituting the fusion data is the same as the method for obtaining the word vector corresponding to each character in the multiple characters constituting the text data, that is, the character word vector, position word vector, and character type word vector corresponding to each character can be obtained by querying the character word vector query table, position word vector query table, and character type query table, respectively, and then by summing the character word vector, position word vector, and character type word vector corresponding to each character, the word vector corresponding to the character is obtained. Finally, according to the word vectors corresponding to each character constituting the fusion data, a second word vector matrix corresponding to the fusion data is generated. The specific implementation method for generating the second word vector matrix corresponding to the fusion data can refer to the description of generating the first word vector matrix corresponding to the text data above, and will not be elaborated here. It is not difficult to understand that after determining the second word vector matrix corresponding to the fusion data, by inputting the second word vector matrix into the encoder, the semantic feature vector output by the encoder can be obtained as the second semantic feature data corresponding to the fusion data.

[0116] S104. Determine the entity corresponding to the second semantic feature data based on the semantic entity decoder to realize entity recognition in the text data associated with the semantic intention.

[0117] In some feasible embodiments, after determining the second semantic feature data corresponding to the fusion data, by inputting the second semantic feature data into the semantic entity decoder, the entity included in the text data associated with the semantic intention can be determined according to the semantic entity decoder. Among them, the semantic entity decoder can satisfy:

[0118] y i = F(W4 * T + b4)

[0119] Among them, yi It represents the output result of the semantic entity decoder. W4 represents the weight matrix of the semantic entity decoder, b4 represents the bias parameter, and T represents the second semantic feature data corresponding to the fused data output by the encoder.

[0120] It should be understood that when the semantic intention corresponding to the text data and the entity associated with the semantic intention are determined based on the above steps, the vehicle can be controlled to perform the target function according to the semantic intention and the entity. For example, please refer to Figure 8 , Figure 8 It is a schematic diagram of the application scenario of entity recognition provided by an embodiment of the present application. As Figure 8 shown, assume that the text data 1 corresponding to the user-entered voice is "I want to listen to Qi Li Xiang". By determining the first word vector matrix corresponding to the text data 1, the first word vector matrix can be input into the encoder 210 to obtain the semantic feature vector output by the encoder 210 as the first semantic feature data corresponding to the text data 1. By inputting the above first semantic feature data into the semantic intention decoder 211, the semantic intention of the user can be determined based on the output result of the semantic intention decoder 211 as "Music", that is, to play music. Further, by splicing the text data 1 and the semantic intention, fused data can be obtained. Further, by determining the second word vector matrix corresponding to the fused data, the second word vector matrix can be input into the above encoder 210, and the second semantic feature data corresponding to the fused data can be determined by the encoder 210. Finally, by inputting the second semantic feature data into the semantic entity decoder 212, the entity recognition result output by the semantic entity decoder 212 can be obtained as "Qi Li Xiang". Therefore, according to the above-determined semantic intention and the entity associated with the semantic intention, the in-vehicle terminal can play music for the user, and the song played is "Qi Li Xiang".

[0121] In the present application, after obtaining the text data converted from the input voice, the first semantic feature data corresponding to the text data can be obtained through the encoder, and then the semantic intention corresponding to the first semantic feature data can be determined according to the semantic intention decoder to realize the intention recognition of the text data. Fused data is generated according to the text data and the semantic intention, and the second semantic feature data corresponding to the fused data can be obtained based on the above encoder, and then the entity corresponding to the second semantic feature data can be determined based on the semantic entity decoder to realize the entity recognition associated with the semantic intention in the text data.

[0122] It should be understood that this application realizes the extraction of semantic features of text data and the extraction of semantic features of fusion data based on the same encoder, which can improve the accuracy of semantic understanding. That is to say, this application trains the encoder through a decoupled pre-training method. Specifically, the decoupled pre-training method can be understood as follows: the encoder in this application is trained according to the first training sample and the second training sample, the semantic intention decoder is trained according to the first training sample, and the semantic entity decoder is trained according to the second training sample. The first training sample includes sample text data and the intention category corresponding to the pre-annotated sample text data. The second training sample includes sample fusion data and the entity corresponding to the pre-annotated sample fusion data. The sample fusion data includes the sample text data and the intention category corresponding to the sample text data. For ease of understanding, the training processes of the encoder, semantic intention decoder, and semantic entity decoder provided in this application will be illustrated by examples below.

[0123] Please refer to Figure 9 , Figure 9 which is another flowchart of the semantic understanding method provided by the embodiment of this application. As Figure 9 shown, the above semantic understanding method may include the following steps:

[0124] S201. Obtain the first training sample and the second training sample.

[0125] In some feasible implementation manners, when training the initial encoder, the initial semantic intention decoder, and the initial semantic entity decoder, the training sample set may be obtained first. Among them, the training sample set includes the first training sample and the second training sample. The first training sample includes sample text data and the intention category corresponding to the pre-labeled sample text data. The second training sample includes sample fusion data and the entity corresponding to the pre-annotated sample fusion data. It can be understood that the sample fusion data is determined according to the sample text data included in the first training sample and the intention category corresponding to the pre-labeled sample text data. Generally speaking, the entity corresponding to the pre-annotated sample fusion data is the entity associated with the intention category included in the sample text data.

[0126] S202. Obtain the third semantic feature data corresponding to the sample text data through the initial encoder, and predict the semantic intention corresponding to the third semantic feature data through the initial semantic intention decoder.

[0127] In some feasible embodiments, when training the initial encoder and the initial semantic intention decoder according to the first training sample, it is possible to adjust the weight parameters of the initial encoder and the weight parameters of the initial semantic intention decoder to obtain the first encoder and the semantic intention decoder. When training the initial encoder and the initial semantic entity decoder according to the second training sample, it is possible to further adjust the weight parameters of the adjusted initial encoder above (i.e., adjust the weight parameters of the first encoder) and adjust the weight parameters of the initial semantic entity decoder. That is to say, the third semantic feature data corresponding to the sample text data can be obtained through the initial encoder, and the semantic intention corresponding to the third semantic feature data can be predicted through the initial semantic intention decoder. Specifically, by inputting the word vector matrix corresponding to the sample text data included in the first training sample into the initial encoder, after encoding by the initial encoder, the semantic feature data output by the initial encoder can be input into the initial semantic intention decoder, and then the semantic intention output by the initial semantic intention decoder can be obtained. Therefore, the weight parameters of the initial encoder and the initial semantic intention decoder can be adjusted according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data to train the initial encoder and the initial semantic intention decoder.

[0128] S203. Adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data to train the initial encoder and the initial semantic intention decoder, and obtain the first encoder and the semantic intention decoder.

[0129] In some feasible embodiments, the weight parameters of the initial encoder and the initial semantic intention decoder can be adjusted according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data to train the initial encoder and the initial semantic intention decoder, and obtain the first encoder and the semantic intention decoder. Specifically, the first loss between the semantic intention output by the initial semantic intention decoder and the intention category corresponding to the pre-marked sample text data can be calculated, and then the weight parameters of the initial encoder and the initial semantic intention decoder can be adjusted according to the calculated first loss to obtain the first encoder and the semantic intention decoder.

[0130] S204. Obtain the fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predict the entity corresponding to the fourth semantic feature data based on the initial semantic entity decoder.

[0131] In some feasible embodiments, the fourth semantic feature data corresponding to the sample fusion data can be obtained through a first encoder, and the entity corresponding to the fourth semantic feature data can be predicted based on an initial semantic entity decoder. Therefore, the weight parameters of the first encoder and the initial semantic entity decoder can be adjusted according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder. Specifically, the word vector matrix corresponding to the sample fusion data composed of the sample text data and the intention category corresponding to the pre-labeled sample text data included in the second training sample can be input into the encoder obtained after being trained by the first training sample, so as to obtain the semantic feature data corresponding to the sample fusion data output by the encoder. By inputting the semantic feature data corresponding to the sample fusion data into the initial semantic entity decoder, the entity output by the initial semantic entity decoder can be obtained. Therefore, the first encoder and the initial semantic entity decoder can be trained according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data.

[0132] S205. Adjust the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder to obtain an encoder and a semantic entity decoder.

[0133] In some feasible embodiments, by adjusting the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, the first encoder and the initial semantic entity decoder can be trained to obtain an encoder and a semantic entity decoder. Specifically, by calculating the second loss between the entity output by the initial semantic entity decoder and the pre-labeled entity, the weight parameters of the encoder and the weight parameters of the initial semantic entity decoder can be further adjusted according to the second loss until it is determined based on the test sample set that the encoder obtained by training with the first training sample and the second training sample, the semantic intention decoder obtained by training with the first training sample, and the semantic entity decoder obtained by training with the second training sample meet the target convergence condition, and then the training process ends.

[0134] Specifically, when determining whether the adjusted encoder, semantic intent decoder, and semantic entity decoder meet the target convergence condition, a test sample set can be obtained first. Among them, the test sample set includes multiple test text data, the intent categories corresponding to each pre-labeled test text data, and the entities included in each test text data. Therefore, each test text in the test sample set can be encoded based on the adjusted encoder to obtain the semantic feature data corresponding to each test text data. Then, intent recognition is performed on the semantic feature data corresponding to each test text data based on the adjusted semantic intent decoder to obtain the user intents corresponding to each test text data. Further, each fusion data composed of each test text and the output user intents is encoded based on the adjusted encoder to obtain the semantic feature data corresponding to each fusion data. Then, entity recognition is performed on the semantic feature data based on the adjusted semantic entity decoder to obtain the entities included in each test text. Among them, if it is determined that the intent recognition accuracy is not less than the first accuracy threshold based on the user intents corresponding to each test text output by the semantic intent decoder and the intent categories corresponding to each pre-labeled test text, and, if it is determined that the entity recognition accuracy is not less than the second accuracy threshold based on the entities included in each test text output by the semantic entity decoder and the entities included in each pre-labeled test text, then it can be determined that the adjusted encoder, semantic intent decoder, and semantic entity decoder meet the target convergence condition. Therefore, the training can be ended, and the trained encoder, semantic intent decoder, and semantic entity decoder can be used as the Figure 3 encoder, semantic intent decoder, and semantic entity decoder used in each of the above Figure 3 steps. Correspondingly, if it is determined based on the test sample set that the adjusted encoder, semantic intent decoder, and semantic entity decoder do not meet the target convergence condition, the training process continues until the target convergence condition is met and the training ends. Among them, the process of processing text data based on the trained encoder, semantic intent decoder, and semantic entity decoder can be referred to the

[0135] implementation process described in each of the above steps and will not be elaborated here.

[0136] Next, the semantic understanding device in the present application will be described.

[0137] In the case of adopting an integrated unit, refer to Figure 10 , Figure 10It is a schematic structural diagram of a semantic understanding device provided by an embodiment of the present application. The semantic understanding device can be a terminal device or a chip in the terminal device, such as an in-vehicle chip, etc. Optionally, the semantic understanding device can also be a server or a chip in the server, etc. For example, the server can be a cloud server, etc., which is not limited here. As Figure 10 shown, the semantic understanding device includes a processing unit 1001 and a transceiver unit 1002. Among them, the transceiver unit 1002 can be a transceiver or a communication interface, and the processing unit 1001 can be one or more processors. The semantic understanding device can be used to implement the functions of the terminal device, chip, or server involved in the above method embodiments.

[0138] Exemplarily, the semantic understanding device can be a terminal device. The terminal device can be either a network element in a hardware device, a software function running on dedicated hardware, or a virtualized function instantiated on a platform (such as a cloud platform). Optionally, the semantic understanding device can also include a storage unit (not shown in the figure) for storing the program code and data of the semantic understanding device.

[0139] Exemplarily, when the semantic understanding device is a chip, the transceiver unit 1002 can be an interface, a pin, or a circuit, etc. The interface can be used to input data to be processed to the processor and can output the processing result of the processor outward. In a specific implementation, the interface can be a general purpose input output (GPIO) interface, which can be connected to multiple peripheral devices (such as a liquid crystal display (LCD), a camera, a radio frequency (RF) module, an antenna, etc.). The interface is connected to the processor through a bus.

[0140] The processing unit 1001 can be a processor, and the processor can execute the computer execution instructions stored in the storage unit to enable the chip to execute Figure 3 the method involved in the embodiment.

[0141] Further, the processor may include a controller, an arithmetic unit, and registers. Exemplarily, the controller is mainly responsible for instruction decoding and issuing control signals for the operations corresponding to the instructions. The arithmetic unit is mainly responsible for performing fixed-point or floating-point arithmetic operations, shift operations, and logical operations, etc., and may also perform address operations and conversions. The registers are mainly responsible for storing register operands and intermediate operation results temporarily stored during the execution of instructions, etc. In a specific implementation, the hardware architecture of the processor may be an application specific integrated circuits (ASIC) architecture, a microprocessor without interlocked piped stages architecture (MIPS), an advanced RISC machines (ARM) architecture, or a network processor (NP) architecture, etc. The processor may be single-core or multi-core.

[0142] The storage unit may be a storage unit within the chip, such as registers, caches, etc. The storage unit may also be a storage unit located outside the chip, such as a Read Only Memory (ROM) or other types of static storage devices that can store static information and instructions, a Random Access Memory (RAM), etc.

[0143] It should be noted that the functions corresponding to the processor and the interface can be implemented through hardware design, software design, or a combination of software and hardware, and there is no limitation here.

[0144] Specifically, in one design, the semantic understanding device can be used to process the obtained text data based on a pre-trained encoder, a semantic intent decoder, and an entity intent decoder. Specifically:

[0145] The transceiver unit 1002 is used to obtain text data;

[0146] The processing unit 1001 is used to obtain the first semantic feature data corresponding to the above text data through the encoder, and determine the semantic intent corresponding to the first semantic feature data through the semantic intent decoder, so as to realize the intent recognition of the above text data;

[0147] The above processing unit 1001 is further used to generate fusion data based on the above text data and the above semantic intent, and obtain the second semantic feature data corresponding to the fusion data based on the above encoder;

[0148] The above processing unit 1001 is further configured to determine an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to implement entity recognition in the text data associated with the semantic intention.

[0149] Optionally, the processing unit 1001 is further configured to:

[0150] Determine a first word vector matrix corresponding to the text data;

[0151] Input the first word vector matrix into an encoder, and obtain a semantic feature vector output by the encoder as the first semantic feature data corresponding to the text data.

[0152] Optionally, the processing unit 1001 is further configured to:

[0153] Perform character splitting on the text data to obtain multiple characters included in the text data;

[0154] Obtain a character word vector, a position word vector, and a character type word vector corresponding to each character in the multiple characters, where the position word vector is used to represent the position of the character in the text data;

[0155] Sum the character word vector, the position word vector, and the character type word vector corresponding to each character to obtain a word vector corresponding to each character;

[0156] Generate a first word vector matrix corresponding to the text data according to the multiple word vectors corresponding to the multiple characters.

[0157] Optionally, the processing unit 1001 is further configured to:

[0158] Obtain a character word vector query table, a position word vector query table, and a character type word vector query table, where the character word vector query table includes n character word vectors corresponding to n characters, the position word vector query table includes m position word vectors corresponding to m positions, and the character type word vector query table includes character type word vectors corresponding to non-padding characters and character type word vectors corresponding to padding characters, and both n and m are integers greater than 0;

[0159] Obtain the character word vector corresponding to each character in the multiple characters from the character word vector query table, obtain the position word vector corresponding to the position of each character in the text data from the position word vector query table, and obtain the character type word vector corresponding to the non-padding character as the character type word vector corresponding to each character from the character type word vector query table.

[0160] Optionally, the above encoder includes j encoding units. The input of the first encoding unit in the above j encoding units is the above first word vector input matrix. The input of any encoding unit after the first encoding unit is the output of the previous encoding unit of the above any encoding unit. The output of the j-th encoding unit in the above j encoding units is the above first semantic feature data, where j is an integer greater than 0.

[0161] Optionally, the above processing unit 1001 is further configured to:

[0162] Concatenate the above text data and the above semantic intention to obtain the fusion data corresponding to the above text data and the above semantic intention.

[0163] Optionally, the above processing unit 1001 is further configured to:

[0164] Obtain the word vector corresponding to each character in the multiple characters that make up the above fusion data;

[0165] Generate a second word vector matrix according to the multiple word vectors corresponding to the multiple characters that make up the above fusion data;

[0166] Input the above second word vector matrix into the above encoder, and obtain the semantic feature vector output by the above encoder as the second semantic feature data corresponding to the above fusion data.

[0167] Optionally, the above text data is obtained by converting the input voice, and the input voice is a vehicle control voice; the above processing unit 1001 is further configured to:

[0168] Control the above vehicle to execute the target function according to the above semantic intention and the above entity.

[0169] Optionally, the above encoder is trained according to a first training sample and a second training sample. The above semantic intention decoder is trained according to the above first training sample. The above semantic entity decoder is trained according to the above second training sample. The above first training sample includes sample text data and the intention category corresponding to the above pre-annotated sample text data. The above second training sample includes sample fusion data and the entity corresponding to the above pre-annotated sample fusion data. The above sample fusion data includes the above sample text data and the intention category corresponding to the above sample text data.

[0170] Optionally, the above processing unit 1001 is further configured to train the above encoder, the above semantic intention decoder, and the above semantic entity decoder through the following steps:

[0171] Obtain the third semantic feature data corresponding to the above sample text data through an initial encoder, and predict the semantic intention corresponding to the above third semantic feature data through an initial semantic intention decoder;

[0172] Adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the above semantic intention obtained by prediction and the intention category corresponding to the above pre-annotated sample text data, so as to train the above initial encoder and the above initial semantic intention decoder to obtain a first encoder and the above semantic intention decoder;

[0173] Obtain the fourth semantic feature data corresponding to the above sample fusion data through the above first encoder, and predict the entity corresponding to the above fourth semantic feature data based on the initial semantic entity decoder;

[0174] Adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the above entity obtained by prediction and the entity corresponding to the above pre-annotated sample fusion data, so as to train the above first encoder and the above initial semantic entity decoder to obtain the above encoder and the above semantic entity decoder.

[0175] In another design, please also refer to Figure 10 The semantic understanding device can also be used to train an encoder, a semantic intention decoder, and a semantic entity decoder. It can be understood that the semantic understanding device used to train the encoder, the semantic intention decoder, and the semantic entity decoder can be the same device as the above semantic understanding device for processing text data, or the semantic understanding device used to train the encoder, the semantic intention decoder, and the semantic entity decoder can be a different device from the above semantic understanding device for processing text data, etc., which is not limited here. Among them, when the semantic understanding device is a device for training an encoder, a semantic intention decoder, and a semantic entity decoder, it includes:

[0176] A transceiver unit 1002, configured to obtain a first training sample and a second training sample, where the above first training sample includes sample text data and the intention category corresponding to the above sample text data, the above second training sample includes sample fusion data and the entity corresponding to the above pre-annotated sample fusion data, and the above sample fusion data includes the above sample text data and the intention category corresponding to the above pre-annotated sample text data;

[0177] A processing unit 1001, configured to obtain third semantic feature data corresponding to the above sample text data through an initial encoder, and predict the semantic intention corresponding to the above third semantic feature data through an initial semantic intention decoder;

[0178] The above processing unit 1001 is further configured to adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the predicted above semantic intention and the intention category corresponding to the above pre-annotated sample text data, so as to train the above initial encoder and the above initial semantic intention decoder to obtain a first encoder and a semantic intention decoder;

[0179] The above processing unit 1001 is further configured to obtain fourth semantic feature data corresponding to the above sample fusion data through the above first encoder, and predict an entity corresponding to the above fourth semantic feature data based on an initial semantic entity decoder;

[0180] The above processing unit 1001 is further configured to adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the predicted above entity and the entity corresponding to the above pre-annotated sample fusion data, so as to train the above first encoder and the above initial semantic entity decoder to obtain an encoder and a semantic entity decoder.

[0181] Optionally, the above processing unit 1001 is specifically configured to:

[0182] Determine a first loss according to the predicted above semantic intention and the intention category corresponding to the above pre-annotated sample text data;

[0183] Adjust the weight parameters of the above initial encoder and the above initial semantic intention decoder according to the above first loss.

[0184] Optionally, the above processing unit 1001 is specifically configured to:

[0185] Determine a second loss according to the predicted above entity and the entity corresponding to the above pre-annotated sample fusion data;

[0186] Adjust the weight parameters of the above first encoder and the above initial semantic entity decoder according to the above second loss.

[0187] Optionally, the above transceiver unit 1002 is further configured to obtain text data;

[0188] The above processing unit 1001 is further configured to obtain first semantic feature data corresponding to the above text data through the encoder, and determine a semantic intention corresponding to the above first semantic feature data through a semantic intention decoder, so as to implement intention recognition of the above text data;

[0189] The above processing unit 1001 is further configured to generate fusion data according to the above text data and the above semantic intention, and obtain second semantic feature data corresponding to the above fusion data based on the above encoder;

[0190] The above processing unit 1001 is further configured to determine the entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to implement entity recognition in the above text data associated with the above semantic intention.

[0191] It should be understood that the above semantic understanding device may correspondingly execute the steps of the foregoing method embodiments, and the above operations or functions of each unit in the semantic understanding device are respectively for implementing the corresponding operations executed by the terminal device in the foregoing method embodiments. Among them, the corresponding beneficial effects can refer to the method embodiments, and for the sake of brevity, they will not be elaborated here.

[0192] The semantic understanding device of the embodiments of the present application has been introduced above. The following introduces possible product forms of the semantic understanding device. It should be understood that any product form Figure 10 with the functions of the above-mentioned semantic understanding device falls within the protection scope of the embodiments of the present application. It should also be understood that the following introduction is only for example and does not limit the product form of the semantic understanding device of the embodiments of the present application to this.

[0193] As a possible product form, the above-mentioned semantic understanding device of the embodiments of the present application can be implemented by a general bus architecture.

[0194] For the sake of convenience of description, refer to Figure 11 , Figure 11 which is another structural schematic diagram of the semantic understanding device provided by the embodiments of the present application. The semantic understanding device can be a terminal device, or a chip in a terminal device, or a server, or a chip in a server, etc. Figure 11 Only the main components of the semantic understanding device are shown. In addition to the processor 1101 and the transceiver 1102, the above semantic understanding device may further include a memory 1103 and an input / output device (not shown in the figure).

[0195] The processor 1101 is mainly used to process communication protocols and communication data, and control the entire semantic understanding device, execute software programs, and process data of software programs. The memory 1103 is mainly used to store software programs and data. The transceiver 1102 may include a control circuit and an antenna. The control circuit is mainly used for the conversion between baseband signals and radio frequency signals and the processing of radio frequency signals. The antenna is mainly used to transmit and receive radio frequency signals in the form of electromagnetic waves. The input / output device, such as a touch screen, a display screen, a keyboard, etc., is mainly used to receive data input by the user and output data to the user.

[0196] When the semantic understanding device is turned on, the processor 1101 can read the software program in the memory 1103, interpret and execute the instructions of the software program, and process the data of the software program. When data needs to be sent wirelessly, the processor 1101 performs baseband processing on the data to be sent, and outputs the baseband signal to the RF circuit. The RF circuit performs RF processing on the baseband signal and then sends the RF signal outward in the form of electromagnetic waves through the antenna. When data is sent to the semantic understanding device, the RF circuit receives the RF signal through the antenna, converts the RF signal into a baseband signal, and outputs the baseband signal to the processor 1101. The processor 1101 converts the baseband signal into data and processes the data.

[0197] In another implementation, the above-mentioned RF circuit and antenna can be set independently of the processor performing baseband processing. For example, in a distributed scenario, the RF circuit and antenna can be arranged remotely from the semantic understanding device.

[0198] The processor 1101 , the transceiver 1102 , and the memory 1103 may be connected via a communication bus.

[0199] In one design, the semantic understanding device can be used to perform the functions of the terminal device in the above method embodiment: the processor 1101 can be used to execute Figure 3 Steps S102 to S104 in the embodiment, and / or for executing Figure 9 Steps S202 to S205 in the embodiment, and / or other processes for performing the technology described herein; the transceiver 1102 may be used to perform Figure 3 Step S101 in, and / or for executing Figure 9 Step S201 in, and / or other processes for the technology described herein.

[0200] In any of the above designs, the processor 1101 may include a transceiver for implementing the receiving and sending functions. For example, the transceiver may be a transceiver circuit, or an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing the receiving and sending functions may be separate or integrated. The above transceiver circuit, interface, or interface circuit may be used for reading and writing code / data, or the above transceiver circuit, interface, or interface circuit may be used for transmitting or delivering signals.

[0201] In any of the above designs, the processor 1101 may store instructions, which may be computer programs. The computer programs run on the processor 1101, and the semantic understanding device may execute the method described in any of the above method embodiments. The computer program may be fixed in the processor 1000, in which case the processor 1101 may be implemented by hardware.

[0202] In one implementation, the semantic understanding device may include circuitry, which may implement the functions of sending, receiving, or obtaining in the foregoing method embodiments. The processors and transceivers described in this application may be implemented on an integrated circuit (IC), analog IC, radio frequency integrated circuit (RFIC), mixed-signal IC, application specific integrated circuit (ASIC), printed circuit board (PCB), electronic device, etc. The processors and transceivers may also be fabricated using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal-oxide-semiconductor (NMOS), P-type metal-oxide semiconductor (PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0203] The scope of the semantic understanding device described in this application is not limited thereto, and the structure of the semantic understanding device may not be limited by Figure 11 . The semantic understanding device may be an independent device or may be part of a larger device. For example, the semantic understanding device may be:

[0204] (1) An independent integrated circuit IC, or chip, or chip system or subsystem;

[0205] (2) A set of one or more ICs, optionally, the IC set may also include storage components for storing data and computer programs;

[0206] (3) An ASIC, such as a modem;

[0207] (4) A module that can be embedded in other devices;

[0208] (5) A receiver, terminal, smart terminal, cellular phone, wireless device, handset, mobile unit, vehicle-mounted device, network device, cloud device, artificial intelligence device, etc.;

[0209] (6) Others, etc.

[0210] As a possible product form, the terminal device described in the embodiments of the present application can be implemented by a general-purpose processor.

[0211] The general-purpose processor for implementing the terminal device includes a processing circuit and an input / output interface that is internally connected and communicates with the processing circuit.

[0212] In one design, the general-purpose processor can be used to execute the functions of the terminal device in the foregoing method embodiments. Specifically, the processing circuit is used to execute Figure 3 the steps S102 to S104 in Figure 9 , and / or used to execute Figure 3 the steps S202 to S205 in Figure 9 , and / or used to execute other processes of the technologies described herein; the input / output interface is used to execute the step S101 in

[0213] , and / or used to execute the step S201 in

[0214] , and / or used to execute other processes of the technologies described herein. Figure 3 Figure 9 It should be understood that the semantic understanding devices of the above various product forms have any functions of the terminal device in the above method embodiments, can correspondingly implement the steps in the above method embodiments, and achieve corresponding technical effects. For the sake of brevity, they will not be elaborated here.

[0215] Figure 3 Figure 9

[0216] The embodiments of the present application further provide a computer-readable storage medium, in which computer program code is stored. When the above processor executes the computer program code, it is used to execute the methods of each step in the foregoing embodiments Figure 3 or Figure 9

[0217] The embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, it causes the computer to execute the methods of each step in the foregoing embodiments Figure 3 or Figure 9

[0217] The embodiments of the present application further provide a semantic understanding device. The device can exist in the product form of a chip. The structure of the device includes a processor and an interface circuit. The processor is used to communicate with other devices through a receiving circuit, so that the device executes the methods of each step in the foregoing embodiments Figure 3 or Figure 9

[0217] The steps of the methods or algorithms described in connection with the disclosure of the present application may be implemented in hardware or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable hard disk, compact disc read-only memory (CD-ROM), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Of course, the storage medium may also be part of the processor. The processor and the storage medium may be located in an ASIC. Additionally, the ASIC may be located in a core network interface device. Of course, the processor and the storage medium may also exist as discrete components in the core network interface device.

[0218] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer-readable storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general or special-purpose computer.

[0219] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc., made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.

Claims

1. A semantic understanding method, characterized in that, The method includes: Obtaining text data; Obtaining first semantic feature data corresponding to the text data through an encoder, and determining a semantic intention corresponding to the first semantic feature data through a semantic intention decoder, so as to realize intention recognition of the text data; Generating fusion data according to the text data and the semantic intention, and obtaining second semantic feature data corresponding to the fusion data based on the encoder; Determining an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to realize entity recognition associated with the semantic intention in the text data; Wherein, before obtaining the first semantic feature data corresponding to the text data through the encoder, the method further includes: Performing character splitting on the text data to obtain multiple characters included in the text data; Obtaining a character word vector, a position word vector, and a character type word vector corresponding to each character in the multiple characters, wherein the position word vector is used to represent the position of the character in the text data; Summing the character word vector, the position word vector, and the character type word vector corresponding to each character to obtain a word vector corresponding to each character; Generating a first word vector matrix corresponding to the text data according to the multiple word vectors corresponding to the multiple characters; Encoding the text data through the encoder to obtain first semantic feature data corresponding to the text data, including: Inputting the first word vector matrix into the encoder, and obtaining a semantic feature vector output by the encoder as the first semantic feature data corresponding to the text data.

2. The method according to claim 1, wherein Obtaining the character word vector, the position word vector, and the character type word vector corresponding to each character in the multiple characters, including: Obtaining a character word vector query table, a position word vector query table, and a character type word vector query table, wherein the character word vector query table includes n character word vectors corresponding to n characters, the position word vector query table includes m position word vectors corresponding to m positions, and the character type word vector query table includes character type word vectors corresponding to non-padding characters and character type word vectors corresponding to padding characters, and both n and m are integers greater than 0; Obtaining the character word vector corresponding to each character in the multiple characters from the character word vector query table, obtaining the position word vector corresponding to the position of each character in the text data from the position word vector query table, and obtaining the character type word vector corresponding to the non-padding character as the character type word vector corresponding to each character from the character type word vector query table.

3. The method according to claim 1 or 2, characterized in that, The encoder includes j layers of encoding units. The input of the first layer of encoding units in the j layers of encoding units is the first word vector matrix, the input of any layer of encoding units after the first layer of encoding units is the output of the previous layer of encoding units of the any layer of encoding units, and the output of the jth layer of encoding units in the j layers of encoding units is the first semantic feature data, where j is an integer greater than 0.

4. The method according to claim 1 or 2, characterized in that, Generating the fusion data according to the text data and the semantic intention, including: Concatenate the text data and the semantic intention to obtain the fusion data corresponding to the text data and the semantic intention.

5. The method according to claim 4, wherein Before obtaining the second semantic feature data corresponding to the fusion data based on the encoder, the method further includes: Obtain the word vector corresponding to each character in the multiple characters that make up the fusion data; Generate a second word vector matrix according to the multiple word vectors corresponding to the multiple characters that make up the fusion data; Based on the encoder, obtaining the second semantic feature data corresponding to the fusion data includes: Input the second word vector matrix into the encoder, and obtain the semantic feature vector output by the encoder as the second semantic feature data corresponding to the fusion data.

6. The method according to claim 1 or 2, characterized in that, The text data is obtained by converting the input voice, and the input voice is the vehicle control voice; the method further includes: Control the vehicle to execute the target function according to the semantic intention and the entity.

7. The method according to claim 1 or 2, characterized in that, The encoder is trained according to the first training sample and the second training sample, the semantic intention decoder is trained according to the first training sample, the semantic entity decoder is trained according to the second training sample, the first training sample includes the sample text data and the intention category corresponding to the pre-annotated sample text data, and the second training sample includes the sample fusion data and the entity corresponding to the pre-annotated sample fusion data, and the sample fusion data includes the sample text data and the intention category corresponding to the sample text data.

8. The method according to claim 7, wherein The encoder, the semantic intention decoder, and the semantic entity decoder are trained through the following steps: Obtain the third semantic feature data corresponding to the sample text data through the initial encoder, and predict the semantic intention corresponding to the third semantic feature data through the initial semantic intention decoder; Adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data, so as to train the initial encoder and the initial semantic intention decoder to obtain the first encoder and the semantic intention decoder; Obtain the fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predict the entity corresponding to the fourth semantic feature data based on the initial semantic entity decoder; Adjust the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder to obtain the encoder and the semantic entity decoder.

9. A semantic understanding method, characterized in that, The method includes: Obtain the first training sample and the second training sample, wherein the first training sample includes the sample text data and the intention category corresponding to the sample text data, and the second training sample includes the sample fusion data and the entity corresponding to the pre-annotated sample fusion data, and the sample fusion data includes the sample text data and the intention category corresponding to the pre-annotated sample text data; Obtain the third semantic feature data corresponding to the sample text data through the initial encoder, and predict the semantic intention corresponding to the third semantic feature data through the initial semantic intention decoder; Adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data, so as to train the initial encoder and the initial semantic intention decoder to obtain a first encoder and a semantic intention decoder; Obtain the fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predict the entity corresponding to the fourth semantic feature data based on the initial semantic entity decoder; Adjust the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder to obtain an encoder and a semantic entity decoder.

10. The method according to claim 9, wherein The adjusting the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data includes: Determine a first loss according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data; Adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the first loss.

11. The method according to claim 9 or 10, characterized in that, The adjusting the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data to train the first encoder and the initial semantic entity decoder includes: Determine a second loss according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data; Adjust the weight parameters of the first encoder and the initial semantic entity decoder according to the second loss.

12. The method according to any one of claims 9 or 10, characterized in that, The method includes: Obtain text data; Obtain the first semantic feature data corresponding to the text data through the encoder, and determine the semantic intention corresponding to the first semantic feature data through the semantic intention decoder, so as to realize the intention recognition of the text data; Generate fusion data according to the text data and the semantic intention, and obtain the second semantic feature data corresponding to the fusion data based on the encoder; Determine the entity corresponding to the second semantic feature data based on the semantic entity decoder, so as to realize the entity recognition associated with the semantic intention in the text data.

13. A semantic understanding device, characterized in that, The device includes: A transceiver unit, configured to obtain text data; A processing unit, configured to obtain the first semantic feature data corresponding to the text data through an encoder, and determine the semantic intention corresponding to the first semantic feature data through a semantic intention decoder, so as to realize the intention recognition of the text data; The processing unit is further configured to generate fusion data according to the text data and the semantic intention, and obtain the second semantic feature data corresponding to the fusion data based on the encoder; The processing unit is further configured to determine an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to implement entity recognition in the text data associated with the semantic intention; Wherein, before obtaining the first semantic feature data corresponding to the text data through the encoder, the processing unit is further configured to: Perform character splitting on the text data to obtain multiple characters included in the text data; Obtain a character word vector, a position word vector, and a character type word vector corresponding to each character in the multiple characters, where the position word vector is used to represent the position of the character in the text data; Sum the character word vector, the position word vector, and the character type word vector corresponding to each character to obtain a word vector corresponding to each character; Generate a first word vector matrix corresponding to the text data according to the multiple word vectors corresponding to the multiple characters; When encoding the text data through the encoder to obtain the first semantic feature data corresponding to the text data, the processing unit specifically is configured to: Input the first word vector matrix into the encoder, and obtain the semantic feature vector output by the encoder as the first semantic feature data corresponding to the text data.

14. The device according to claim 13, characterized in that, The processing unit is further configured to: Obtain a character word vector query table, a position word vector query table, and a character type word vector query table, where the character word vector query table includes n character word vectors corresponding to n characters, the position word vector query table includes m position word vectors corresponding to m positions, and the character type word vector query table includes character type word vectors corresponding to non-padding characters and character type word vectors corresponding to padding characters, and both n and m are integers greater than 0; Obtain the character word vector corresponding to each character in the multiple characters from the character word vector query table, obtain the position word vector corresponding to the position of each character in the text data from the position word vector query table, and obtain the character type word vector corresponding to the non-padding character as the character type word vector corresponding to each character from the character type word vector query table.

15. The device according to claim 13 or 14, characterized in that, The encoder includes j encoding units. The input of the first encoding unit in the j encoding units is the first word vector matrix, the input of any encoding unit after the first encoding unit is the output of the upper encoding unit of the any encoding unit, and the output of the jth encoding unit in the j encoding units is the first semantic feature data, where j is an integer greater than 0.

16. The device according to claim 13 or 14, characterized in that The processing unit is further configured to: Concatenate the text data and the semantic intention to obtain fusion data corresponding to the text data and the semantic intention.

17. The device according to claim 16, characterized in that, The processing unit is further configured to: Obtain a word vector corresponding to each character in the multiple characters constituting the fusion data; Generate a second word vector matrix according to the multiple word vectors corresponding to the multiple characters constituting the fusion data; Input the second word vector matrix into the encoder, and obtain the semantic feature vector output by the encoder as the second semantic feature data corresponding to the fusion data.

18. The device according to claim 13 or 14, characterized in that, The text data is obtained by converting the input voice, and the input voice is the vehicle control voice; the processing unit is further configured to: Control the vehicle to perform a target function according to the semantic intention and the entity.

19. The device according to claim 13 or 14, characterized in that, The encoder is trained according to a first training sample and a second training sample, the semantic intention decoder is trained according to the first training sample, the semantic entity decoder is trained according to the second training sample, the first training sample includes sample text data and the intention category corresponding to the pre-annotated sample text data, the second training sample includes sample fusion data and the entity corresponding to the pre-annotated sample fusion data, and the sample fusion data includes the sample text data and the intention category corresponding to the sample text data.

20. The device according to claim 19, wherein The processing unit is further configured to train the encoder, the semantic intention decoder, and the semantic entity decoder through the following steps: Obtain third semantic feature data corresponding to the sample text data through an initial encoder, and predict the semantic intention corresponding to the third semantic feature data through an initial semantic intention decoder; Adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data, so as to train the initial encoder and the initial semantic intention decoder to obtain a first encoder and the semantic intention decoder; Obtain fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predict the entity corresponding to the fourth semantic feature data based on an initial semantic entity decoder; Adjust the weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder to obtain the encoder and the semantic entity decoder.

21. A semantic understanding device, characterized in that, The device includes: A transceiver unit, configured to obtain a first training sample and a second training sample, wherein the first training sample includes sample text data and the intention category corresponding to the sample text data, the second training sample includes sample fusion data and the entity corresponding to the pre-annotated sample fusion data, and the sample fusion data includes the sample text data and the intention category corresponding to the pre-annotated sample text data; A processing unit, configured to obtain third semantic feature data corresponding to the sample text data through an initial encoder, and predict the semantic intention corresponding to the third semantic feature data through an initial semantic intention decoder; The processing unit is further configured to adjust the weight parameters of the initial encoder and the initial semantic intention decoder according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data, so as to train the initial encoder and the initial semantic intention decoder to obtain a first encoder and a semantic intention decoder; The processing unit is further configured to obtain fourth semantic feature data corresponding to the sample fusion data through the first encoder, and predict an entity corresponding to the fourth semantic feature data based on an initial semantic entity decoder; The processing unit is further configured to adjust weight parameters of the first encoder and the initial semantic entity decoder according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data, so as to train the first encoder and the initial semantic entity decoder to obtain an encoder and a semantic entity decoder.

22. The device according to claim 21, characterized in that, Specifically, the processing unit is configured to: Determine a first loss according to the predicted semantic intention and the intention category corresponding to the pre-annotated sample text data; Adjust weight parameters of the initial encoder and the initial semantic intention decoder according to the first loss.

23. The device according to claim 21 or 22, characterized in that, Specifically, the processing unit is configured to: Determine a second loss according to the predicted entity and the entity corresponding to the pre-annotated sample fusion data; Adjust weight parameters of the first encoder and the initial semantic entity decoder according to the second loss.

24. The apparatus according to claim 21 or 22, wherein The transceiver unit is further configured to obtain text data; The processing unit is further configured to obtain first semantic feature data corresponding to the text data through an encoder, and determine a semantic intention corresponding to the first semantic feature data through a semantic intention decoder, so as to implement intention recognition of the text data; The processing unit is further configured to generate fusion data according to the text data and the semantic intention, and obtain second semantic feature data corresponding to the fusion data based on the encoder; The processing unit is further configured to determine an entity corresponding to the second semantic feature data based on a semantic entity decoder, so as to implement entity recognition associated with the semantic intention in the text data.

25. A terminal device, characterized in that, The terminal device includes a processor, a transceiver, and a memory; The processor and the transceiver are used to be coupled to the memory, read and run instructions in the memory, so as to implement the method according to any one of claims 1-8, or the method according to any one of claims 9-12.

26. A computer program product comprising instructions, characterized in that, When the computer program product runs on a terminal, the terminal is caused to execute the method according to any one of claims 1-8, or the method according to any one of claims 9-12.

27. A computer-readable storage medium, characterized in that, Program instructions are stored in the computer-readable storage medium, and when the program instructions run, the method according to any one of claims 1-8 is caused to be executed, or the method according to any one of claims 9-12 is caused to be executed.

Citation Information

Patent Citations

  • Voice processing method and device, computer storage medium and electronic equipment

    CN110890097A