Server and Multilingual Text Semantic Understanding Method
By training on the server and fine-tuning the multilingual text semantic understanding model using meta-learning, the problem of multilingual voice command intent analysis is solved, and a complete process of unified understanding of multilingual voice commands and adapting downstream tasks is realized.
Patent Information
- Application Number
- CN202210263106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-17
AI Technical Summary
The existing technology cannot effectively understand the semantics of multilingual languages, which leads to high cost of multilingual multi-model development and difficult maintenance, especially in the process of training small language models due to lack of data.
By training multilingual text semantics to understand the model on the server and fine-tuning the model using meta-learning methods, the received voice commands are recognized as text data and input into the model for analysis, obtaining user intentions and key slot information, and after error correction and unified processing, standard entity information is finally generated for parameter encapsulation, and sent to the intelligent terminal to respond to voice commands.
The intention analysis of multilingual voice commands is realized, and the problem of missing corpus in small languages is solved. Through a model, a unified understanding of voice commands containing multilinguals is achieved, and the complete process from text input to parameter encapsulation output is realized.
Smart Images

Figure CN114706944B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to a server and a multi - language text semantic understanding method. Background Art
[0002] With the continuous development of human - machine interaction, speech is being reshaped into a new paradigm of human - machine interaction. The operation of humans on machines has evolved from physical handle buttons, to physical keyboards and mice, and then to touchscreens, and now speech has become an important interaction method. More and more people search for weather and maps through voice, and control smart home appliances such as lights and TVs by issuing voice commands. Along with the wave of economic globalization, smart home appliances with intelligent voice services as the main selling point are sold in multiple countries and regions at the same time. Therefore, the voice interaction function therein faces challenges of multiple languages.
[0003] For semantic understanding of multiple languages, usually, single - language models for various languages are trained and developed separately, and finally, the results output by multiple single - language models are fused. However, the models corresponding to different languages need to be optimized separately and error problems need to be processed, making the development cost of multiple languages and multiple models high and the maintenance difficult. In addition, when training the models corresponding to minority languages, due to the small number of users and scarce language materials, the lack of data makes the training process of the models corresponding to minority languages difficult. Therefore, at present, it is impossible to uniformly perform semantic understanding on multiple languages for different downstream tasks. Summary of the Invention
[0004] This application provides a server and a multi - language text semantic understanding method to solve the technical problem that it is impossible to effectively perform unified semantic understanding on multiple languages in the prior art.
[0005] In a first aspect, this application provides a server, which is configured to:
[0006] Recognize the received voice command as text data;
[0007] Input the text data into a trained multi - language text semantic understanding model to obtain the user intention and key slot values represented by the text data, wherein the training process of the multi - language text semantic understanding model is fine - tuned using a meta - learning training method;
[0008] Parse the key slot values into key - value pair forms to obtain initial entity information, and correct the initial entity information to obtain final entity information;
[0009] Unify the final entity information to obtain standard entity information;
[0010] Perform parameter encapsulation based on the user intention and the standard entity information to obtain encapsulated information, and send the encapsulated information to the intelligent terminal so that the intelligent terminal responds to the voice command according to the encapsulated information.
[0011] In a second aspect, the present application provides a multi - language text semantic understanding method, and the method includes:
[0012] Recognize the received voice command as text data;
[0013] Input the text data into a trained multi - language text semantic understanding model to obtain the user intention and key slot values represented by the text data. Among them, the training process of the multi - language text semantic understanding model is fine - tuned using a meta - learning training method;
[0014] Parse the key slot values into key - value pair form to obtain initial entity information, and correct the initial entity information to obtain final entity information;
[0015] Unify the final entity information to obtain standard entity information;
[0016] Perform parameter encapsulation according to the user intention and the standard entity information to obtain encapsulated information, and send the encapsulated information to the intelligent terminal so that the intelligent terminal responds to the voice command according to the encapsulated information.
[0017] Compared with the prior art, the beneficial effects of the present application are:
[0018] The present application provides a server and a multi - language text semantic understanding method. The server trains a multi - language text semantic understanding model based on the collected feature data containing multiple languages, and fine - tunes the multi - language text semantic understanding model using a meta - learning training method. When the user outputs a voice command, the server converts the received voice command of the user into text data, and inputs the text data into the multi - language text semantic understanding model, so that the model analyzes the text data to obtain the user intention and key slot information represented by the text data. The server stores the initial entity information by parsing the key slot value into a key - value pair form, and corrects the content and unifies the form of the initial entity information to obtain standard entity information. Finally, the server can obtain the corresponding encapsulated information according to the user intention and the standard entity information, and feedback it to the intelligent terminal, so that the intelligent terminal responds to the voice command input by the user. In the present application, the meta - learning is used to fine - tune the multi - language text semantic understanding model to enhance the effect of few - shot learning. An intention analysis is performed on the voice command containing multiple languages through one model, and the problem of lack of small - language corpus is solved, that is, the intention analysis of the multi - language voice command including small languages is realized. In addition, in order to make the multi - language semantic understanding adapt to downstream tasks, the obtained entity information is unified, and a unified and complete process from text input to parameter encapsulation output of multi - language semantic understanding is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 FIG. schematically shows the system architecture of a speech recognition method and a speech recognition device according to some embodiments;
[0021] Figure 2 FIG. schematically shows the hardware configuration block diagram of the intelligent device 200 according to some embodiments;
[0022] Figure 3 FIG. schematically shows the configuration diagram of the intelligent device 200 according to some embodiments;
[0023] Figure 4 FIG. schematically shows a schematic diagram of a voice interaction network architecture according to some embodiments;
[0024] Figure 5 FIG. schematically shows a flowchart of the fine - tuning process of a multi - language text semantic understanding model according to some embodiments;
[0025] Figure 6Exemplarily shown is a schematic diagram of the distribution of data volumes in various languages according to some embodiments;
[0026] Figure 7 Exemplarily shown is a schematic diagram of the division of meta - learning task data according to some embodiments;
[0027] Figure 8 Exemplarily shown is a schematic flowchart of a multi - language text semantic understanding method according to some embodiments;
[0028] Figure 9 Exemplarily shown is an application scenario diagram of a user sending a voice command according to some embodiments;
[0029] Figure 10 Exemplarily shown is a schematic structural diagram of a multi - language text semantic understanding model according to some embodiments;
[0030] Figure 11 Exemplarily shown is a display schematic diagram of an intelligent device 200 responding to a voice command according to some embodiments. Detailed implementation manners
[0031] To make the purpose and implementation manners of this application clearer, the following will clearly and completely describe the exemplary implementation manners of this application with reference to the accompanying drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only a part of the embodiments of this application, rather than all of the embodiments.
[0032] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of this application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.
[0033] The terms "first", "second", "third", etc. in the description, claims and the above - mentioned accompanying drawings of this application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0034] Figure 1 Shows an exemplary system architecture to which the speech recognition method and speech recognition device of this application can be applied. As Figure 1 shown, where 10 is a server, 200 is a terminal device, and is exemplarily included (smart TV 200a, mobile device 200b, smart speaker 200c).
[0035] In this application, the server 10 and the intelligent device 200 perform data communication through various communication methods. The intelligent device 200 is allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 10 can provide various contents and interactions to the terminal device 20. Exemplarily, the intelligent device 200 and the server 10 can send and receive information, and receive software program updates.
[0036] The server 10 can be a server that provides various services. For example, it can be a background server that supports the audio data collected by the intelligent device 200. The background server can analyze and process the received audio and other data, and feedback the processing results (such as endpoint information) to the terminal device. The server 10 can be a server cluster or multiple server clusters, and can include one or more types of servers.
[0037] The intelligent device 200 can be hardware or software. When the intelligent device 200 is hardware, it can be various electronic devices with a sound collection function, including but not limited to smart speakers, smartphones, TVs, tablets, e-book readers, smart watches, players, computers, AI devices, robots, smart vehicles, and so on. When the intelligent devices 200, 201, 202 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as those used to provide sound collection services), or as a single software or software module. No specific limitation is made here.
[0038] It should be noted that the multi-language text semantic understanding method provided by the embodiments of this application can be executed by the server 10, or by the terminal device 20, or jointly executed by the server 10 and the terminal device 20. This application does not make a limitation on this.
[0039] Figure 2 The hardware configuration block diagram of the intelligent device 200 in accordance with an exemplary embodiment is shown. As Figure 2 shown, the intelligent device 200 includes at least one of a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, an audio processor, a graphics processor, RAM, ROM, and first to n interfaces for input / output.
[0040] The display 260 includes a display screen component for presenting a picture, and a driving component for driving image display, a component for receiving an image signal output from the controller and displaying video content, image content, and a menu control interface, as well as a user control UI interface.
[0041] The display 260 can be a liquid crystal display, an OLED display, and a projection display, and can also be a projection device and a projection screen.
[0042] The communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example, the communicator can include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module and other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The smart device 200 can establish the sending and receiving of control signals and data signals with the server 10 through the communicator 220.
[0043] The user interface can be used to receive external control signals.
[0044] The detector 230 is used to collect signals from the external environment or for external interactions. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0045] The sound collector can be a microphone, also known as a "microphone" or "transmitter", which can be used to receive the user's voice and convert the voice signal into an electrical signal. The smart device 200 can be provided with at least one microphone. In some other embodiments, the smart device 200 can be provided with two microphones, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the smart device 200 can also be provided with three, four or more microphones to collect sound signals, reduce noise, and can also identify the sound source to implement functions such as directional recording.
[0046] In addition, the microphone can be built into the smart device 200, or the microphone is connected to the smart device 200 in a wired or wireless manner. Of course, the embodiments of the present application do not limit the position of the microphone on the smart device 200. Or, the smart device 200 may not include a microphone, that is, the above microphone is not provided in the smart device 200. The smart device 200 can externally connect a microphone (which can also be called a microphone) through an interface (such as the USB interface 130). The externally connected microphone can be fixed on the smart device 200 through an external fixing member (such as a camera bracket with a clip).
[0047] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the smart device 200.
[0048] Exemplarily, the controller includes at least one of a Central Processing Unit (CPU), an audio processor, a Graphics Processing Unit (GPU), a Random Access Memory (RAM), a Read-Only Memory (ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.
[0049] In some examples, taking the Android system as an example of the operating system of the smart device, as Figure 3 shown, the smart TV 200-1 can be logically divided into an Applications layer (abbreviation "application layer") 21, a kernel layer 22, and a hardware layer 23.
[0050] Among them, as Figure 3 shown, the hardware layer may include Figure 2 shown, a controller 250, a communicator 220, a detector 230, etc. The application layer 21 includes one or more applications. The application can be a system application or a third-party application. For example, the application layer 21 includes a voice recognition application, and the voice recognition application can provide a voice interaction interface and service for realizing the connection between the smart TV 200-1 and the server 10.
[0051] The kernel layer 22, as a software middleware between the hardware layer and the application layer 21, is used to manage and control hardware and software resources.
[0052] In some examples, the kernel layer 22 includes a detector driver, and the detector driver is used to send the voice data collected by the detector 230 to the voice recognition application. Exemplarily, when the voice recognition application in the smart device 200 is started and the smart device 200 has established a communication connection with the server 10, the detector driver is used to send the voice data of the user input collected by the detector 230 to the voice recognition application. Then, the voice recognition application sends the query information containing the voice data to the intent recognition module 202 in the server. The intent recognition module 202 is used to input the voice data sent by the smart device 200 into the intent recognition model.
[0053] To clearly illustrate the embodiments of the present application, the following combines Figure 4 to describe a voice recognition network architecture provided by the embodiments of the present application.
[0054] Refer to Figure 4 , Figure 4 which is a schematic diagram of a voice interaction network architecture provided by the embodiments of the present application. Figure 4Among them, the intelligent device is used to receive the input information and output the processing result of the information. The speech recognition module deploys a speech recognition service for recognizing audio as text; the semantic understanding module deploys a semantic understanding service for semantic parsing of the text; the service management module deploys a service instruction management service for providing service instructions; the language generation module deploys a language generation service (NLG) for converting the instructions indicating the execution of the intelligent device into text language; the speech synthesis module deploys a text-to-speech (TTS) service for processing the text language corresponding to the instructions and sending it to the speaker for broadcast. In one embodiment, Figure 4 In the architecture shown, there may be multiple entity service devices deploying different business services, or one or more entity service devices may integrate one or more functional services.
[0055] In some embodiments, the following takes the Figure 4 architecture shown as an example to describe the process of processing the information input to the intelligent device. Taking the information input to the intelligent device as a query statement input by voice as an example:
[0056] [Speech Recognition]
[0057] After receiving the query statement input by voice, the intelligent device can perform noise reduction processing and feature extraction on the audio of the query statement. The noise reduction processing here may include steps such as removing echoes and environmental noises.
[0058] [Semantic Understanding]
[0059] Using the acoustic model and language model, perform natural language understanding on the recognized candidate text and related context information, parse the text into structured and machine-readable information, such as information on business domain, intent, slot, etc. to express semantics. Obtain the executable intent and determine the intent confidence score. The semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.
[0060] [Service Management]
[0061] Based on the semantic parsing result of the text of the query statement, the semantic understanding module issues a query instruction to the corresponding service management module to obtain the query result given by the service, and execute the actions required to "complete" the user's final request, and feedback the device execution instruction corresponding to the query result.
[0062] It should be noted that Figure 4 the architecture shown is only an example and not a limitation on the protection scope of this application. In the embodiments of this application, other architectures may also be used to implement similar functions. For example, all or part of the above process may be completed by an intelligent terminal, which will not be elaborated here.
[0063] With the widespread application of intelligent devices and the continuous development of human-computer interaction, controlling corresponding intelligent devices through voice commands has won the favor of a large number of users. Since intelligent devices are sold in multiple countries and regions at the same time, the voice interaction function therein faces the challenge of multiple languages. Currently, usually, separate single-language models for various languages are trained and developed for multilingual semantic understanding. However, the variety of languages increases the maintenance difficulty and slows down the expansion speed. The models corresponding to different languages need to be optimized separately and error problems need to be handled, making the development cost of multilingual multiple models high and the maintenance difficulty large. In addition, when training models corresponding to minority languages, due to the small number of users and scarce language materials, the lack of data makes the training process of models corresponding to minority languages difficult. Therefore, currently, it is impossible to perform unified semantic understanding on multiple languages for different downstream tasks. In view of this, in some embodiments of the present application, a server is provided, and the server is configured to execute a multilingual text semantic understanding process. The multilingual text semantic understanding process will be described below with reference to the accompanying drawings.
[0064] In some embodiments, the multilingual text semantic understanding process executed by the server 10 can be executed in another device. Subsequently, the execution by the server 10 will be taken as an example. Before performing multilingual text semantic understanding, a multilingual text semantic understanding model can be trained first.
[0065] In some embodiments, the multilingual text semantic understanding model is obtained based on LaBSE (Language-agnostic BERT Sentence Embedding, multilingual BERT embedding vector model). The pre-training of LaBSE includes the alignment of word vectors and the alignment of sentence vectors. For the alignment of word vectors, multilingual BERT (pre-trained language representation model) is used to train on the language materials of multiple languages. Among them, in order to map the encodings of different languages to the same space and achieve the alignment of word vectors of multiple languages, MMLM (Multilingual Masked Language Model) and TLM (Translation Language Model) are used for hybrid pre-training to represent the encodings of different languages into the same semantic space. For sentence alignment, a contrastive learning method is used to further train on multilingual parallel language materials according to LaBSE, so that the sentence vectors of different languages are aligned in a unified semantic space. When LaBSE is trained, dual encoders are used to encode the source language and the target language respectively, and the two encoders share parameters and are initialized with the parameters of the BERT model pre-trained by the MMLM and TLM methods.
[0066] In some embodiments, the server 10 builds a multilingual text semantic understanding model based on the pre-trained LaBSE model. For the input text, the LaBSE model first performs encoding. The input is fed into LaBSE, and after multiple layers of transformers, the encoding result is output. The encoding corresponding to [CLS] represents the encoding of the entire sentence, and its result is added with a softmax layer for intent classification. The formula is as follows:
[0067] y i = softmax(W i h [CLS] + b i )
[0068] In the formula, h [CLS] represents the embedding output result corresponding to [CLS], W i corresponds to the weight matrix of the linear layer, b i is the bias vector of the linear layer, and y i is the output vector.
[0069] The encoding output of the corresponding word is added with a softmax layer for slot prediction. The formula is as follows:
[0070] n ∈ 1...N
[0071] In the formula, h [n] represents the embedding output result corresponding to the nth word in the sentence s, W s corresponds to the weight matrix of the linear layer, b s is the bias vector of the linear layer, the output vector corresponding to the nth word in the sentence s.
[0072] In some embodiments, the loss function of the LaBSE model is as follows:
[0073]
[0074] In the formula, represents the loss function, N represents the number of sentences in a batch, x i , y i respectively represent the output vectors of [CLS] of the ith sentences in the first language and the second language, and φ(x i , y i ) is the similarity of the two language sentence vectors, and m is the Margin.
[0075] In some embodiments, when the LaBSE model is fine-tuned, the loss function includes two parts: the cross-entropy loss between the intent classification prediction and the true classification is L I, the cross-entropy loss between the slot prediction and the true slot is L S , the total loss is L = L I +L S , and the total loss is minimized by training and fine-tuning.
[0076] The following describes the fine-tuning process of the LaBSE model with reference to the accompanying drawings.
[0077] Figure 5 FIG. shows a schematic flow diagram of the fine-tuning process of a multilingual text semantic understanding model according to some embodiments. As Figure 5 shown, the fine-tuning process of the multilingual text semantic understanding model is as follows:
[0078] S501: Identify the language of the original data, and obtain feature data according to the original data of the identified language. Among them, the original data includes corpora in multiple languages.
[0079] In some embodiments, in order to achieve multilingual semantic understanding, the server 10 uses corpora in 21 languages during fine-tuning. Since there are significant differences in grammar habits and writing structures among different languages. For example, Chinese, Japanese, and Thai are written in characters, and there are no spaces between characters, while English, French, etc. use words as the smallest unit, and there are spaces between words. If a sentence contains multiple languages at the same time, for example, "Has the Dow Jones index dropped?", there are both Chinese and English in the current sentence. When the server 10 prepares the feature data, it first needs to identify the languages in the original data to facilitate the division of words (characters) before tokenize (token parsing) for different languages.
[0080] In some embodiments, the server 10 uses the regular matching of Unicode encoding, the word list of each language, and the python language detection toolkit langdetect to identify the language in the sentence in turn. The recognition process is as follows: First, the server 10 uses the regular matching of Unicode encoding to detect the original data in units of characters. Since the Unicode encoding of letters in different languages is different, certain languages, such as Chinese, Japanese, Thai and Arabic, can be roughly judged using the Unicode encoding. However, for English and French that share Latin letters, they are languages with the same alphabet system and cannot be distinguished using Unicode encoding. Therefore, after the regular matching detection using Unicode encoding is completed, if there is original data in the test result that does not recognize the language due to the same alphabet system, the server 10 can then use the word list of each language to detect the original data in units of words. Due to incomplete word lists in various languages and the existence of common words in various languages, in a few cases, there are situations where complete distinction cannot be made. At this time, if there is raw data whose language is not identified due to the presence of identical words, the server 10 can use the Python language detection toolkit langdetect for detection, and give the final detection result in combination with the previous detection results.
[0081] In some embodiments, through the above detection process, the server 10 identifies the language in the original data, and then segments the original data of the identified language with the smallest unit of each language (for example, English as a word, Chinese as a character). For example, for "Did the Dow Jones index fall?", it will be segmented into [Dow, Jones, of, index, fell, has it]. After segmentation, the server 10 performs sequence annotation [B_index_name, I_index_name, O, O, O, O, O, O] according to the sequence annotation method of BIO1. Here, the server 10 uses BIO1 to label the words in the text with different tags, annotates "O" for non-critical words, and labels for critical information start with "B-Tag", and the subsequent part is annotated with "I-Tag".
[0082] It should be noted that for languages written from right to left, the text is from right to left and the annotation is from left to right when annotating data, such as Arabic sentences. The annotation for (playing a live football match) is [O, B_MediaType, I_MediaType, O, B_title, I_title].
[0083] Through the above step S501, the characteristic data finally sorted out is shown in Table 1.
[0084] Table 1:
[0085]
[0086] S502: Divide the training process into a training task and an adaptation task. Among them, both the training task and the adaptation task are respectively carried out through training data and test data, and the training data and the test data are selected from the feature data.
[0087] Figure 6 Exemplarily shows a schematic diagram of the data volume distribution of each language according to some embodiments. For regions with a large number of users and a relatively developed Internet, it is more convenient to collect language materials, so the language materials are relatively sufficient, while the language materials in regions with a small number of users are relatively scarce. As Figure 6 shown, in some embodiments, in order to enhance the semantic recognition effect of languages lacking data, the server 10 adopts a meta-learning training method to achieve the transfer of knowledge of common languages to languages lacking data.
[0088] In some embodiments, the server 10 divides the training into two tasks: a training task (train task) and an adaptation task (adapt task). The training task and the adaptation task are respectively assigned training data (suport set) and test data (query set). Among them, the training data and the test data are selected from the feature data, and the feature data includes large-language feature data and small-language feature data.
[0089] In some embodiments, the server 10 realizes the transfer of knowledge between languages through different implementations of dividing the corpus languages of different tasks. Figure 7 Exemplarily shows a schematic diagram of the division of meta-learning task data according to some embodiments. Combining Figure 7 shown, the training data of the training task only includes large-language feature data with sufficient corpus, such as English, Chinese, and French, while the training data of the adaptation task only includes small-language feature data with insufficient corpus, such as Arabic and Malay. The test data for both the training task and the adaptation task also only includes small-language feature data with insufficient corpus.
[0090] S503: Initialize the parameters in the meta-network structure, and perform iterative training based on the training task and the adaptation task to update the parameters in the meta-network structure.
[0091] In some embodiments, initialize the parameter φ in the meta-network structure 0 , and start iterative training. The training process is as follows:
[0092] a. First perform the training task, and assign the parameter φ in the meta-network structure 0 to the network of the training task to obtain θ t; b. Use the training data of the training task with a learning rate α t to iteratively optimize θ t and update θ t ; c. Calculate the loss l using the test data of the training task t , and calculate the gradient of l t with respect to θ t ; d. Update the parameter φ in the meta - network structure using this gradient 0 to obtain φ 1 ; e. Assign φ 1 to the network of the adaptation task to obtain θ a , and then update θ using the training data of the adaptation task a ; f. Calculate the gradient of the loss on the test data of the adaptation task using the updated θ a ; g. Update φ using the gradient obtained in step f 1 ; h. Repeat the process of steps a to g above.
[0093] Above, the server 10's multi - language text semantic understanding model established based on LaBSE takes advantage of LaBSE's superiority in aligning sentence vectors in different languages, which helps the effect of intent recognition. The server 10 uses meta - learning to fine - tune the LaBSE model, enabling the knowledge of languages with rich corpora to be transferred to low - resource languages, which can improve the accuracy of intent recognition and slot filling in low - resource languages.
[0094] The following describes the multi - language text semantic understanding process provided by some embodiments of the present application with reference to the accompanying drawings.
[0095] Figure 8 FIG. [Exemplarily shows a flowchart of a multi - language text semantic understanding method according to some embodiments. As Figure 8 shown, the method includes the following steps:
[0096] S801: Recognize the received voice command as text data.
[0097] Figure 9 FIG. [Exemplarily shows an application scenario diagram of a user sending a voice command according to some embodiments. When a user uses a smart TV, they can send voice commands related to searching for media resources or settings to the smart TV. For example, when the user wants to search for a movie they want to watch, they can send a corresponding voice command to the smart TV. Combining Figure 9, the user sends voice commands in different languages to the smart TV, such as "Search for XXX movie season 3 (Chinese)", or "search for XXX movie (English)" or "Rechercher xxx Films (French)" according to his / her own language habits. In some embodiments, the smart TV can send the voice command to the server 10, and the server 10 converts the user's voice input into text according to ASR. Here, the smart TV can also directly convert the user's voice command into text data through ASR, and then send the text data to the server.
[0098] S802: Input the text data into a trained multilingual text semantic understanding model to obtain the user intent and key slot values represented by the text data, wherein the training process of the multilingual text semantic understanding model is fine-tuned using a meta-learning training method.
[0099] In some embodiments, the server 10 inputs the text data into a multilingual text semantic understanding model, which performs intent recognition and slot filling on the text data, and obtains information representing user intent and key slot values from the text data.
[0100] Figure 10 FIG. 4 is a schematic diagram showing a structure of a multilingual text semantic understanding model according to some embodiments. Figure 10 , input "search for XX'movie XX three" into the model, and use the model to identify the intent and fill the slots. The model predicts that the user intent is the movie and TV search intent, and the label sequence predicted by the model is "OOB I-actor I-actor B-Mediatype B-title I-title B-season".
[0101] The training results of the model are as follows:
[0102] The indicators of the model on the validation set are shown in Table 2.
[0103] Table 2:
[0104]
[0105] For example, the text data "search movie XX" is input into the model, and the parsed result is the user intention "video.search", the media type is "movie", and the movie name is "XX".
[0106] S803: Parse the key slot values into key-value pairs to obtain initial entity information, and correct the errors in the initial entity information to obtain the final entity information.
[0107] In some embodiments, to facilitate the subsequent process to obtain corresponding data more conveniently and accurately, the server 10 parses the result of slot annotation into the form of key-value pairs. Continuing with the above example, it is {title: XX, actor: XX, Mediatype: movie, Season: three}.
[0108] In some embodiments, the server 10 can obtain some entities from the initial entity information by taking keys, such as movie names, person names, etc., for multi-language entity error correction. The server 10 controls the establishment of a knowledge graph related to the user's intention in the entity library of the combined knowledge graph, and corrects the initial entity information according to the knowledge graph related to the user's intention. Establish and maintain in real time a multi-language knowledge graph of movies, music, encyclopedias, etc. for the intention, which is used to provide multi-language knowledge support for the semantic understanding module.
[0109] The entity error correction here mainly targets two aspects of errors. The first aspect is the error of ASR speech recognition. During the ASR recognition process, since the information of the entity has a low relevance to the context, there are many errors in the entity that have not been corrected by speech recognition. For example, in the text "I want to watch the movie XXX by X Jing", the person name "X Jing" in this text should be the director "X Jing" according to the knowledge graph related to the user's intention. The second aspect is the error of slot filling: such as "play xxx audio", the filled result of the movie name may be "xxx audio".
[0110] S804: Unify the final entity information to obtain standard entity information.
[0111] In some embodiments, since the multi-language text semantic understanding model performs unified semantic understanding on each voice command, after parsing the intention category and slot label, the server 10 can perform certain entity unification processing on some final entity information to obtain standard entity information, which is convenient for downstream business processing. Here, the standard entity information refers to standard English entity information. For example, date and time, duration in the case of fast forward and rewind intentions, season period in the video.search (movie search) intention, signal source entity in the control.tv_input_source (control TV input source) intention, etc.
[0112] In some embodiments, if the final entity information representation has date and time, the control uses the date and time regular expression to unify the final entity information. For example, there may be date and time entities under the intents such as weather query, alarm / timer setting, and EPG query. The solution for unifying date and time entities is: based on the open source time parsing library dateparser and customized development according to the scenario requirements of this project, the key task is to define corresponding date and time regular expressions for different languages, and parse the fields of the date_time tag in the slot annotation to meet the parsing of the date and time content in the user's colloquial expression. Here, dateparser supports almost all existing date formats: absolute date, relative date ("two weeks ago" or "tomorrow"), timestamp, etc. It supports more than 200 language regions, supports automatic language detection, supports dates with time zone abbreviations or UTC offsets ("August 14,2015 EST", "21 July 2013 10:15 pm+0500"...), and supports custom parsing formats.
[0113] In some embodiments, if the final entity information represents a duration, the control uses a duration regular expression to unify the final entity information. Duration entity resolution mainly appears in fast forward and rewind intents. The duration resolution module independently developed in the server 10 is used to define duration regular expressions in different languages, and then the duration_time tag field in the slot marking result is parsed. For example, the duration regular expression in simplified Chinese is (\d+)(hours|hours)(\d+)(minutes|minutes)|(\d+)(seconds), and the duration regular expression in English is (\d+)(hou)|(\d+)(minute)|(\d+)(second).
[0114] In some embodiments, if the final entity information represents the existence of a setting entity, a search entity, etc., the control uses the self-defined entity synonyms to unify the final entity information. Since setting entities, search entities, etc. have specific writing methods in smart terminals and are limited in number, for example, the entity reset to factory in the setting intent may be expressed in different languages as restore factory settings (Chinese), reset to factory (English), réinitialisation auxréglages usine (French), and entities in different languages need to be unified into English.
[0115] The final unified results of entity information are shown in Table 3.
[0116] Table 3:
[0117]
[0118] S805: Perform parameter encapsulation according to the user intention and the standard entity information to obtain encapsulated information, and send the encapsulated information to the intelligent terminal, so that the intelligent terminal responds to the voice command according to the encapsulated information.
[0119] In some embodiments, the server 10 sets and selects a preset parameter encapsulation template according to the user intention, and selects corresponding information parameters from the standard entity information and adds them to the parameter encapsulation template to obtain encapsulated information. Figure 11 FIG. shows a display schematic diagram of an intelligent device 200 responding to a voice command according to some embodiments. As Figure 11 shown, after the user sends a voice of "search for the third season of XXX movie" to the smart TV, the encapsulated information fed back to the smart TV via the server 10 is used for corresponding display by the smart TV.
[0120] In this application, a multi-language text semantic understanding model is fine-tuned using meta-learning to enhance the effect of few-shot learning. An intention analysis of voice commands containing multiple languages is performed through a single model, and the problem of lack of small-language corpora is solved, that is, intention analysis of multi-language voice commands including small languages is achieved. Additionally, in order to make the multi-language semantic understanding adapt to downstream tasks, the obtained entity information is unified, realizing a unified and complete process from text input to parameter encapsulation output for multi-language semantic understanding.
[0121] Corresponding to the above server, this application also provides a multi-language text semantic understanding method, which includes: The server 10 recognizes the received voice command as text data. The server 10 inputs the text data into a trained multi-language text semantic understanding model to obtain the user intention and key slot values represented by the text data, where the training process of the multi-language text semantic understanding model is fine-tuned using a meta-learning training method. The server 10 parses the key slot values into key-value pairs to obtain initial entity information, and corrects the initial entity information to obtain final entity information. The server 10 unifies the final entity information to obtain standard entity information. The server 10 performs parameter encapsulation according to the user intention and the standard entity information to obtain encapsulated information, and sends the encapsulated information to the intelligent terminal, so that the intelligent terminal responds to the voice command according to the encapsulated information.
[0122] In some embodiments, the multilingual text semantic understanding model is established based on LaBSE, and LaBSE is fine-tuned using the training method of meta-learning, including: The server 10 identifies the language of the original data, and obtains feature data according to the original data of the identified language, wherein the original data contains corpora in multiple languages. The server 10 divides the training process into training tasks and adaptation tasks, wherein both the training tasks and the adaptation tasks are respectively carried out through training data and test data, and the training data and the test data are selected from the feature data. The server 10 initializes the parameters in the meta-network structure, and performs iterative training based on the training tasks and the adaptation tasks to update the parameters in the meta-network structure.
[0123] Since the above embodiments are all described by reference and combination on the basis of other methods, and there are identical parts among different embodiments, the same or similar parts among the various embodiments in this specification can be referred to each other. Details are not elaborated herein again.
[0124] It should be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such circuit structure, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the circuit structure, article or device including the element.
[0125] Those skilled in the art will readily think of other implementation schemes of this application after considering the specification and practicing the disclosure of this invention. This application is intended to cover any variations, uses or adaptations of this invention, and these variations, uses or adaptations follow the general principles of this application and include common general knowledge or conventional technical means in the technical field not disclosed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this application are pointed out by the content of the claims.
[0126] The above embodiments of this application do not constitute a limitation on the protection scope of this application.
Claims
1. A server, characterized in that, The server is configured to: Recognize the received voice command as text data; Input the text data into a trained multi - language text semantic understanding model to obtain the user intention and key slot values represented by the text data. Among them, the multi - language text semantic understanding model is fine - tuned using a meta - learning training method, and the training method includes: Identify the language of the original data, and obtain feature data according to the original data of the identified language, where the original data contains corpora of multiple languages; Divide the training process of the multi - language text semantic understanding model into a training task and an adaptation task. Both the training task and the adaptation task are respectively carried out through training data and test data, and the training data and test data are selected from the feature data containing large - language feature data and small - language feature data; the training data of the training task is large - language feature data, and the training data of the adaptation task is small - language feature data; the test data of the training task and the test data of the adaptation task are both small - language feature data; Assign the initialization parameters of the meta - network structure to the network of the training task to obtain the first result parameters of the training task; Iteratively optimize the first result parameters through the learning rate to obtain updated first result parameters; Through the test data of the training task, calculate the first gradient of the network of the training task with respect to the result parameters, and update the initialization parameters to target training parameters according to the first gradient; Assign the target training parameters to the network of the adaptation task to obtain the second result parameters of the adaptation task; Update the second result parameters through the training data of the adaptation task to obtain updated second result parameters; Through the test data of the adaptation task, calculate the second gradient of the network of the adaptation task with respect to the second result parameters, and update the target training parameters in the network of the adaptation task according to the second gradient; Based on the training task and the adaptation task, perform iterative training to update the parameters in the meta - network structure; Parse the key slot values into key - value pair form to obtain initial entity information, and correct the initial entity information according to the knowledge graph related to the user intention to obtain the final entity information; Unify the final entity information to obtain standard entity information; Perform parameter encapsulation according to the user intention and the standard entity information to obtain encapsulated information, and send the encapsulated information to the intelligent terminal so that the intelligent terminal responds to the voice command according to the encapsulated information.
2. The server according to claim 1, wherein In the step of unifying the final entity information to obtain standard entity information, the server is configured to: If the final entity information represents the existence of a date and time, control to unify the final entity information using a date - time regular expression; If the final entity information represents the existence of a duration, control to unify the final entity information using a duration regular expression; If the final entity information represents the existence of a set entity, control to unify the final entity information by means of self - defined entity synonyms.
3. The server according to claim 1, wherein In the step of obtaining feature data from the original data according to the identified language, the server is configured to: Segment and label the original data in the identified language by the minimum unit according to the sequence labeling method to obtain the feature data.
4. The server according to claim 1, wherein In the step of identifying the language of the original data, the server is configured to: Detect the original data character by character using regular matching of Unicode encoding; When there is original data in the detection result that has the same alphabet system and whose language has not been identified, detect the original data word by word using the word list of each language; When there is original data in the detection result that has the same word and whose language has not been identified, detect the original data using the Python language detection toolkit.
5. The server according to claim 1, wherein In the step of correcting the initial entity information, the server is configured to: Control the entity library in the joint knowledge graph to establish a knowledge graph related to the user's intention.
6. The server according to claim 1, wherein In the step of performing parameter encapsulation according to the user's intention and the standard entity information to obtain encapsulated information, the server is configured to: Set and select a preset parameter encapsulation template according to the user's intention; Select corresponding information parameters from the standard entity information and add them to the parameter encapsulation template to obtain the encapsulated information.
7. A multi - language text semantic understanding method, characterized in that, The method includes: Recognize the received voice instruction as text data; Input the text data into a trained multi - language text semantic understanding model to obtain the user's intention and key slot values represented by the text data, where the multi - language text semantic understanding model is fine - tuned using the training method of meta - learning, and the training method includes: Identify the language of the original data, and obtain feature data from the original data in the identified language, where the original data contains corpora in multiple languages; Divide the training process of the multi - language text semantic understanding model into a training task and an adaptation task. Both the training task and the adaptation task are respectively carried out through training data and test data. The training data and test data are selected from the feature data including large - language feature data and small - language feature data; the training data of the training task is large - language feature data, the training data of the adaptation task is small - language feature data; the test data of the training task and the test data of the adaptation task are both small - language feature data; Assign the initialization parameters of the meta - network structure to the network of the training task to obtain the first result parameters of the training task; Iteratively optimize the first result parameters through the learning rate to obtain the updated first result parameters; Through the test data of the training task, calculate the first gradient of the network of the training task with respect to the result parameters, and update the initialization parameters to target training parameters according to the first gradient; Assign the target training parameters to the network of the adaptation task to obtain the second result parameters of the adaptation task; Update the second result parameters through the training data of the adaptation task to obtain the updated second result parameters; Calculate a second gradient of the network for the adaptation task with respect to a second result parameter using the test data for the adaptation task, and update target training parameters in the network for the adaptation task according to the second gradient; Perform iterative training based on the training task and the adaptation task to update parameters in the meta-network structure; Parse the key slot values into key-value pairs to obtain initial entity information, and correct the initial entity information according to a knowledge graph related to the user intention to obtain final entity information; Unify the final entity information to obtain standard entity information; Perform parameter encapsulation according to the user intention and the standard entity information to obtain encapsulated information, and send the encapsulated information to the intelligent terminal so that the intelligent terminal responds to the voice command according to the encapsulated information.
Citation Information
Patent Citations
Multi-language model training method and device, storage medium and electronic equipment
CN112749556A
Adversarial sampling training method and device based on meta-learning
CN112786030A
Model training method and device, knowledge extraction method and device, equipment and medium
CN114186533A