A cross-lingual natural language understanding method based on global-local contrastive learning

Through the global-local comparison learning method, the sentence representations of different languages ​​are explicitly aligned and fine-grained alignment information is learned, and the problem of opacity of the alignment mechanism and mismatch of the sub-task level in zero-sample learning is solved, achieving more efficient cross-language understanding.

CN116227498BActive Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211571399.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-05-13
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing zero-sample learning method has problems such as opacity in the alignment mechanism and mismatch in the fine-grained sub-task hierarchy in cross-language natural language understanding, resulting in the impact of model prediction performance.

Method used

The global-local contrast learning method is adopted, and the similar sentence representations of different languages ​​are explicitly aligned, and the fine-grained alignment information at different levels is learned using the local contrast learning module, and a global semantic-level intent-slot comparison learning module is constructed to realize the interactive channel of intent and slots.

Benefits of technology

It realizes transparent analysis of alignment, learns rich semantic features, reduces the prediction differences between the original language and the target language, and improves the performance of cross-language understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227498B_ABST
    Figure CN116227498B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-language natural language understanding method based on global-local contrastive learning, and belongs to the technical field of natural language processing. Aiming at the high-performance cross-language migration requirements of natural language understanding models, the method studies a cross-language natural language understanding method based on a global-local contrastive learning network, which mainly includes three modules: a local sentence-level intention contrastive learning module, which realizes cross-language sentence representation alignment for intention detection tasks; a local character-level slot contrastive learning module, which realizes cross-language character representation alignment for slot filling tasks; and a semantic-level global intention-slot contrastive learning module, which realizes representation alignment between intentions and slots. The present invention can learn fine-grained alignment information at different levels, dig out rich semantic features, and narrow the prediction difference between the original language and the target language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and relates to a cross-language natural language understanding method based on global-local contrastive learning. Background Art

[0002] At present, language is still the primary carrier of information exchange among people, and it is the most effective and convenient way. As the most natural and direct way of interaction in human-computer communication, voice interaction has natural advantages. As one of the key technologies, natural language understanding usually includes two subtasks: intent detection and slot filling. In order to make natural language understanding models better applicable to low-resource languages ​​that lack a large amount of labeled data, many studies have focused on building networks using zero-shot learning. The zero-shot learning method can use labeled data in high-resource languages ​​to train models and transfer them to target low-resource languages ​​for application.

[0003] Although zero-shot learning can greatly reduce the workload of manually labeling data and has achieved good results in the field, this method only relies on shared parameters and can only perform implicit cross-language alignment. This mechanism brings two problems. First, this implicit alignment process is still a black box at present, which not only seriously affects the alignment representation, but also makes it difficult to analyze its alignment mechanism; second, many research works do not fully consider the different fine-grained levels of the two subtasks. For example, intent detection is at the sentence level, while slot filling is at the character level, which will result in the inability of intent and slot to receive some migration information from different granularity levels, affecting the prediction performance of the model. Therefore, the current work is to make up for the defects of existing natural language understanding models based on zero-shot learning in terms of alignment mechanism and subtask interaction. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a cross-language natural language understanding method based on global-local contrastive learning, which explicitly aligns similar sentence representations in different languages ​​through a contrastive learning method, uses local contrastive learning to learn fine-grained alignment information at different levels, realizes global-local information fusion, and completes cross-language understanding.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A cross-language natural language understanding method based on global-local contrastive learning, the method comprising the following steps:

[0007] S1. Generate an original speech sequence, translate the original speech sequence into a positive sample according to a cross-language dictionary, and input the positive sample into a cross-language pre-training model to obtain the corresponding encoding representation;

[0008] S2. Generate a negative sample queue based on the encoded original speech sequence, positive samples, and negative samples at the previous moment, and input the negative sample queue into the cross-language pre-training model to obtain the corresponding encoding representation;

[0009] S3, build a local sentence-level intention comparison learning module to achieve cross-language sentence representation alignment;

[0010] S4, build a local character-level slot comparison learning module to achieve cross-language character representation alignment;

[0011] S5. Build a global semantic-level intent-slot comparison learning module to align the representation of intent and slot and complete cross-language understanding.

[0012] Further, in step S1, for each character in the original speech sequence, a corresponding translation character is randomly selected in the cross-language dictionary for replacement to generate a positive sample;

[0013] The positive sample is input into the pre-trained model, and the hidden layer state representation h is generated through the bidirectional recurrent neural network. i =BiLSTM(θ emb (x i ), h i-1 ,h i+1 ), where θ emb Represents the vectorized function, and finally the encoding representation for the positive sample can be obtained as:

[0014]

[0015] In the formula, Respectively represent the vector representation of the start flag of the positive sample and the vector representation of the end flag, Represents the vector representation formed after each character in the positive sample is encoded.

[0016] Further, in step S2, a method similar to that in step S1 is used to input the negative sample into the pre-trained model, and the encoding obtained is represented as:

[0017]

[0018] In the formula, K represents the maximum capacity of the negative sample queue, The vector representation of the start flag of the negative sample queue, Represents the vector representation formed after each character in the negative sample queue is encoded.

[0019] Further, in step S3, a local sentence-level intention contrastive learning module is constructed by designing a loss function, and the loss function is as follows:

[0020]

[0021] In the formula, s([], []) represents the dot product operation, h CLS The vector representation of the start flag of the original speech sequence, The vector representation of the start flag in the negative sample queue, K represents the maximum capacity of the negative sample queue.

[0022] Further, in step S4, a local character-level slot contrast learning module is constructed by designing a loss function, and the loss function is as follows:

[0023]

[0024] In the formula, n represents the length of the original speech sequence, h i represents the vector representation of the character at position i in the original utterance sequence, represents the vector representation of the character at position j in the positive sample, Represents the vector representation of the character at position j in the negative sample queue.

[0025] Furthermore, in the local character-level slot contrastive learning module, the final loss is the sum of all character loss functions.

[0026] Further, in step S5, a global semantic level intent-slot contrast learning module is constructed by designing a loss function, and the loss function is designed as follows:

[0027] First, we construct the loss functions for the original speech sequence and the positive sample sequence respectively:

[0028]

[0029]

[0030] In the formula, represents the loss function for the original speech sequence, represents the loss function for positive samples, h j represents the vector representation of the character at position j in the original utterance sequence.

[0031] Combining loss functions and Get the loss function for semantic-level contrastive learning

[0032]

[0033] Then construct the loss functions for intent detection respectively And the loss function for slot filling

[0034]

[0035]

[0036] In the formula, σ Y =softmax(W Y h CLS +b Y ), where W Y and b Y represents the trainable parameters, Indicates the intent tag, represents the predicted value of the slot label for the character at position i, Represents the actual value of the slot label for the character at position i, n Y Indicates the number of intent labels, n C Indicates the number of slot labels.

[0037] Finally, combined with the loss function and The overall loss function of global semantic-level intent-slot contrast learning is obtained:

[0038]

[0039] Where λ represents the trainable parameters, which are all scalars.

[0040] The beneficial effects of the present invention are as follows: the present invention explicitly aligns similar sentence representations in different languages ​​through a contrastive learning method, and supports analysis of the alignment method; a local contrastive learning module is used to learn fine-grained alignment information at different levels in intents and slots, and a global contrastive learning module is used to build an interactive channel between intents and slots to mine richer semantic features, thereby achieving global-local information fusion, completing cross-language understanding, and narrowing the prediction difference between the original language and the target language.

[0041] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0043] Figure 1 This is the model framework of the present invention. DETAILED DESCRIPTION

[0044] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0045] The present invention is proposed for the high-performance cross-language migration requirements of natural language understanding models, and mainly includes three modules: a local sentence-level intention contrast learning module, a local character-level slot contrast learning module, and a semantic-level global intention-slot contrast learning module. Its framework model is as Figure 1 shown. Specifically, the main functions of each module are: the local sentence-level intention contrast learning module realizes cross-language sentence representation alignment for the intention detection task; the local character-level slot contrast learning module realizes cross-language character representation alignment for the slot filling task; the global semantic-level intention-slot contrast learning module realizes the alignment representation between the intention and the slot.

[0046] The implementation steps of the present invention are as follows:

[0047] 1. Generate positive and negative samples. For contrast learning, the key operation is to select appropriate positive and negative sample pairs for the original (anchor) utterance.

[0048] 1) Generation of positive samples based on a cross-language dictionary

[0049] Compared with the original utterance, the positive sample should retain the same semantics. For the original utterance sequence of length n:

[0050] u = ([CLS], u1, u2,..., [SEP])

[0051] where [CLS] and [SEP] both belong to flag bits. The former is placed at the beginning of the entire sequence, and the latter is used to separate non-continuous sequences. A cross-language dictionary is constructed, and the positive sample u is generated under the action of a character conversion generator + . Specifically, for each ui in the sequence u + , a corresponding translated character is randomly selected in the cross-language dictionary for replacement to generate a positive sample sequence. For example, for the original utterance "watch a comedy movie" in Chinese, a positive sample sequence containing English, Japanese, and Korean, such as "watch (看), コメディ (喜剧), (电影)", can be regarded as a cross-language view with the same meaning. Regarding u +Input into the cross-language pre-training model (mBERT), and generate the hidden state representation h through the bidirectional recurrent neural network i =BiLSTM(θ emb (x i ), h i-1 ,h i+1 ), where θ emb Represents the vectorized function, and finally the encoding representation for the positive sample can be obtained as:

[0052]

[0053] In the formula, are the vector representation of the start flag of the positive sample and the vector representation of the end flag, respectively. It is the vector representation formed after each character in the positive sample is encoded.

[0054] 2) Negative sample generation based on queue mechanism

[0055] Considering that the traditional method of generating negative samples is not efficient (for example, selecting other characters in the current batch), a negative sample queue is maintained, which contains the encoded original speech sequence u, the positive sample sequence u + and the negative sample sequence u at the previous moment - , which can gradually reuse the samples of the previous batch to reduce unnecessary encoding processes. Similar to the generation method of positive sample encoding representation, the negative sample queue and word representation for [CLS] are:

[0056]

[0057] Where K represents the maximum capacity of the negative sample queue, is the vector representation of the start flag of the negative sample queue, It is the vector representation formed after each character in the negative sample queue is encoded.

[0058] 2. Build local modules.

[0059] 1) Local sentence-level intention contrast learning module

[0060] Considering that intent detection is a sentence-level classification task, and cross-language sentence representation alignment is the goal of the zero-shot cross-language intent detection task, a dedicated loss function is designed to drive the training model to align similar sentence representations to the same local space across languages ​​for intent detection. The loss function is as follows:

[0061]

[0062] In the formula, s([], []) represents the dot product operation, h CLsis the vector representation of the start flag of the original speech sequence.

[0063] 2) Local character-level slot comparison learning module

[0064] Considering that slot filling is a character-level annotation task, a dedicated loss function is also designed to drive the training model to complete the character alignment for slot filling and realize the cross-language transfer of fine-grained information. For the character at position i, the loss function is:

[0065]

[0066] In the formula, n represents the length of the original speech sequence, h i is the vector representation of the character at position i in the original speech sequence, is the vector representation of the character at position j in the positive sample, is the vector representation of the character at position j in the negative sample queue.

[0067] The final loss is the sum of all character loss functions.

[0068] 3. Build a global semantic level intention-slot comparison learning module

[0069] When slots and intents belong to the same user query, they are usually highly semantically related, so the intent expressed by an utterance and the slots it contains can naturally form a positive sample pair, while the corresponding slots in other utterances can form a negative sample pair. In this regard, we further designed a loss function for global semantic-level intent-slot comparison learning to simulate the semantic interaction between intents and slots, and further improve their cross-language transmission performance.

[0070] First, we construct the loss functions for the original speech sequence and the positive sample sequence respectively:

[0071]

[0072]

[0073] In the formula, represents the loss function for the original speech sequence, represents the loss function for the positive sample sequence, h j is the vector representation of the character at position j in the original utterance sequence.

[0074] Combining loss functions and Get the loss function for semantic-level contrastive learning

[0075]

[0076] Then construct the loss functions for intent detection respectively And the loss function for slot filling

[0077]

[0078]

[0079] In the formula, σ Y =softmax(W Y h CLS +b Y ), where W Y and b Y represents the trainable parameters, Indicates the intent tag, represents the slot label for the character at position i, Represents the actual value of the slot label for the character at position i, n Y Indicates the number of intent labels, n C Indicates the number of slot labels;

[0080] Finally, combined with the loss function and The overall loss function for global semantic-level intent-slot contrastive learning is formed by tuning linear combinations:

[0081]

[0082] Where λ represents a trainable parameter, which is a scalar.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A cross-language natural language understanding method based on global-local contrastive learning, characterized by: The method comprises the following steps: S1. Generate an original speech sequence, translate the original speech sequence into a positive sample according to a cross-language dictionary, and input the positive sample into a cross-language pre-training model to obtain the corresponding encoding representation; S2. Generate a negative sample queue based on the encoded original speech sequence, positive samples, and negative samples at the previous moment, and input the negative sample queue into the cross-language pre-training model to obtain the corresponding encoding representation; S3. Construct a local sentence-level intent contrast learning module by establishing a loss function to achieve cross-language sentence representation alignment. S4. Construct a local character-level slot comparison learning module by establishing a loss function to achieve cross-language character representation alignment; S5. Construct a global semantic-level intent-slot comparison learning module by establishing a loss function to align the representation of intent and slot and complete cross-language understanding. In step S5, a global semantic level intent-slot contrast learning module is constructed by designing a loss function, and the loss function is designed as follows: First, we construct the loss functions for the original speech sequence and the positive sample sequence respectively: In the formula, represents the loss function for the original speech sequence, represents the loss function for positive samples, n represents the length of the original speech sequence, h CLS The vector representation of the start flag of the original speech sequence, h j represents the vector representation of the character at position j in the original utterance sequence, Represents the vector representation of the character at position j in the negative sample queue, represents the vector representation of the character at position j in the positive sample, and K represents the maximum capacity of the negative sample queue; Combining loss functions and Get the loss function for semantic-level contrastive learning Then construct the loss functions for intent detection respectively And the loss function for slot filling In the formula, σ Y =softmax(W Y h CLS +b Y ), where W Y and b Y represents the trainable parameters, Indicates the intent tag, represents the predicted value of the slot label for the character at position i, Represents the actual value of the slot label for the character at position i, n Y Indicates the number of intent labels, n C Indicates the number of slot labels; Finally, combined with the loss function and The overall loss function of global semantic-level intent-slot contrast learning is obtained: Where λ represents a trainable parameter.

2. The cross-language natural language understanding method according to claim 1, characterized in that: In step S1, for each character in the original speech sequence, a corresponding translation character is randomly selected in the cross-language dictionary for replacement to generate a positive sample; The positive sample is input into the pre-trained model, and the hidden layer state representation h is generated through the bidirectional recurrent neural network. i =BiLSTM(θ emb (x i ),h i-1 ,h i+1 ), where θ emb Represents the vectorized function, and finally the encoding representation for the positive sample can be obtained as: In the formula, Respectively represent the vector representation of the start flag of the positive sample and the vector representation of the end flag, Represents the vector representation formed after each character in the positive sample is encoded.

3. The cross-language natural language understanding method according to claim 1, characterized in that: In step S2, the encoding obtained by inputting the negative sample into the pre-trained model is represented as: In the formula, K represents the maximum capacity of the negative sample queue, The vector representation of the start flag of the negative sample queue, Represents the vector representation formed after each character in the negative sample queue is encoded.

4. The cross-language natural language understanding method according to claim 1, characterized in that: In step S3, a local sentence-level intention contrastive learning module is constructed by designing a loss function, and the loss function is as follows: In the formula, s([],[]) represents the dot product operation, h CLS The vector representation of the start flag of the original speech sequence, Represents the vector representation of the start flag in the positive sample, The vector representation of the start flag in the negative sample queue, K represents the maximum capacity of the negative sample queue.

5. The cross-language natural language understanding method according to claim 1, characterized in that: In step S4, a local character-level slot contrast learning module is constructed by designing a loss function, and the loss function is as follows: In the formula, s([],[]) represents the dot product operation, n represents the length of the original speech sequence, and h i represents the vector representation of the character at position i in the original utterance sequence, represents the vector representation of the character at position j in the positive sample, represents the vector representation of the character at position j in the negative sample queue, and K represents the maximum capacity of the negative sample queue.

6. The cross-language natural language understanding method according to claim 5, characterized in that: In the local character-level slot contrastive learning module, the final loss is the sum of all character loss functions.

Citation Information

Patent Citations

  • Chinese-Vietnamese parallel sentence pair extraction method based on cross-language bilingual pre-training and Bi-LSTM

    CN112287695A

  • Minimization of computational demands in model agnostic cross-lingual transfer with neural task representations as weak supervision

    CN112840344A