A Chinese-Vietnamese cross-language event retrieval method integrating event knowledge

By constructing the Hanyue cross-language event pre-training module and using the comparative learning difference discrimination technology, the problems of poor language alignment and weak event comprehension in the Hanyue cross-language event retrieval are solved, and the performance and event attention of the model in the cross-language event retrieval task are improved.

CN117009458BActive Publication Date: 2025-05-16KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310849936.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2025-05-16
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

The prior art has problems in the Hanyue cross-language event retrieval of low-resource language alignment and weak event comprehension ability of multilingual pre-trained models.

Method used

By constructing the Hanyue cross-language event pre-training module, the model is continuously pre-trained on the Hanyue dataset, and the shared representation of Chinese and Vietnamese in the high-level semantic space of the model is learned. The pre-training module covers the prediction of event knowledge and compares the difference discrimination, so as to improve the model's attention to event knowledge.

Benefits of technology

The performance of the model in cross-language tasks such as Chinese-Free cross-language event retrieval is improved, attention to event elements is enhanced, consistency between target search events and query, and the precise monitoring ability of public opinion events is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117009458B_ABST
    Figure CN117009458B_ABST
Patent Text Reader

Abstract

The present invention relates to a Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge, and belongs to the technical field of natural language processing. The present invention includes: constructing a Chinese-Vietnamese cross-language event pre-training module for continuous pre-training to improve the representation effect of the model on the low-resource languages ​​of Chinese and Vietnamese, and distinguishing the difference between the masked predicted value and the true value of the event knowledge based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge. The Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge proposed by the present invention, compared with information retrieval, cross-language event retrieval enhances the focus on event elements in the query to ensure the consistency between the target retrieval event and the query, and plays an important role in tasks such as accurate monitoring of public opinion events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a Chinese-Vietnamese cross-language event retrieval method integrating event knowledge, and belongs to the technical field of natural language processing. Background Art

[0002] The Chinese-Vietnamese cross-language event retrieval task is to retrieve documents expressing the same event in the target language (such as Vietnamese) based on the input source language (such as Chinese) information. Unlike general information retrieval tasks, cross-language event retrieval not only focuses on the semantic similarity between the source language and the target language, but also requires that the query and the result must express the same event and have the same event elements. Compared with information retrieval, cross-language event retrieval increases the focus on event elements in the query to ensure consistency between the target retrieval event and the query, and plays an important role in tasks such as accurate monitoring of public opinion events.

[0003] The Chinese-Vietnamese cross-language event retrieval of the present invention is a special cross-language information retrieval task. Traditional cross-language retrieval usually has two methods: one is to use machine translation to translate the query or the document to be queried into the same language for retrieval, and convert the cross-language retrieval task into the retrieval task of the same language, but its effect is limited by the performance of machine translation. For resource-rich languages ​​such as Chinese and English, the translation error is small and the effect is good, while for low-resource languages ​​such as Vietnamese, the translation performance is limited, and important entities of events such as names, place names, and organization names may be translated incorrectly, and the retrieval performance is not ideal. The second is to use multilingual word embedding or multilingual pre-training models to represent different languages ​​into the same semantic space, and realize cross-language information retrieval by calculating the semantic similarity between the query and the document. However, there are some multilingual pre-training models at present, and their data distribution between different languages ​​is unbalanced. This imbalance causes these models to show strong representation capabilities when processing resource-rich languages ​​such as Chinese and English, and can achieve good alignment effects. However, the representation effect for low-resource languages ​​such as Chinese-Vietnamese is still unsatisfactory, and directly using multilingual pre-training models to realize cross-language retrieval of Chinese-Vietnamese still faces great challenges. In addition, cross-language event retrieval is different from traditional information retrieval. Traditional information retrieval only considers the similarity between query and document at the semantic level, while event retrieval should consider whether the sentence and query are describing the same event, that is, whether they have the same event subject, trigger words, event time and other element information. For cross-language event retrieval tasks, it is necessary to improve pre-training models such as mBert, and increase the ability to understand events on the basis of existing text representation. The present invention proposes a Chinese-Vietnamese cross-language event retrieval method that incorporates event knowledge. On the one hand, in order to solve the problem of poor alignment effect of Chinese-Vietnamese low-resource languages, a Chinese-Vietnamese cross-language event pre-training module is constructed to continuously pre-train the model on the Chinese-Vietnamese data set, and in this process, the shared representation of Chinese and Vietnamese in the high-level semantic space of the model is learned; on the other hand, in order to solve the problem of weak event understanding ability of multilingual pre-training models, the pre-training module continuously masks the prediction of event knowledge, and distinguishes the masked prediction value from the true value based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge and improve the model's attention to event knowledge. Finally, through the fusion of the two strategies, the performance of the model on cross-language tasks such as Chinese-Vietnamese cross-language event retrieval is improved. Summary of the invention

[0004] The present invention provides a Chinese-Vietnamese cross-language event retrieval method that incorporates event knowledge, so as to solve the problem that the model has poor alignment effect in Chinese-Vietnamese low-resource retrieval, and that simple semantic matching retrieval is difficult to understand the event semantic information of complex queries.

[0005] The technical solution of the present invention is: a Chinese-Vietnamese cross-language event retrieval method integrating event knowledge, and the specific steps of the method are as follows:

[0006] Step 1: Collect corresponding alignment data from Chinese-Vietnamese Wikipedia and Vietnam News Network, preprocess the data set, and construct the experimental data set;

[0007] Step 2: Construct a Chinese-Vietnamese cross-language event pre-training module for continuous pre-training to improve the model's representation effect on the low-resource languages ​​of Chinese and Vietnamese, and distinguish the masked predicted values ​​and true values ​​of event knowledge based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge;

[0008] Step 3. Based on Step 2, fine-tune the model on the Chinese-Vietnamese retrieval dataset to further improve the performance of the model in Chinese-Vietnamese cross-language retrieval and detect Vietnamese documents that are consistent with the events mentioned in the query.

[0009] As a further solution of the present invention, the specific steps of Step 1 are:

[0010] Step 1.1, collect relevant data from Wikipedia and Vietnam News Network, in order to complete the cross-language event retrieval between Chinese and Vietnamese, and finally build an aligned Chinese-Vietnamese language pair;

[0011] Step 1.2: For Wikipedia, extract the document data DocC and DocV with the same month from the chronology of Chinese and Vietnamese Wikipedia pages; obtain the Chinese query document QueC(i) by concatenating the hyperlinks of the Chinese document data DocC, and calculate the similarity between the Chinese query document QueC(i) and the Vietnamese document DocV(j) using cosine similarity on cross-language word embeddings. If the result is greater than the preset threshold β, they are regarded as similar language pairs, thus realizing the alignment of the relevant language pairs of Chinese and Vietnamese.

[0012] Step 1.3. For news network data, select the paragraph header text SentV(i) in the Vietnamese news text ArtV as the document part of the language pair, and generate the Chinese summary SentC as the query part of the language pair through the output of the Chinese-Vietnamese cross-language summary generation model.

[0013] As a further solution of the present invention, the Step 2 comprises:

[0014] Step 2.1, constructing the Chinese-Vietnamese cross-language event pre-training module includes adding two new pre-training task modules to the basic bert model; the two new pre-training task modules include the trigger word mask prediction module TMP and the retrieval modeling ranking module RS. The model uses different input language pairs for different pre-training modules;

[0015] The trigger words marked in the language pair are masked through the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model;

[0016] The retrieval modeling and ranking module RS is used to model the query and recall document process, construct positive and negative language pairs, continuously optimize the model's discrimination of positive and negative examples, and improve the model's ranking accuracy in downstream retrieval tasks.

[0017] Step 2.2. Based on Step 2.1, a generative discriminator module GDC based on contrastive learning is constructed to map the prediction results to the trigger word list. The masked prediction value and the true value of the event knowledge are distinguished based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge.

[0018] As a further solution of the present invention, in Step 2.1, the trigger words marked in the language pair are masked by the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model. The specific steps are as follows:

[0019] Step 2.1.1, for the input sequence x of trigger word mask prediction, the Bert model prediction head linear layer component in the MLM pre-training task of the Bert model is used to predict the masked trigger words and map the results to the full vocabulary label space, and obtain the final full vocabulary label prediction probability distribution through the softmax layer:

[0020] p(x t |x)=Softmax(f δ (h J (x t )));

[0021] Among them, p(x t |x) is the label prediction probability distribution of the entire vocabulary, x t is the masked trigger word of the language pair indexed at position t in the input sequence x, h J (x) is the output of the Bert model prediction head, f δ is a linear layer module that maps the masked word vectors to the full vocabulary label space δ;

[0022] Step 2.1.2, use cross entropy loss to calculate the target of each trigger word mask prediction; during the training process, randomly sample one, two or three trigger words from the matched trigger word list with a random probability of 6:2:2, and mask them in the query text; in the matching and masking process, there is still a probability that the number of masked trigger words is set to 2 or 3;

[0023]

[0024] Among them l tmp is the loss of the TMP module, is the real label of the trigger word in the full vocabulary label space, is the corresponding true probability label distribution of the masked trigger word in the full vocabulary label space.

[0025] As a further solution of the present invention, in Step 2.1, the process of querying and recalling documents is modeled by the retrieval modeling and ranking module RS, positive and negative language pairs are constructed, the model is continuously optimized for distinguishing positive and negative examples, and the specific steps of improving the ranking accuracy of the model in the downstream retrieval task are as follows:

[0026] Step 2.1.3, for the selection of irrelevant documents in the input sequence y of the retrieval modeling ranking module RS, in order to avoid the model falling into a fixed gradient and causing the optimization process to converge to a local optimal solution, random sampling is used to match irrelevant documents DV during the training process; the retrieval modeling ranking module RS encodes the positive and negative examples of the input sequence y respectively, and obtains the encoded positive and negative examples y{QD+}k and y{QD-}k, where k∈batch_size, batch_size represents the number of samples selected for one training;

[0027] Step 2.1.4, in the ranking score generation layer, a learnable parameter weight matrix W is used to multiply the sequence positive and negative example encoding output y{QD+}k and y{QD-}k to obtain the ranking scores S+ and S- respectively: and the model is optimized using cross entropy loss; by continuously optimizing W and minimizing the cross entropy loss function l rs , the model will improve the ranking score S+ for positive examples and continue to reduce the ranking score S- for negative examples, so that the model can make better judgments on positive and negative examples in the retrieval scenario and improve the ranking process of recalled documents; the solution process is:

[0028] S + / - =y{QD + / -} k *W.

[0029] As a further solution of the present invention, the specific steps of Step 2.2 are:

[0030] Step 2.2.1. Use the real language pair as the positive sample. When the trigger word predicted by the model is inconsistent with the real trigger word, map the predicted trigger word from the full vocabulary to the trigger word vocabulary label space η, and use the generated predicted text as the negative sample; generate a discriminator to compare the positive and negative samples to continuously optimize the accuracy of the model's prediction of the trigger word;

[0031]

[0032] Among them l d is the loss of the GDC module, D(x t |x)=Sigmoid(f η (h D (x t ), a is a binary label used to determine whether the model prediction is correct, f is a linear layer module, which maps the masked word vector to the trigger word vocabulary label space η, h D and the trigger word mask prediction module h J It is the shared Bert model prediction head;

[0033] Step 2.2.2, in the process of generating positive and negative examples, there is a p% probability that the trigger words predicted by the model are not used. Instead, the time trigger word prediction task and the event trigger word prediction task are first classified, and then trigger words of the same length are randomly extracted from the trigger word list to replace the predicted trigger words; p is set to 50 to ensure the relative balance of positive and negative samples;

[0034] The loss function of the trigger word mask prediction module TMP and the generative discriminator module GDC based on contrastive learning in the Chinese-Vietnamese cross-language event pre-training module is to mix the cross entropy loss and contrast loss obtained by the trigger word mask prediction module and iterate and optimize the model through the forward propagation layer. At the same time, the loss of the retrieval modeling ranking module adopts the cross entropy loss.

[0035] As a further solution of the present invention, the specific steps of Step 3 are:

[0036] Step 3.1. The question part in the question-answering dataset is used as the query part in the retrieval task. The correct answer corresponding to the question and the multiple predicted answers generated by the question-answering task are used as the relevant documents in the retrieval. The other question answers and the corresponding multiple predicted answers generated by the question-answering task are used as the irrelevant documents corresponding to the query. Finally, an evaluation dataset of queries and relevant and irrelevant documents in different languages ​​is obtained.

[0037] Step 3.2: Based on Step 3.1, the average accuracy and MAP indicators are used to evaluate the performance of the model in the retrieval task:

[0038]

[0039] Where N represents the total number of relevant documents, position(i) represents the position of the i-th relevant document in the search result list, and the average accuracy is the average of the average correct rates AP of multiple queries. As a common evaluation indicator for retrieval tasks, the average accuracy can reflect the retrieval performance of the model as a whole.

[0040] Step 3.3. Based on Step 3.1, the retrieval scenario is optimized through the retrieval modeling and ranking module. Therefore, only the last three encoding layers are fine-tuned during the fine-tuning process to avoid over-matching. The model is optimized using cross entropy loss during the evaluation process, and the maximum number of training times is set to 20 times. When the model obtains the best MAP score on the development set, the program evaluates it on the test set and records the corresponding MAP score, and finally selects the best score to save as the result.

[0041] The beneficial effects of the present invention are:

[0042] 1. Aiming at the problem of poor alignment of Chinese and Vietnamese low-resource languages, the present invention constructs a Chinese-Vietnamese cross-language event pre-training module to continuously pre-train the model on Chinese-Vietnamese datasets, and in this process learns the shared representation of Chinese and Vietnamese in the high-level semantic space of the model;

[0043] 2. The present invention aims to solve the problem that the multilingual pre-trained model has a weak ability to understand events. The pre-trained module continuously masks the prediction of event knowledge, and distinguishes the masked prediction value from the true value based on comparative learning, so as to enable the model to better understand and capture the characteristics of event knowledge and improve the model's attention to event knowledge.

[0044] 3. Compared with information retrieval, cross-language event retrieval in this invention increases the focus on event elements in the query to ensure the consistency between the target retrieval event and the query, and plays an important role in tasks such as accurate monitoring of public opinion events;

[0045] 4. The Chinese-Vietnamese cross-language event retrieval method that incorporates event knowledge proposed in the present invention outperforms the traditional baseline method on the constructed dataset, verifying the effectiveness of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the overall flow chart of the model in the present invention;

[0047] Figure 2 As an example of the present invention;

[0048] Figure 3 This is a model diagram of the present invention. DETAILED DESCRIPTION

[0049] Example 1: Figure 1-Figure 3 As shown, the specific steps of the Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge are as follows:

[0050] Step 1: Collect corresponding alignment data from Chinese-Vietnamese Wikipedia and Vietnam News Network, preprocess the data set, and construct the experimental data set;

[0051] As a further solution of the present invention, the specific steps of Step 1 are:

[0052] Step 1.1, collect and obtain relevant data from Wikipedia and Vietnam News Network, and finally construct 108,868 aligned Chinese-Vietnamese language pairs in order to complete the Chinese-Vietnamese cross-language event retrieval;

[0053] Step 1.2: For Wikipedia, extract the document data DocC and DocV with the same month from the chronology of Chinese and Vietnamese Wikipedia pages; obtain the Chinese query document QueC(i) by concatenating the hyperlinks of the Chinese document data DocC, and calculate the similarity between the Chinese query document QueC(i) and the Vietnamese document DocV(j) using cosine similarity on cross-language word embeddings. If the result is greater than the preset threshold β, they are regarded as similar language pairs, thus realizing the alignment of the relevant language pairs of Chinese and Vietnamese.

[0054] Step 1.3: For the news network data, we select the paragraph header text SentV(i) in the Vietnamese news text ArtV as the document part of the language pair, and generate the Chinese summary SentC as the query part of the language pair through the output of the Chinese-Vietnamese cross-language summary generation model. The final data set statistics are shown in Table 1.

[0055] Table 1 Dataset statistics

[0056]

[0057] Step 2: Construct a Chinese-Vietnamese cross-language event pre-training module for continuous pre-training to improve the model’s representation effect on the Chinese and Vietnamese low-resource languages, and distinguish the masked predicted values ​​and true values ​​of event knowledge based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge; Figure 3 The figure shows an example of this embodiment.

[0058] As a further solution of the present invention, Step 2 includes:

[0059] Step 2.1, constructing the Chinese-Vietnamese cross-language event pre-training module includes adding two new pre-training task modules to the basic bert model; the two new pre-training task modules include the trigger words mask prediction module (Trigger Words Mask Prediction, TMP) and the retrieval modeling sorting module (Retrieve Modeling Sort, RS), and the model uses different input language pairs for different pre-training modules;

[0060] The trigger words marked in the language pair are masked through the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model;

[0061] The retrieval modeling and ranking module RS is used to model the query and recall document process, construct positive and negative language pairs, continuously optimize the model's discrimination of positive and negative examples, and improve the model's ranking accuracy in downstream retrieval tasks.

[0062] Step 2.2. Based on Step 2.1, a generative discriminator for contrastive learning (GDC) module is constructed to map the prediction results to the trigger word list, and the masked prediction value and the true value of the event knowledge are distinguished based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge.

[0063] As a further solution of the present invention, in Step 2.1, the trigger words marked in the language pair are masked by the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model. The specific steps are as follows:

[0064] Step 2.1.1. For the input sequence x of trigger word mask prediction, the Bert model prediction head (BertLMPredictionHead) linear layer component in the MLM pre-training task of the Bert model is used to predict the masked trigger words and map the results to the full vocabulary label space, and obtain the final full vocabulary label prediction probability distribution through the softmax layer:

[0065] p(x t |x)=Softmax(f δ (h J (x t )));

[0066] Among them, p(x t |x) is the label prediction probability distribution of the entire vocabulary, x t is the masked trigger word of the language pair indexed at position t in the input sequence x, h J (x) is the output of the Bert model prediction head, f δ is a linear layer module that maps the masked word vectors to the full vocabulary label space δ;

[0067] Step 2.1.2, use cross entropy loss to calculate the target of each trigger word mask prediction; during the training process, randomly sample one, two or three trigger words from the matched trigger word list with a random probability of 6:2:2, and mask them in the query text; since the query text is usually short, the length ratio of some trigger words may have exceeded 30% of the query text length. Also, considering that some trigger words may be short and there may be three or more trigger words in some query texts, there is still a probability that the number of masked trigger words will be set to 2 or 3 during the matching and masking process;

[0068]

[0069] Among them l tmp is the loss of the TMP module, is the real label of the trigger word in the full vocabulary label space, is the corresponding true probability label distribution of the masked trigger word in the full vocabulary label space.

[0070] As a further solution of the present invention, in Step 2.1, the process of querying and recalling documents is modeled by the retrieval modeling and ranking module RS, positive and negative language pairs are constructed, the model is continuously optimized for distinguishing positive and negative examples, and the specific steps of improving the ranking accuracy of the model in the downstream retrieval task are as follows:

[0071] Step 2.1.3 is similar to the recall ranking stage in the retrieval task, but compared with the fine-grained data of the general retrieval task, the semantics of the task data has a coarser granularity; for the selection of irrelevant documents in the input sequence y of the retrieval modeling ranking module RS, in order to avoid the model falling into a fixed gradient and causing the optimization process to converge to the local optimal solution, random sampling is used to match irrelevant documents DV during the training process; the retrieval modeling ranking module RS encodes the positive and negative examples of the input sequence y respectively, and obtains the encoded positive and negative examples y{QD+}k and y{QD-}k, where k∈batch_size, batch_size represents the number of samples selected for one training;

[0072] Step 2.1.4, in the ranking score generation layer, a learnable parameter weight matrix W is used to multiply the sequence positive and negative example encoding output y{QD+}k and y{QD-}k to obtain the ranking scores S+ and S- respectively: and the model is optimized using cross entropy loss; by continuously optimizing W and minimizing the cross entropy loss function l rs , the model will improve the ranking score S+ for positive examples and continue to reduce the ranking score S- for negative examples, so that the model can make better judgments on positive and negative examples in the retrieval scenario and improve the ranking process of recalled documents; the solution process is:

[0073] S+ / - =y{QD + / -} k *W.

[0074] As a further solution of the present invention, the specific steps of Step 2.2 are:

[0075] Step 2.2.1. Use the real language pair as the positive sample. When the trigger word predicted by the model is inconsistent with the real trigger word, map the predicted trigger word from the full vocabulary to the trigger word vocabulary label space η, and use the generated predicted text as the negative sample; generate a discriminator to compare the positive and negative samples to continuously optimize the accuracy of the model's prediction of the trigger word;

[0076]

[0077] Among them l d is the loss of the GDC module, D(x t |x)=Sigmoid(f η (h D (x t ), a is a binary label used to determine whether the model prediction is correct, f is a linear layer module, which maps the masked word vector to the trigger word vocabulary label space η, h D and the trigger word mask prediction module h J It is the shared Bert model prediction head;

[0078] Step 2.2.2, in the process of generating positive and negative examples, there is a p% probability that the trigger words predicted by the model are not used. Instead, the time trigger word prediction task and the event trigger word prediction task are first classified, and then trigger words of the same length are randomly extracted from the trigger word list to replace the predicted trigger words; p is set to 50 to ensure the relative balance of positive and negative samples;

[0079] The loss function of the trigger word mask prediction module TMP and the generative discriminator module GDC based on contrastive learning in the Chinese-Vietnamese cross-language event pre-training module is to mix the cross entropy loss and contrast loss obtained by the trigger word mask prediction module and iterate and optimize the model through the forward propagation layer. At the same time, the loss of the retrieval modeling ranking module adopts the cross entropy loss.

[0080] Step 3. Based on Step 2, fine-tune the model on the Chinese-Vietnamese retrieval dataset to further improve the performance of the model in Chinese-Vietnamese cross-language retrieval and detect Vietnamese documents that are consistent with the events mentioned in the query.

[0081] As a further solution of the present invention, the specific steps of Step 3 are:

[0082] Step 3.1. The question part in the question-answering dataset is used as the query part in the retrieval task. The correct answer corresponding to the question and the multiple predicted answers generated by the question-answering task are used as the relevant documents in the retrieval. The other question answers and the corresponding multiple predicted answers generated by the question-answering task are used as the irrelevant documents corresponding to the query. Finally, an evaluation dataset of queries and relevant and irrelevant documents in different languages ​​is obtained.

[0083] Step 3.2: Based on Step 3.1, the average accuracy and MAP indicators are used to evaluate the performance of the model in the retrieval task:

[0084]

[0085] Where N represents the total number of relevant documents, position(i) represents the position of the i-th relevant document in the search result list, and the average accuracy is the average of the average correct rates AP of multiple queries. As a common evaluation indicator for retrieval tasks, the average accuracy can reflect the retrieval performance of the model as a whole.

[0086] Step 3.3. Based on Step 3.1, the retrieval scenario is optimized through the retrieval modeling and ranking module. Therefore, only the last three encoding layers are fine-tuned during the fine-tuning process to avoid over-matching. The model is optimized using cross entropy loss during the evaluation process, and the maximum number of training times is set to 20 times. When the model obtains the best MAP score on the development set, the program evaluates it on the test set and records the corresponding MAP score, and finally selects the best score to save as the result.

[0087] The present invention uses the Mean Average Precision (MAP) indicator to evaluate the performance of the model in the retrieval task.

[0088] In order to illustrate the detection effect of the present invention, a baseline method is used to compare with the detection results of the present invention, specifically, compared with the following core event detection method.

[0089] mBert is characterized by its strong cross-lingual capabilities, that is, it can achieve word-level and sentence-level alignment between different languages, so that the model has good performance and generalization capabilities when processing tasks in different languages. Its design allows mBert to be used as a general basic model for transfer learning and fine-tuning in natural language processing tasks in different languages, thereby reducing the additional training cost for specific languages.

[0090] XLM-R is a multilingual pre-trained model based on the RoBERTa architecture. It effectively learns the semantic associations and commonalities between various languages ​​through self-supervised learning on large-scale cross-lingual datasets, enabling it to perform understanding and generation tasks in a multilingual environment.

[0091] LaBSE is a language-independent BERT sentence embedding model. It implements cross-language sentence similarity calculation and text matching tasks by mapping sentences in different languages ​​into a shared semantic space. The LaBSE model learns the semantic alignment and similarity between various languages ​​through self-supervised learning on large-scale cross-lingual data, enabling it to generate language-independent sentence embedding representations.

[0092] mBert, MLM, ZV In order to explore the impact of the idea of ​​improving the attention to event knowledge proposed in this invention on the model performance, we continuously pre-trained mBert on the MLM task on the Chinese-Vietnamese pre-training dataset to explore the impact of different tasks on the model when the training expectations are the same.

[0093] The experimental results of the Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge are shown in Table 2:

[0094] Table 2 Performance of the model on the cross-language event retrieval task - using MAP as the evaluation indicator

[0095]

[0096] The results of the experiment are shown in Table 2: For the baseline model, although XLM-R and mBert have similar model structures and pre-training objectives, and use more training data for pre-training, the performance of the XLM-R model in the three cross-language retrieval tasks is far inferior to that of the mBert model. The present invention speculates that the reason may be due to the way the pre-training data is input into the model. The XLM-R model pre-training task accepts word-level input, which may be suitable for word-level tasks, but for tasks that require alignment to represent long texts, such as cross-language event retrieval, it may cause model confusion and confusion. The LaBSE model performs better than other baseline models in Chinese-English cross-language tasks, but performs poorly in Vietnamese cross-language tasks. This may be caused by the lack of Vietnamese in the pre-training data during the model pre-training process. The mBert model, which was re-pretrained with Chinese-Vietnamese data, achieved the best score among the four baseline models in Chinese-Vietnamese cross-language event retrieval, which shows that the use of Chinese-Vietnamese cross-language texts for pre-training has a certain improvement and promotion in the alignment of Chinese and Vietnamese languages ​​in the high-level semantic space of the model. For the model after using the Chinese-Vietnamese cross-language event retrieval method, the MAP score on the downstream task of Chinese-Vietnamese cross-language event retrieval is improved by 32% compared with the best performing baseline model, which proves that the pre-training strategy and generated discriminator proposed in the present invention have significant improvements for the Chinese-Vietnamese cross-language event retrieval task. At the same time, in zero-shot learning tasks such as Chinese-English, Vietnamese-English, etc. that have not been cross-language trained, the model scores after using the Chinese-Vietnamese cross-language event retrieval method are also higher than other baseline models, which proves that the pre-training method proposed in the present invention also has different degrees of positive effects on the alignment of other languages ​​in high-level semantic space.

[0097] In order to further verify the effectiveness of the method proposed in this invention, an ablation experiment was conducted on the proposed method: two new pre-training modules and a generative discriminator based on contrastive learning, and the experimental results were analyzed. The experimental results are shown in Table 3.

[0098] Table 3 Ablation experiment - using MAP as the evaluation index

[0099]

[0100] The present invention retrained multiple models without a single pre-training strategy module to study the impact of each pre-training module on the model and downstream task performance. As shown in Table 3, the new pre-training strategy has played a positive role in improving the performance of downstream tasks, but the two pre-training modules have different effects on model performance.

[0101] mBert, TMP+GDC is based on the mBert model and uses the trigger word mask prediction module and the contrastive learning-based generative discriminator module for continuous pre-training, removing the retrieval modeling ranking module.

[0102] mBert,RS is based on the mBert model and uses the retrieval modeling ranking module for continuous pre-training, removing the trigger word masking prediction module and the generative discriminator module based on contrastive learning.

[0103] mBert, TMP+RS is based on the mBert model and uses the trigger word mask prediction module and the retrieval modeling ranking module for continuous pre-training, removing the generative discriminator module based on contrastive learning.

[0104] Both the retrieval modeling ranking module and the trigger word masking prediction module provide positive benefits for Chinese-Vietnamese cross-language event retrieval, and the former is much more effective than the latter. Considering that the retrieval modeling ranking module was designed with reference to the scenario of using language models for retrieval tasks, this may be the reason why the retrieval modeling module is more effective in retrieval. In the zero-shot task, the trigger word masking module has a more significant impact on downstream tasks than the retrieval modeling ranking module. Both the trigger word masking task and the retrieval modeling ranking task have a positive impact on downstream cross-language retrieval tasks, and these pre-training tasks complement each other. Because by retraining with two new pre-training tasks, the model performs better than the model retrained with only a single pre-training task in most cases.

[0105] A specific example of the present invention is Figure 2 As shown, the present invention understands and pays attention to the trigger words in the query, and the Vietnamese text queried more accurately understands the query semantics, rather than simply matching the semantics like the bert basic model; the present invention constructs a good event and time trigger word list, corrects the prediction of the trigger words for the pre-training task of trigger word masking prediction and the generation discriminator, and continuously optimizes the accuracy of the model's prediction of trigger words to improve the model's attention to event knowledge. As the pre-training task loss and contrast loss continue to decrease, this shows that the model gradually increases its attention to event knowledge and the prediction of trigger words is gradually accurate. In order to evaluate the impact of the model's attention to event knowledge on the performance of downstream tasks, the present invention sequentially adopts the MLM task to replace the trigger word strategy masking module and the strategy of removing the generation discriminator during training. The retrieval modeling ranking module is introduced by default each time training, and the model is re-pre-trained multiple times. In the Chinese-Vietnamese cross-language event retrieval task, as the training rounds deepen, the model's attention to events increases, and the retrained model performs significantly better in downstream tasks than the cross-language model using the MLM task and the cross-language model without the generation discriminator. This shows that improving the model's focus on event knowledge is of great help in improving the performance of downstream tasks such as Chinese-Vietnamese cross-language event retrieval.

[0106] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge, characterized by: The specific steps of the method are as follows: Step 1: Collect the corresponding aligned data from Chinese-Vietnamese Wikipedia and Vietnam News Network, preprocess the data set, and construct the experimental data set; Step 2: Construct a Chinese-Vietnamese cross-language event pre-training module for continuous pre-training to improve the model's representation effect on the low-resource languages ​​of Chinese and Vietnamese, and distinguish the masked predicted values ​​and true values ​​of event knowledge based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge; Step 3: Based on Step 2, fine-tune the model on the Chinese-Vietnamese retrieval dataset to further improve the performance of the model in Chinese-Vietnamese cross-language retrieval and detect Vietnamese documents that are consistent with the events mentioned in the query; The Step 2 includes: Step 2.1, constructing the Chinese-Vietnamese cross-language event pre-training module includes adding two new pre-training task modules to the basic bert model; the two new pre-training task modules include the trigger word mask prediction module TMP and the retrieval modeling ranking module RS. The model uses different input language pairs for different pre-training modules; The trigger words marked in the language pair are masked through the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model; The retrieval modeling and ranking module RS is used to model the query and recall document process, construct positive and negative language pairs, continuously optimize the model's discrimination of positive and negative examples, and improve the model's ranking accuracy in downstream retrieval tasks. Step 2.2, based on Step 2.1, construct a generative discriminator module GDC based on contrastive learning, map the prediction results to the trigger word list, and distinguish the masked prediction value and the true value of the event knowledge based on contrastive learning, so as to enable the model to better understand and capture the characteristics of event knowledge; In the above Step 2.1, the trigger words marked in the language pair are masked by the trigger word masking prediction module TMP, and the prediction results are mapped to the full word list to continuously correct the prediction accuracy of the model. The specific steps are as follows: Step 2.1.1, for the input sequence x of trigger word mask prediction, the Bert model prediction head linear layer component in the MLM pre-training task of the Bert model is used to predict the masked trigger words and map the results to the full vocabulary label space, and obtain the final full vocabulary label prediction probability distribution through the softmax layer: Among them, p(x t |x) is the label prediction probability distribution of the entire vocabulary, x t is the masked trigger word of the language pair indexed at position t in the input sequence x, h J (x) is the output of the Bert model prediction head, f δ is a linear layer module that maps the masked word vectors to the full vocabulary label space δ; Step 2.1.2, use cross entropy loss to calculate the target of each trigger word mask prediction; during the training process, randomly sample one, two or three trigger words from the matched trigger word list with a random probability of 6:2:2, and mask them in the query text; in the matching and masking process, there is still a probability that the number of masked trigger words is set to 2 or 3; Among them l tmp is the loss of the TMP module, is the real label of the trigger word in the full vocabulary label space, is the corresponding true probability label distribution of the masked trigger word in the full vocabulary label space.

2. The Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge according to claim 1 is characterized by: The specific steps of Step 1 are: Step 1.1, collect relevant data from Wikipedia and Vietnam News Network, and finally build an aligned Chinese-Vietnamese language pair in order to complete the cross-language event retrieval between Chinese and Vietnamese; Step 1.2, for Wikipedia, extract the document data DocC and DocV with the same month from the chronicles of Chinese and Vietnamese Wikipedia pages; only the Chinese document data DocC is concatenated by hyperlinks to obtain the Chinese query document QueC, and the similarity between the Chinese query document QueC and the Vietnamese document DocV is calculated using cosine similarity on cross-language word embeddings. If the result is greater than the preset threshold β, they are regarded as similar language pairs, thus realizing the alignment of the relevant language pairs of Chinese and Vietnamese; Step 1.3: For news network data, select the paragraph header text SentV in the Vietnamese news text ArtV as the document part of the language pair, and generate the Chinese summary SentC as the query part of the language pair through the output of the Chinese-Vietnamese cross-language summary generation model.

3. The Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge according to claim 1 is characterized by: In Step 2.1, the process of querying and recalling documents is modeled through the retrieval modeling and ranking module RS, positive and negative language pairs are constructed, the model's discrimination of positive and negative examples is continuously optimized, and the specific steps for improving the model's ranking accuracy in downstream retrieval tasks are as follows: Step 2.1.3, for the selection of irrelevant documents in the input sequence y of the retrieval modeling ranking module RS, in order to avoid the model falling into a fixed gradient and causing the optimization process to converge to a local optimal solution, random sampling is used to match irrelevant documents DV during the training process; the retrieval modeling ranking module RS encodes the positive and negative examples of the input sequence y respectively, and obtains the encoded positive and negative examples y{QD+} k and y{QD-} k , where k∈batch_size, batch_size represents the number of samples selected for one training; Step 2.1.4, in the ranking score generation layer, use a learnable parameter weight matrix W to multiply the sequence positive and negative example encoding output y{QD+} k and y{QD-} k , respectively get the ranking scores S+ and S-: and use the cross entropy loss to optimize the model; by continuously optimizing W and minimizing the cross entropy loss function l rs , the model will improve the ranking score S+ for positive examples and continue to reduce the ranking score S- for negative examples, so that the model can make better judgments on positive and negative examples in the retrieval scenario and improve the ranking process of recalled documents; the solution process is: S + / - =y{QD + / - } k *IN.

4. The Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge according to claim 1 is characterized by: The specific steps of Step 2.2 are: Step 2.2.

1. Use the real language pair as the positive sample. When the trigger word predicted by the model is inconsistent with the real trigger word, map the predicted trigger word from the full vocabulary to the trigger word vocabulary label space η, and use the generated predicted text as the negative sample; generate a discriminator to compare the positive and negative samples to continuously optimize the accuracy of the model's prediction of the trigger word; Among them, l d is the loss of the GDC module, D(x t |x)=Sigmoid(f η (h D (x t ), a is a binary label used to determine whether the model prediction is correct, f is a linear layer module, which maps the masked word vector to the trigger word vocabulary label space η, h D and the trigger word mask prediction module h J It is the shared Bert model prediction head; Step 2.2.2, in the process of generating positive and negative examples, there is a p% probability that the trigger words predicted by the model are not used. Instead, the time trigger word prediction task and the event trigger word prediction task are first classified, and then trigger words of the same length are randomly extracted from the trigger word list to replace the predicted trigger words; p is set to 50 to ensure the relative balance of positive and negative samples; The loss function of the trigger word mask prediction module TMP and the generative discriminator module GDC based on contrastive learning in the Chinese-Vietnamese cross-language event pre-training module is to mix the cross entropy loss and contrast loss obtained by the trigger word mask prediction module and iterate and optimize the model through the forward propagation layer. At the same time, the loss of the retrieval modeling ranking module adopts the cross entropy loss.

5. The Chinese-Vietnamese cross-language event retrieval method incorporating event knowledge according to claim 1 is characterized by: The specific steps of Step 3 are: Step 3.

1. The question part in the question-answering dataset is used as the query part in the retrieval task. The correct answer corresponding to the question and the multiple predicted answers generated by the question-answering task are used as the relevant documents in the retrieval. The other question answers and the corresponding multiple predicted answers generated by the question-answering task are used as the irrelevant documents corresponding to the query. Finally, an evaluation dataset of queries and relevant and irrelevant documents in different languages ​​is obtained. Step 3.2: Based on Step 3.1, the average accuracy and MAP indicators are used to evaluate the performance of the model in the retrieval task: Where N represents the total number of relevant documents, position(i) represents the position of the i-th relevant document in the search result list, and the average accuracy is the average of the average correct rates (AP) of multiple queries. The average accuracy is a common evaluation indicator for retrieval tasks and can reflect the retrieval performance of the model as a whole. Step 3.

3. Based on Step 3.1, the retrieval scenario is optimized through the retrieval modeling and ranking module. Therefore, only the last three encoding layers are fine-tuned during the fine-tuning process to avoid over-matching. The model is optimized using cross entropy loss during the evaluation process, and the maximum number of training times is set to 20 times. When the model obtains the best MAP score on the development set, the program evaluates it on the test set and records the corresponding MAP score, and finally selects the best score to save as the result.

Citation Information

Patent Citations

  • Pre-training method and device for cross-language language model

    CN115204408A

  • Event pre-training method for Chinese cross-language event retrieval

    CN115470393A