Cross-language short text matching method introducing bilingual topic knowledge
By introducing bilingual topic knowledge, updating the topic number of short texts, and combining bilingual LDA model and bidirectional LSTM, the problem of insufficient feature extraction in cross-language short text matching is solved, and the accuracy of matching is improved.
Patent Information
- Application Number
- CN202311577546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, in cross-language short text matching, the feature extraction is insufficient due to the short text, making it difficult to achieve good results.
Introduce bilingual topic knowledge, update the topic number of each word through Gibbs probability distribution, obtain the document topic distribution and topic word distribution, and use bilingual LDA model, embedding and bidirectional LSTM modeling to calculate the document vector similarity.
By leveraging bilingual topic knowledge, expanding short text semantic information, alleviating the semantic gap between short texts across languages and improving the accuracy of short text matching across languages.
Smart Images

Figure SMS_6 
Figure SMS_8 
Figure SMS_10
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information data, and particularly relates to a cross - language short text matching method introducing bilingual topic knowledge. Background Art
[0002] With the successful application of deep learning in computer vision, speech recognition, and recommendation systems, in recent years, many studies have been dedicated to applying deep neural network models to natural language processing tasks to reduce the cost of feature engineering and mine implicit semantic features. The earliest to apply deep learning to text matching was Microsoft Research Redmond, and the DSSM [1] and related series of models released by it are relatively influential in deep - learning text - matching models.
[0003] Text matching is a core issue in natural language understanding, and it can be applied to a large number of natural language processing tasks, such as information retrieval, question - answering systems, paraphrasing questions, dialogue systems, machine translation, and so on. These natural language processing tasks can be abstracted into text - matching problems to a large extent. For example, information retrieval can be reduced to the matching of search terms and document resources, question - answering systems can be reduced to the matching of questions and candidate answers, paraphrasing questions can be reduced to the matching of two synonymous sentences, dialogue systems can be reduced to the matching of the previous sentence and the reply, and machine translation can be reduced to the matching of two languages.
[0004] In traditional text - matching technologies, methods such as topic models, word - matching models, and VSM mainly focus on keyword - based matching problems. This type of model requires a large number of manually defined and extracted features as the basis, and these features are task - related and cannot be directly applied to other tasks.
[0005] Currently, methods based on deep neural networks extract short - text vector representations and then calculate the vector similarity between the texts to be matched. There is an obvious problem with this approach, that is, the text is too short, resulting in too few features that can be extracted. Therefore, it is difficult to achieve good results by simply applying deep neural network models. Therefore, the present invention proposes a cross - language short - text matching method introducing bilingual topic knowledge. Summary of the Invention
[0006] Aiming at the above - mentioned disadvantages of the prior art, the first object of the present invention is to provide a cross - language short - text matching method introducing bilingual topic knowledge to solve the problems in the above - mentioned background art.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A cross - language short - text matching method introducing bilingual topic knowledge, comprising the following steps:
[0009] 1) Select an appropriate number of topics K and an appropriate hyperparameter vector
[0010] 2) For each word in each document in the corresponding corpus, randomly assign a topic number z;
[0011] 3) Rescan the corpus and update the topic number of each word using the Gibbs probability distribution;
[0012] 4) Repeat the sampling process in step 3) until the Gibbs sampling converges;
[0013] 5) Count the topics of each word in each document in the corpus to obtain the document-topic distribution and the topic-word distribution of each language and
[0014] By adopting the above technical solution: In the present invention, by updating the topic number of each word using the Gibbs probability distribution, each word can be retrieved, thereby collecting a more comprehensive subject sample.
[0015] Furthermore, the specific calculation method is as follows:
[0016]
[0017]
[0018]
[0019]
[0020]
[0021]
[0022] By adopting the above technical solution: By using the above calculation method for calculation, the data can be expressed more intuitively, thereby understanding the topic-word distribution of each language.
[0023] Furthermore, in step 3), the prediction process of the Gibbs sampling algorithm is as follows:
[0024] a) For each word in the current document, randomly assign a topic number z ;
[0025] b) Rescan the current document and update the topic number of each word using the Gibbs sampling formula;
[0026] c) Repeat the sampling process in step 2 until the Gibbs sampling converges;
[0027] d) Count the topics of each word in the document to obtain the topic distribution of the document.
[0028] By adopting the above technical solution: each word in the document is renumbered, and then compared with the previous document, so as to realize the scanning and sampling of each word, thus making the obtained data more accurate.
[0029] Furthermore, the method framework:
[0030] Input: short text pair <d r , d y >
[0031] Output: whether the two short texts are similar
[0032] A) Input the two documents into the trained bilingual LDA model respectively to obtain their corresponding topic representations;
[0033] B) Pass the two documents through embedding and bidirectional LSTM modeling respectively to obtain their corresponding semantic representations;
[0034] C) Concatenate the output vector at the last moment of the LSTM with the topic expression corresponding to each document to obtain the final vector representation;
[0035] D) Calculate the vector similarity of the two documents. When the similarity is greater than 0.5, the two documents are considered similar; otherwise, the two documents are not similar.
[0036] By adopting the above technical solution: by using bilingual topic knowledge, the semantic information of short texts is extended, the semantic gap of cross-language short texts is alleviated, and the accuracy of cross-language short text matching is improved.
[0037] Beneficial effects
[0038] By adopting the technical solution provided by the present invention, compared with the known public technology, the following beneficial effects are obtained:
[0039] In the present invention, the two documents are respectively input into the trained bilingual LDA model to obtain their corresponding topic representations, and then the two documents are respectively passed through embedding and bidirectional LSTM modeling to obtain their corresponding semantic representations. Then, the output vector at the last moment of the LSTM is concatenated with the topic expression corresponding to each document to obtain the final vector representation. Finally, the vector similarity of the two documents is calculated. When the similarity is greater than 0.5, the two documents are considered similar; otherwise, the two documents are not similar. By using bilingual topic knowledge, the semantic information of short texts is extended, the semantic gap of cross-language short texts is alleviated, and the accuracy of cross-language short text matching is improved. Detailed implementation manners
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without any creative efforts shall fall within the protection scope of the present invention.
[0041] The present invention will be further described below in conjunction with embodiments.
[0042] Embodiment 1
[0043] A cross-language short text matching method for introducing bilingual topic knowledge includes the following steps:
[0044] 1) Select a suitable number of topics K and a suitable hyperparameter vector
[0045] 2) Randomly assign a topic number to each word in each document in the corresponding corpus z ;
[0046] 3) Rescan the corpus and update the topic number of each word using the Gibbs probability distribution;
[0047] 4) Repeat the sampling process in step 3) until the Gibbs sampling converges;
[0048] 5) Count the topics of each word in each document in the corpus to obtain the document topic distribution and the topic word distribution of each language and
[0049] Among them, the specific calculation method is as follows:
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] Among them, in the said step 3), the prediction process of the Gibbs sampling algorithm is as follows:
[0057] a) Randomly assign a topic number to each word corresponding to the current document z ;
[0058] b) Rescan the current document and update the topic number of each word using the Gibbs sampling formula;
[0059] c) Repeat the sampling process in step 2 until the Gibbs sampling converges;
[0060] d) Count the topics of each word in the document to obtain the topic distribution of the document.
[0061] Among them, the method framework:
[0062] Input: Short text pair <d r , d y >
[0063] Output: Whether two short texts are similar
[0064] A) Input the two documents into the trained bilingual LDA model respectively to obtain their corresponding topic representations;
[0065] B) Model the two documents through embedding and bidirectional LSTM respectively to obtain their corresponding semantic representations;
[0066] C) Concatenate the output vector at the last moment of the LSTM with the topic expression corresponding to each document to obtain the final vector representation;
[0067] D) Calculate the vector similarity of the two documents. When the similarity is greater than 0.5, the two documents are considered similar; otherwise, the two documents are not similar.
[0068] Comparative Example 1
[0069] This embodiment is substantially the same as the method of Embodiment 1 provided, and the main difference is that: in step 3), topic numbers are not assigned to each word in the document;
[0070] Comparative Example 2
[0071] This embodiment is substantially the same as the method of Embodiment 1 provided, and the main difference is that: in step b), the previous document is not scanned.
[0072] Comparative Example 3
[0073] This embodiment is substantially the same as the method of Embodiment 1 provided, and the main difference is that: in step 5), the learning algorithm is not used.
[0074] Performance test
[0075] Respectively take equal amounts of the data accuracy and matching result error values of the cross - language short - text matching method introducing bilingual theme knowledge provided in Example 1 and Comparative Examples 1 - 3:
[0076]
[0077]
[0078] By analyzing the relevant data in the above tables, the specific steps of the matching method of the present invention are as follows:
[0079] 1) Select an appropriate number of topics K and select an appropriate hyper - parameter vector
[0080] 2) For each word in each document in the corresponding corpus, randomly assign a topic number z;
[0081] 3) Rescan the corpus. For each word, update its topic number using the Gibbs probability distribution;
[0082] 4) Repeat the sampling process in step 3) until the Gibbs sampling converges;
[0083] 5) Count the topics of each word in each document in the corpus to obtain the document - topic distribution and the topic - word distribution of each language and
[0084] The beneficial effects achieved: In the present invention, by updating the topic number of each word using the Gibbs probability distribution, each word can be retrieved, so as to collect a more comprehensive subject sample.
[0085] Furthermore, the specific calculation method is as follows:
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] The beneficial effects achieved: By using the above - mentioned calculation method for calculation, the data can be expressed more intuitively, so as to understand the topic - word distribution of each language.
[0093] Furthermore, in step 3), the prediction process of the Gibbs sampling algorithm is as follows:
[0094] a) For each word corresponding to the current document, randomly assign a topic number z ;
[0095] b) Rescan the current document, and for each word, update its topic number using the Gibbs sampling formula;
[0096] c) Repeat the sampling process in step 2 until the Gibbs sampling converges;
[0097] d) Count the topics of each word in the document to obtain the topic distribution of the document.
[0098] The beneficial effects achieved: By re - numbering each word in the document and then comparing it with the previous document, scanning and sampling for each word are realized, so that the obtained data is more accurate.
[0099] Furthermore, the method framework:
[0100] Input: Short text pair <d r , d y >
[0101] Output: Whether two short texts are similar
[0102] A) Input the two documents into the trained bilingual LDA model respectively to obtain their corresponding topic representations;
[0103] B) Pass the two documents through embedding and bidirectional LSTM modeling respectively to obtain their corresponding semantic representations;
[0104] C) Concatenate the output vector at the last moment of the LSTM with the topic expression corresponding to each document to obtain the final vector representation;
[0105] D) Calculate the vector similarity of the two documents. When the similarity is greater than 0.5, the two documents are considered similar; otherwise, the two documents are not similar.
[0106] The beneficial effects achieved: By using bilingual topic knowledge, the semantic information of short texts is expanded, the semantic gap of cross - language short texts is alleviated, and the accuracy of cross - language short text matching is improved.
[0107] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A cross - language short text matching method introducing bilingual topic knowledge, characterized in that, it includes the following steps: 1) Select an appropriate number of topics K and select an appropriate hyperparameter vector 2) For each word in each document of the corresponding corpus, randomly assign a topic number z; 3) Rescan the corpus. For each word, update its topic number using the Gibbs probability distribution; 4) Repeat the sampling process in step 3) until the Gibbs sampling converges; 5) Statistically analyze the topics of each word in each document in the corpus to obtain the document topic distribution and the topic word distribution of each language and 2. The cross - language short text matching method introducing bilingual topic knowledge according to claim 1, characterized in that: In the said step 5), the specific calculation method is as follows:
3. The cross - language short text matching method introducing bilingual topic knowledge according to claim 1, characterized in that: In the said step 3), the prediction process of the Gibbs sampling algorithm is as follows: a) For each word corresponding to the current document, randomly assign a topic number z; b) Rescan the current document. For each word, update its topic number using the Gibbs sampling formula; c) Repeat the sampling process in step 2) until the Gibbs sampling converges; d) Count the topics of each word in the document to obtain the topic distribution of the document.
4. The cross - language short text matching method introducing bilingual topic knowledge according to claim 1, characterized in that, Method framework: Input: Short text pair <d x , d y > Output: Whether two short texts are similar A) Input the two documents into the trained bilingual LDA model respectively to obtain their corresponding topic representations; B) Subject the two documents to embedding and bidirectional LSTM modeling respectively to obtain their corresponding semantic representations; C) Concatenate the output vector at the last moment of the LSTM with the topic expression corresponding to each document to obtain the final vector representation; D) Calculate the vector similarity of the two documents. When the similarity is greater than 0.5, the two documents are considered similar; otherwise, the two documents are not similar.