Cross-language document retrieval method and system based on knowledge graph, medium and equipment

By constructing and optimizing a cross-language document retrieval model based on knowledge graph, the problem of insufficient utilization of knowledge graph information in cross-language document retrieval is solved, and the accuracy and efficiency of retrieval is improved.

CN120011500APending Publication Date: 2025-05-16YUNNAN POWER GRID CO LTD ELECTRIC POWER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510004761.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In cross-language document retrieval, it is difficult for the prior art to effectively utilize knowledge graph information, resulting in poor cross-language retrieval performance, especially in low-resource language environments.

Method used

Documents of the first and second languages ​​are obtained through cross-language text data, and knowledge interaction diagrams of query documents and candidate document diagrams are constructed, feature encoding is used using graph convolutional networks, and the model is optimized through loss functions to make full use of knowledge graph information.

Benefits of technology

It improves the accuracy and efficiency of cross-language document retrieval, enhances the model's ability to learn semantic information on different languages, and reduces the noise impact caused by the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011500A_ABST
    Figure CN120011500A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a knowledge graph-based cross-language document retrieval method and system, a medium and equipment. The method comprises the following steps: acquiring a document of a first language and a document of a second language through cross-language text data; constructing a query document knowledge interaction graph of the first language and a candidate document graph of the second language according to document keywords in the document of the first language and the document of the second language; performing feature coding on the query document knowledge interaction graph and the candidate document graph by utilizing a graph convolutional network, and constructing a cross-language document retrieval model; optimizing the cross-language document retrieval model according to the model loss function to obtain a cross-language document optimal retrieval model; by inputting the query statement of the first language into the trained cross-language document optimal retrieval model, the document retrieval sorting result of the second language is obtained, and a convenient and rapid cross-language document retrieval service is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language processing technology, and in particular to a cross-language document retrieval method, system, medium and device based on a knowledge graph. Background Art

[0002] Cross-lingual document retrieval involves the process of using documents in one language as queries and then retrieving relevant documents from a document set in another language. However, when facing the cross-lingual document retrieval task in two languages, the lack of sufficient annotated data and obvious language differences make it difficult to align query and document matching.

[0003] Traditional methods based on machine translation and pre-trained language models perform well in resource-rich bilingual translation, but when it comes to some niche languages, they are limited by translation performance and their results are not satisfactory. In addition, the alignment effect of cross-language pre-trained language models in low-resource languages ​​is limited, which also limits the development of cross-language retrieval methods based on pre-trained language models. Therefore, in recent years, the introduction of knowledge graphs to enhance query semantic information has made significant progress, but the current model does not fully utilize the knowledge graph information, resulting in unsatisfactory cross-language retrieval performance. Summary of the invention

[0004] Based on this, it is necessary to propose a cross-language document retrieval method, device, medium and equipment based on knowledge graph to address the above problems.

[0005] A cross-language document retrieval method based on a knowledge graph, the method comprising:

[0006] Obtain documents in the first language and documents in the second language through cross-language text data.

[0007] A query document knowledge interaction graph in the first language and a candidate document graph in the second language are constructed according to document keywords in the first language and the second language.

[0008] A graph convolutional network is used to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model.

[0009] The cross-language document retrieval model is optimized according to the loss function to obtain the optimal cross-language document retrieval model.

[0010] The query statement in the first language is input into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

[0011] The step of acquiring a document in a first language and a document in a second language through cross-language text data specifically includes:

[0012] Determine the relevance of candidate documents based on cross-lingual text data.

[0013] Documents whose candidate document relevance is a preset value are selected as documents in the first language and documents in the second language.

[0014] The document in the second language is translated into the first language, and the translation results are screened using evaluation indicators to obtain high-quality data and save the high-quality data to the document in the first language.

[0015] The step of constructing a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on the document keywords in the first language and the document in the second language specifically includes:

[0016] A query document graph in the first language and a candidate document graph in the second language are respectively constructed according to document keywords of the documents in the first language and the documents in the second language.

[0017] The multilingual entity linking model is used to annotate the Wikipedia query identifiers of the entities corresponding to the document keywords in the first language.

[0018] The external knowledge information of the entity is queried through the Wikipedia query identifier.

[0019] The external knowledge information is integrated into the query document graph of the first language to obtain a query document knowledge interaction graph.

[0020] The method of using a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model specifically includes:

[0021] By randomly selecting data augmentation operators to perform comparative learning on the query document knowledge interaction graph, different document knowledge graph structure views are obtained.

[0022] A graph convolutional network is used to perform feature encoding on the document knowledge graph structure view and the candidate document graph respectively, to obtain feature vectors of knowledge interaction graphs and candidate document graphs with different structures.

[0023] A cross-language document retrieval model is constructed based on the knowledge interaction graph feature vector and the candidate document graph feature vector.

[0024] The step of optimizing the cross-language document retrieval model according to the model loss function to obtain the optimal cross-language document retrieval model specifically includes:

[0025] The feature vectors of the knowledge interaction graphs of different structures constitute positive samples, the feature vectors of the candidate document graphs constitute negative samples, and the document knowledge graph comparison loss function is determined based on the positive samples and the negative samples.

[0026] A similarity loss function is determined based on the document knowledge graph structure view and the candidate document graph.

[0027] The document knowledge graph comparison loss and similarity loss are added together to determine the model loss function.

[0028] The cross-language document retrieval model is optimized according to the model loss function to obtain an optimal cross-language document retrieval model.

[0029] The feature vectors of the knowledge interaction graphs of different structures constitute positive samples, and the feature vectors of the candidate document graphs constitute negative samples. The document knowledge graph contrast loss function is determined according to the positive samples and the negative samples, specifically including:

[0030] The document knowledge graph comparison loss function is:

[0031]

[0032] Among them, L DSCL is the document knowledge graph comparison loss function, i is the i-th candidate document graph, A d (i) is the sample set of positive and negative samples of query i, ad is the sample set of current positive and negative samples, j is the anchor instance, Pd(i) is the positive sample set of query i, pd is the current positive sample, |Pd(i)| is the number of positive samples of query i, z j is the feature vector representation corresponding to the anchor instance, · is the inner product, τ is a temperature coefficient that controls the distance between samples, and R+ is a positive real number.

[0033] The determining of the similarity loss function according to the document knowledge graph structure view and the candidate document graph specifically includes:

[0034] The similarity loss function is:

[0035]

[0036] Among them, L qd is the similarity loss, G k is the document knowledge graph structure view, G kd is the candidate document graph, is the set of relevant documents corresponding to the query, is the set of irrelevant documents corresponding to the query, max{0,·} is the maximum value, is the graph of irrelevant documents corresponding to the query.

[0037] The step of inputting the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language specifically includes:

[0038] The cross-language document optimal retrieval model is loaded into the memory.

[0039] The loaded cross-language document optimal retrieval model is deployed as an API interface using the Flask framework.

[0040] A query statement in a first language is received through the API interface and input into the cross-language document optimal retrieval model.

[0041] The cross-language document optimal retrieval model performs retrieval according to a query statement in a first language to obtain a document retrieval ranking result in a second language.

[0042] A cross-language document retrieval system based on a knowledge graph, the system comprising:

[0043] The module for acquiring documents in the first language and documents in the second language is used to acquire documents in the first language and documents in the second language through cross-language text data.

[0044] The query document knowledge interaction graph and candidate document graph construction module is used to construct the query document knowledge interaction graph in the first language and the candidate document graph in the second language based on the document keywords in the documents in the first language and the documents in the second language.

[0045] The cross-language document retrieval model construction module is used to use a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model.

[0046] The cross-language document optimal retrieval model acquisition module is used to optimize the cross-language document retrieval model according to the loss function to obtain the cross-language document optimal retrieval model.

[0047] The document retrieval module is used to input the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

[0048] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method described above.

[0049] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0050] The embodiments of the present invention have the following beneficial effects:

[0051] The present invention constructs a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on the document keywords in the first language and the document in the second language, and represents the keywords in the document and their relationships in a graphical manner, which helps to better capture the semantic connections and contextual information between documents, thereby improving the accuracy and efficiency of retrieval. Furthermore, a graph convolutional network is used to perform feature encoding on the query document knowledge interaction graph and the candidate document graph, and a cross-language document retrieval model is constructed. The loss function is used to optimize the model to ensure that the model continuously reduces the prediction error during the training process, so that the query document knowledge interaction graph and the candidate document graph in the second language are matched and aligned, the semantic information of the first language and the second language is fully learned, and the influence of noise brought by the knowledge graph is reduced, thereby improving the retrieval accuracy. Finally, by inputting the query statement in the first language into the trained cross-language document optimal retrieval model, the document retrieval ranking result in the second language is obtained, providing a convenient and fast cross-language document retrieval service. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0053] in:

[0054] Figure 1 A flowchart of an embodiment of a Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement provided by the present invention;

[0055] Figure 2 A flowchart of another embodiment of a Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement provided by the present invention;

[0056] Figure 3 A schematic diagram of the structure of an embodiment of a Chinese-Vietnamese cross-language document retrieval system based on knowledge graph enhancement provided by the present invention;

[0057] Figure 4 A schematic structural diagram of an embodiment of the device provided by the present invention;

[0058] Figure 5 A schematic structural diagram of an embodiment of the medium provided by the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] like Figure 1 As shown, Figure 1 A flow chart of an embodiment of a Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement provided by the present invention. A Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement, the method comprising:

[0061] S101: Acquire a document in a first language and a document in a second language through cross-language text data.

[0062] Exemplarily, the relevance of candidate documents is determined based on publicly available cross-language text data; documents with a relevance of 6 for the queried candidate documents are taken as documents in the first language and documents in the second language, and other irrelevant documents are taken as irrelevant documents; then the documents in the second language are translated into the first language, and the ROUGE evaluation index is used to screen the translation results, select high-quality data, and save the high-quality data to documents in the first language.

[0063] S102: Constructing a query document knowledge interaction graph in the first language and a candidate document graph in the second language according to document keywords in the first language and the second language.

[0064] Exemplarily, a query document graph in the first language and a candidate document graph in the second language are respectively constructed based on document keywords in the first language and document in the second language, and a Wikipedia query identifier of an entity corresponding to the document keyword in the first language is annotated using a multilingual entity linking model; external knowledge information of the entity is queried through the Wikipedia query identifier; the external knowledge information is integrated into the query document graph in the first language to obtain a query document knowledge interaction graph.

[0065] S103: Use graph convolutional networks to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to build a cross-language document retrieval model.

[0066] Exemplarily, by randomly selecting data augmentation operators to perform comparative learning on the query document knowledge interaction graph, different document knowledge graph structural views are obtained; a graph convolutional network is used to perform feature encoding on the document knowledge graph structural view and the candidate document graph, respectively, to obtain knowledge interaction graph feature vectors and candidate document graph feature vectors of different structures; and a cross-language document retrieval model is constructed based on the knowledge interaction graph feature vectors and the candidate document graph feature vectors.

[0067] Specifically, we first randomly select the data augmentation operator Sq(·) to perform comparative learning on the query document knowledge interaction graph, generate different document knowledge graph structure views for comparative learning, and reflect the invariance of the document knowledge graph to knowledge noise through the consistency of different structure views, and finally generate document knowledge graph structure views with different expansion structures. The specific formula is as follows:

[0068]

[0069] Among them, G k1 , G k2 G is a document knowledge graph structure view with two different structures obtained by randomly selecting augmentation operators. k is the document knowledge interaction graph, ⊙ is the product operation, For the masking vectors corresponding to two knowledge interaction graphs with different structures, define the masking vector It is a binary indicator, indicating whether a specific knowledge node is selected.

[0070] Furthermore, a graph convolutional network is used to perform feature encoding on the constructed document knowledge graph structure view and candidate document graph, where the document knowledge graph structure view and candidate document graph share a common encoder Ed. The specific encoding formula is as follows:

[0071] v q1 =E d (G k1 );

[0072] v q2 =E d (G k2 );

[0073]

[0074] Among them, v q1 、v q2 are the knowledge interaction graph feature vectors of two different knowledge interaction graphs, is the candidate document graph feature vector of all candidate document graphs, i is the i-th candidate document graph, E d For the encoder, is the candidate document graph, G k1 , G k2 Structural views of document knowledge graphs with two different structures obtained by randomly selecting augmentation operators.

[0075] Furthermore, a knowledge graph enhanced cross-language document retrieval model is constructed based on the knowledge interaction graph feature vector and the candidate document graph feature vector.

[0076] S104: Optimize the cross-language document retrieval model according to the loss function to obtain the optimal cross-language document retrieval model.

[0077] Exemplarily, feature vectors of knowledge interaction graphs of different structures constitute positive samples, and feature vectors of candidate document graphs constitute negative samples. The document knowledge graph comparison loss function is determined based on the positive samples and negative samples; the similarity loss function is determined based on the document knowledge graph structural view and the candidate document graph; the document knowledge graph comparison loss and the similarity loss are added to determine the model loss function; the cross-language document retrieval model is optimized based on the model loss function to obtain the optimal cross-language document retrieval model.

[0078] S105: Inputting the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

[0079] Exemplarily, a cross-language document optimal retrieval model is loaded into memory; the loaded cross-language document optimal retrieval model is deployed as an API interface using the Flask framework; a query statement in a first language is received through the API interface and input into the cross-language document optimal retrieval model; the cross-language document optimal retrieval model performs a search based on the query statement in the first language to obtain a document retrieval ranking result in a second language.

[0080] From the above description, it can be seen that the present invention constructs a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on the document keywords in the first language and the document in the second language, and represents the keywords in the document and their relationships in a graphical manner, which helps to better capture the semantic connection and context information between documents, thereby improving the accuracy and efficiency of retrieval. Further, the query document knowledge interaction graph and the candidate document graph are feature encoded using a graph convolutional network to construct a cross-language document retrieval model, and the model is optimized through a loss function to ensure that the model continuously reduces the prediction error during the training process, so that the query document knowledge interaction graph and the candidate document graph in the second language are matched and aligned, and the semantic information of the first language and the second language is fully learned, and the influence of noise brought by the knowledge graph is reduced, thereby improving the retrieval accuracy. Finally, by inputting the query statement in the first language into the trained cross-language document optimal retrieval model, the document retrieval ranking result in the second language is obtained, providing a convenient and fast cross-language document retrieval service.

[0081] like Figure 2 As shown, Figure 2 A flow chart of another embodiment of a Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement provided by the present invention. A Chinese-Vietnamese cross-language document retrieval method based on knowledge graph enhancement, the method comprising:

[0082] S201: Determine the relevance of candidate documents based on cross-language text data.

[0083] S202: Selecting documents whose candidate document relevance is a preset value as documents in the first language and documents in the second language.

[0084] S203: Translate the document in the second language into the first language, screen the translation results using evaluation indicators, obtain high-quality data, and save the high-quality data to the document in the first language.

[0085] It should be noted that steps S201-S203 are Figure 1 This has been discussed in detail in the implementation scenario shown and will not be repeated here.

[0086] S204: Constructing a query document graph in the first language and a candidate document graph in the second language respectively according to the document keywords of the document in the first language and the document in the second language.

[0087] S205: Using a multilingual entity linking model, annotate the Wikipedia query identifier of the entity corresponding to the document keyword in the first language.

[0088] S206: Query external knowledge information of the entity through the Wikipedia query identifier.

[0089] S207: Integrate the external knowledge information into the query document graph of the first language to obtain a query document knowledge interaction graph.

[0090] Exemplarily, based on the document keywords of the document in the first language and the document in the second language, the query document graph in the first language and the candidate document graph in the second language are respectively constructed, and then the multilingual entity linking model mGENRE is used to annotate the Wikipedia query identifier (QID) of the entity corresponding to the document keyword, and then the Wikipedia multilingual knowledge graph API is used to query the source language and target language related entities, aliases, entity descriptions and other external knowledge information of the entity through the QID. The external knowledge information is integrated into the query document graph in the first language to obtain the query document knowledge interaction graph.

[0091] S208: Perform comparative learning on the query document knowledge interaction graph by randomly selecting data augmentation operators to obtain different document knowledge graph structure views.

[0092] For example, by randomly selecting the data augmentation operator Sq(·) to perform comparative learning on the query document knowledge interaction graph, different document knowledge graph structure views are generated for comparative learning. The consistency of different structure views can reflect the invariance of the document knowledge graph to knowledge noise, and finally generate document knowledge graph structure views with different expansion structures. The specific formula is as follows:

[0093]

[0094] Among them, G k1 , G k2 G is the document knowledge graph structure view with two different structures obtained by randomly selecting augmentation operators. k is the document knowledge interaction graph, ⊙ is the product operation, For the masking vectors corresponding to two knowledge interaction graphs with different structures, define the masking vector It is a binary indicator, indicating whether a specific knowledge node is selected.

[0095] S209: Using a graph convolutional network to perform feature encoding on the document knowledge graph structure view and the candidate document graph respectively, and obtaining feature vectors of knowledge interaction graphs and candidate document graphs with different structures.

[0096] Exemplarily, a graph convolutional network is used to perform feature encoding on the constructed document knowledge graph structure view and the candidate document graph, wherein the document knowledge graph structure view and the candidate document graph share an encoder Ed, and the specific encoding formula is as follows:

[0097] v q1 =E d (G k1 );

[0098] v q2 =E d (G k2 );

[0099]

[0100] Among them, v q1 、v q2 are the knowledge interaction graph feature vectors of two different knowledge interaction graphs, is the candidate document graph feature vector of all candidate document graphs, i is the i-th candidate document graph, E d For the encoder, is the candidate document graph, G k1 , G k2 Structural views of document knowledge graphs with two different structures obtained by randomly selecting augmentation operators.

[0101] S210: Construct a cross-language document retrieval model based on the knowledge interaction graph feature vector and the candidate document graph feature vector.

[0102] S211: Feature vectors of knowledge interaction graphs of different structures constitute positive samples, feature vectors of candidate document graphs constitute negative samples, and a document knowledge graph comparison loss function is determined based on the positive samples and negative samples.

[0103] Exemplarily, the feature vectors of knowledge interaction graphs of different structures constitute positive samples, and the feature vectors of candidate document graphs constitute negative samples. The document knowledge graph contrast loss function is determined according to the positive samples and the negative samples. The document knowledge graph contrast loss function is:

[0104]

[0105] Among them, L DSCL is the document knowledge graph comparison loss function, i is the i-th candidate document graph, A d (i) is the sample set of positive and negative samples of query i, ad is the sample set of current positive and negative samples, j is the anchor instance, Pd(i) is the positive sample set of query i, pd is the current positive sample, |Pd(i)| is the number of positive samples of query i, z j is the feature vector representation corresponding to the anchor instance, · is the inner product, τ is a temperature coefficient that controls the distance between samples, and R+ is a positive real number.

[0106] S212: Determine a similarity loss function based on the document knowledge graph structure view and the candidate document graph.

[0107] Exemplarily, a similarity loss function is determined according to the document knowledge graph structure view and the candidate document graph, and the similarity loss function is:

[0108]

[0109] Among them, L qd is the similarity loss, G k is the document knowledge graph structure view, G kd is the candidate document graph, is the set of relevant documents corresponding to the query, is the set of irrelevant documents corresponding to the query, max{0,·} is the maximum value, is the graph of irrelevant documents corresponding to the query.

[0110] S213: Add the document knowledge graph comparison loss and similarity loss to determine the model loss function.

[0111] Exemplarily, the model loss function L is:

[0112] L=L DSCL +L qd .

[0113] S214: Optimize the cross-language document retrieval model according to the model loss function to obtain the optimal cross-language document retrieval model.

[0114] S215: The optimal cross-language document retrieval model is loaded into the memory.

[0115] Exemplarily, the trained cross-language document optimal retrieval model is saved as a ".pth" file, and the model is loaded into memory through the Flask framework to avoid frequent model loading processes caused by request results and improve the running speed of the retrieval model.

[0116] S216: Use the Flask framework to deploy the loaded cross-language document optimal retrieval model as an API interface.

[0117] Exemplarily, the loaded cross-language document optimal retrieval model is deployed as an API interface using the Flask framework, thereby realizing the function of multiple concurrent requests of the Web port.

[0118] S217: Receive the query statement in the first language through the API interface and input it into the cross-language document optimal retrieval model.

[0119] S218: The cross-language document optimal retrieval model performs retrieval according to the query statement in the first language to obtain a document retrieval ranking result in the second language.

[0120] Exemplarily, the cross-language document optimal retrieval model deployed on the server side is called on the Web side to test the query statement input in the first language, obtain the document retrieval ranking result in the second language and display it on the front-end interface.

[0121] From the above description, it can be seen that the present invention first constructs a query document graph in the first language and a candidate document graph in the second language according to document keywords of the first language document and the second language document respectively, uses a multilingual entity linking model to annotate the Wikipedia query identifier of the entity corresponding to the document keyword in the first language, queries the external knowledge information of the entity through the Wikipedia query identifier, integrates the external knowledge information into the query document graph in the first language, effectively utilizes external knowledge to enrich the knowledge information of the query document, enhances the representation alignment ability and knowledge feature fusion ability of the query document graph in the first language, and enables the obtained document knowledge interaction graph to make up for the scarcity of first and second language annotation data, thereby enhancing the model's representation alignment ability for Chinese-Vietnamese cross-language and improving the performance of the cross-language document retrieval model.

[0122] like Figure 3 As shown, Figure 3 A schematic diagram of a structure of an embodiment of a Chinese-Vietnamese cross-language document retrieval system based on knowledge graph enhancement provided by the present invention. A cross-language document retrieval system 10 based on knowledge graph, the system comprises:

[0123] The first language document and second language document acquisition module 11 is used to acquire the first language document and the second language document through the cross-language text data.

[0124] The query document knowledge interaction graph and candidate document graph construction module 12 is used to construct a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on document keywords in the first language document and the second language document.

[0125] The cross-language document retrieval model construction module 13 is used to use a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model.

[0126] The cross-language document optimal retrieval model acquisition module 14 is used to optimize the cross-language document retrieval model according to the loss function to obtain the cross-language document optimal retrieval model.

[0127] The document retrieval module 15 is used to input the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

[0128] Exemplarily, in the first language document and second language document acquisition module 11, the candidate document relevance is determined based on the cross-language text data; the document with a candidate document relevance of a preset value is selected as the first language document and the second language document; the second language document is translated into the first language, and the translation results are screened using evaluation indicators to obtain high-quality data and save the high-quality data to the first language document. In the query document knowledge interaction graph and candidate document graph construction module 12, the query document graph of the first language and the candidate document graph of the second language are respectively constructed based on the document keywords of the first language document and the second language document; the Wikipedia query identifier of the entity corresponding to the document keyword of the first language is annotated using a multilingual entity linking model; the external knowledge information of the entity is queried through the Wikipedia query identifier; the external knowledge information is integrated into the query document graph of the first language to obtain the query document knowledge interaction graph. In the cross-language document retrieval model construction module 13, the query document knowledge interaction graph is subjected to comparative learning by randomly selecting data augmentation operators to obtain different document knowledge graph structure views; the document knowledge graph structure view and the candidate document graph are feature encoded by the graph convolutional network to obtain knowledge interaction graph feature vectors and candidate document graph feature vectors of different structures; a cross-language document retrieval model is constructed based on the knowledge interaction graph feature vectors and the candidate document graph feature vectors. In the cross-language document optimal retrieval model acquisition module 14, the knowledge interaction graph feature vectors of different structures constitute positive samples, and the candidate document graph feature vectors constitute negative samples. The document knowledge graph comparison loss function is determined based on the positive samples and the negative samples; the similarity loss function is determined based on the document knowledge graph structure view and the candidate document graph; the document knowledge graph comparison loss and the similarity loss are added to determine the model loss function; the cross-language document retrieval model is optimized based on the model loss function to obtain the cross-language document optimal retrieval model. In the document retrieval module 15, the cross-language document optimal retrieval model is loaded into the memory; the loaded cross-language document optimal retrieval model is deployed as an API interface using the Flask framework; the query statement in the first language is received through the API interface and input into the cross-language document optimal retrieval model; the cross-language document optimal retrieval model performs retrieval according to the query statement in the first language to obtain the document retrieval ranking result in the second language.

[0129] like Figure 4 As shown, Figure 4 The device 20 includes a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 executes the computer program when working to implement the following. Figure 1 and Figure 2 The method shown.

[0130] The specific technical details of a knowledge graph-based cross-language document retrieval method implemented when the above-mentioned device 20 executes a computer program have been discussed in detail in the aforementioned method steps, so they will not be repeated here.

[0131] like Figure 5 As shown, Figure 5 The structure diagram of an embodiment of the medium provided by the present invention is shown in FIG. The medium 30 stores at least one computer program 31, and the computer program 31 is executed by the processor 22 to implement the following Figure 1 and Figure 2 In one embodiment, the medium 30 may be a storage chip, a hard disk, a mobile hard disk, a USB flash drive, an optical disk, or other readable and writable storage tools, or a server, etc.

[0132] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0133] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer-readable storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0134] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification correspond to each other, and therefore, the apparatus, device, and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be repeated here.

[0135] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0136] For the convenience of description, the above device is described by being divided into various units according to their functions and described separately. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware. It should be understood by those skilled in the art that this specification embodiment can be provided as a method, system, or computer program product. Therefore, this specification embodiment can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification embodiment can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0138] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0139] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A cross-language document retrieval method based on knowledge graph, characterized in that: The method comprises: Obtain documents in a first language and documents in a second language through cross-language text data; constructing a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on document keywords in the document in the first language and the document in the second language; Using a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph, and constructing a cross-language document retrieval model; Optimizing the cross-language document retrieval model according to the model loss function to obtain an optimal cross-language document retrieval model; The query statement in the first language is input into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

2. The cross-language document retrieval method based on knowledge graph according to claim 1, characterized in that: The acquiring of the document in the first language and the document in the second language through the cross-language text data specifically includes: Determining the relevance of candidate documents based on cross-language text data; Selecting documents whose candidate document relevance is a preset value as documents in the first language and documents in the second language; The document in the second language is translated into the first language, and the translation results are screened using evaluation indicators to obtain high-quality data and save the high-quality data to the document in the first language.

3. The cross-language document retrieval method based on knowledge graph according to claim 2, characterized in that: The step of constructing a query document knowledge interaction graph in the first language and a candidate document graph in the second language based on the document keywords in the first language and the document in the second language specifically includes: constructing a query document graph in the first language and a candidate document graph in the second language respectively according to document keywords of the documents in the first language and the documents in the second language; Using a multilingual entity linking model, annotating the Wikipedia query identifiers of entities corresponding to the document keywords in the first language; Querying external knowledge information of the entity through the Wikipedia query identifier; The external knowledge information is integrated into the query document graph of the first language to obtain a query document knowledge interaction graph.

4. The cross-language document retrieval method based on knowledge graph according to claim 3 is characterized in that: The method of using a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model specifically includes: By randomly selecting a data augmentation operator, a comparative study is performed on the query document knowledge interaction graph to obtain different document knowledge graph structure views; Using a graph convolutional network to perform feature encoding on the document knowledge graph structure view and the candidate document graph respectively, to obtain feature vectors of knowledge interaction graphs and candidate document graphs of different structures; A cross-language document retrieval model is constructed based on the knowledge interaction graph feature vector and the candidate document graph feature vector.

5. The cross-language document retrieval method based on knowledge graph according to claim 4 is characterized in that: The optimizing the cross-language document retrieval model according to the model loss function to obtain the cross-language document optimal retrieval model specifically includes: The feature vectors of the knowledge interaction graphs with different structures constitute positive samples, the feature vectors of the candidate document graphs constitute negative samples, and the document knowledge graph comparison loss function is determined according to the positive samples and the negative samples; Determine a similarity loss function based on the document knowledge graph structure view and the candidate document graph; Adding the document knowledge graph comparison loss and similarity loss to determine the model loss function; The cross-language document retrieval model is optimized according to the model loss function to obtain an optimal cross-language document retrieval model.

6. The cross-language document retrieval method based on knowledge graph according to claim 5, characterized in that: The feature vectors of the knowledge interaction graphs of different structures constitute positive samples, and the feature vectors of the candidate document graphs constitute negative samples. The document knowledge graph comparison loss function is determined according to the positive samples and the negative samples, specifically including: The document knowledge graph comparison loss function is: Among them, L DSCL is the document knowledge graph comparison loss function, i is the i-th candidate document graph, A d (i) is the sample set of positive and negative samples of query i, a d is the current sample set of positive and negative samples, j is the anchor instance, Pd(i) is the positive sample set of query i, p d is the current positive sample, |Pd(i)| is the number of positive samples of query i, z j is the feature vector representation corresponding to the anchor instance, · is the inner product, τ is a temperature coefficient that controls the distance between samples, and R+ is a positive real number.

7. The cross-language document retrieval method based on knowledge graph according to claim 5, characterized in that: Determining the similarity loss function according to the document knowledge graph structure view and the candidate document graph specifically includes: The similarity loss function is: Among them, L qd is the similarity loss, G k is the document knowledge graph structure view, G kd is the candidate document graph, is the set of relevant documents corresponding to the query, is the set of irrelevant documents corresponding to the query, max{0,·} is the maximum value, is the graph of irrelevant documents corresponding to the query.

8. The cross-language document retrieval method based on knowledge graph according to claim 5, characterized in that: The step of inputting the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language specifically includes: The cross-language document optimal retrieval model is loaded into the memory; Deploy the loaded cross-language document optimal retrieval model as an API interface using the Flask framework; Receiving a query statement in a first language through the API interface and inputting it into the cross-language document optimal retrieval model; The cross-language document optimal retrieval model performs retrieval according to a query statement in a first language to obtain a document retrieval ranking result in a second language.

9. A cross-language document retrieval system based on knowledge graph, characterized in that: The system comprises: A module for acquiring documents in a first language and documents in a second language, used for acquiring documents in a first language and documents in a second language through cross-language text data; A query document knowledge interaction graph and candidate document graph construction module, configured to construct a query document knowledge interaction graph in a first language and a candidate document graph in a second language based on document keywords in the first language and the second language; A cross-language document retrieval model construction module is used to use a graph convolutional network to perform feature encoding on the query document knowledge interaction graph and the candidate document graph to construct a cross-language document retrieval model; A cross-language document optimal retrieval model acquisition module, used to optimize the cross-language document retrieval model according to the loss function to obtain the cross-language document optimal retrieval model; The document retrieval module is used to input the query statement in the first language into the cross-language document optimal retrieval model to obtain the document retrieval ranking result in the second language.

10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.