Multi-language generative retrieval method based on cross-language semantic compression

By constructing a multilingual document retrieval dataset and performing semantic similarity clustering and dynamic multi-step constraint decoding, the semantic expression and adaptation problems of multilingual generative retrieval methods in multilingual environments are solved, achieving more efficient multilingual information retrieval.

CN120892582AActive Publication Date: 2025-11-04KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510851760.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-11-04
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing multilingual generative retrieval methods struggle to effectively express semantics and adapt across languages ​​in multilingual environments, resulting in poor retrieval performance.

Method used

By constructing a multilingual document retrieval dataset, a keyword extraction model is used to extract keywords and perform semantic similarity clustering, which are then mapped to a shared semantic space. In the inference stage, a dynamic multi-step constraint decoding method is adopted to narrow the decoding range and generate document identifiers.

Benefits of technology

It significantly improves the adaptability and generalization effect of multilingual generative retrieval models in multilingual environments, and enhances retrieval accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892582A_ABST
    Figure CN120892582A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-language generative retrieval method based on cross-language semantic compression, and belongs to the technical field of information retrieval. The method comprises the following steps: constructing a multilingual document retrieval data set; the method comprises the following steps: extracting keywords of a multilingual document from multiple angles through a keyword extraction model, calculating the extracted keywords by using semantic similarity, and constructing a similarity matrix; semantic clustering is carried out according to the similarity matrix, an atomic ID is used for representing a cluster, and then a document identifier is distributed to each multilingual document through the cluster where the keyword is located; and in the reasoning stage, after input query, gradually reducing the decoding range of the document identifier in the current step according to the decoding result of the previous step by adopting a dynamic multi-complement constraint decoding mode so as to obtain the final document identifier. Compared with other models, the retrieval capability of the method is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a multilingual generative retrieval method based on cross-language semantic compression, and belongs to the technical field of information retrieval. BACKGROUND

[0002] Multilingual Information Retrieval (MIR) aims to achieve semantic alignment between different languages, enabling users to access information resources across languages. The core of this task is to establish a semantic connection between queries and documents through language modeling or representation learning, even if they use different languages. In the early stage, MIR mainly relied on translation technology to achieve cross-language retrieval. These methods usually include translating queries into target languages, translating documents into source languages, or translating both queries and documents. Through translation, the semantic differences caused by language inconsistency can be alleviated to some extent, so as to balance the retrieval efficiency and accuracy in actual systems. With the rise of multilingual pre-training language models such as mBERT and XLM-R, the research paradigm of MIR has undergone significant changes. These models have the ability to model multilingual text uniformly, making vector representation the mainstream. The new generation of vector retrieval methods no longer rely on traditional translation steps, but encode queries and documents into vector representations in the same semantic space, and then complete retrieval by calculating the similarity between vectors. However, although vectorization methods have promoted the progress of multilingual information retrieval, they still face a series of challenges. Most current methods still follow the fixed "encoding-matching-sorting" process, which lacks overall optimization mechanism, limiting the performance of the model in actual tasks. In addition, many methods rely on contrastive learning for training, while contrastive learning usually requires a large amount of high-quality cross-language parallel data, which is extremely scarce in many languages.

[0003] As a new information retrieval paradigm, generative information retrieval (GIR) directly generates document identifiers (DocIDs) in the reasoning stage by taking full advantage of the powerful memory of pre-trained language models, rather than relying on encoding and similarity matching as traditional retrieval methods do. Currently, the representation of DocIDs can be mainly divided into two categories: atomic DocIDs and string DocIDs. Atomic DocIDs usually use randomly assigned identifiers or construct document representations based on embedding layers of document clustering. This approach pursues compact and efficient indexing, but lacks semantic interpretability. To enhance semantic expression, some studies introduce a step-by-step training strategy to learn discrete semantic representations, enabling document vectors to better capture content features. In contrast, string DocIDs use text with semantic meaning as document identifiers, such as document titles, keywords, etc. This approach not only improves the interpretability of document representation, but also facilitates the use of existing language knowledge for generative reasoning. Although GIR has made significant progress in monolingual environments, existing methods are mostly limited to English or other monolingual corpora, and their adaptability in multilingual scenarios is still insufficient. In the face of semantic diversity and cross-language inconsistency in multilingual environments, existing models often struggle to migrate directly, resulting in a significant drop in retrieval effectiveness.

[0004] Therefore, to solve this challenge, the present application proposes a multilingual generative retrieval method based on cross-language semantic compression, aiming to map lexical representations in different languages to a shared semantic space, thereby enabling effective modeling of multilingual GIR. SUMMARY

[0005] The technical problem solved by the present application is that the present application provides a multilingual generative retrieval method based on cross-language semantic compression, which clusters keywords of documents in different languages by semantic similarity and maps them to atomic IDs with the same semantics, and then assigns them to documents as document identifiers. In the reasoning stage, a dynamic multi-step constraint decoding method is used to gradually narrow the decoding range, thereby obtaining the final document identifier and improving the capability of multilingual information retrieval.

[0006] The technical solution of the present application is a multilingual generative retrieval method based on cross-language semantic compression, which comprises:

[0007] Step 1, constructing a multilingual document retrieval dataset;

[0008] Step2, extracting keywords from multi-language documents from multiple angles through a keyword extraction model, and calculating the extracted keywords using semantic similarity to construct a similarity matrix;

[0009] Step3, semantic clustering according to the similarity matrix, and using atomic ID to represent the clustering cluster, and then assigning a document identifier to each multi-language document according to the clustering cluster of the keyword;

[0010] Step4, in the reasoning stage, after inputting the query, a dynamic multi-constraint decoding method is adopted, and according to the decoding results of the previous steps, the decoding range of the current document identifier is gradually narrowed, so as to obtain the final document identifier.

[0011] Further, in Step1, the multi-language document retrieval data set is derived from Wikipedia data and related web page data, by extracting the title and multiple parts of the content in the web page, and filtering invalid information, forming a title-content pair, thereby obtaining a multi-language document retrieval data set.

[0012] Further, Step1 includes:

[0013] Step1.1, automatically obtaining Wikipedia web page content through crawler technology, using the Scarpy framework in the crawler technology to write a Python crawler program, obtaining documents of different languages by modifying the language field of the Wikipedia URL;

[0014] Step1.2, using code to construct and send a disguised web page request, injecting key information including user agent and access key in the request header, thereby successfully establishing a secure connection with the target website server;

[0015] Step1.3, after obtaining the web page source code, the page structure is parsed using the lxml.etree module in Python; then, the anchor point position of the required information is located through the XPath path extraction technology, which is used to accurately collect the target text content in the page;

[0016] Step1.4, finally, the target text content obtained is constructed into a title-content information pair, thereby obtaining a multi-language document retrieval data set.

[0017] Further, Step2 includes:

[0018] Using a large language model (LLM) based on prompts as a keyword extraction model to extract a fixed number of keywords from each multi-language document in the multi-language document set as a semantic compact representation of the multi-language document content;

[0019] The extracted keyword set is in a unified format in all languages;

[0020] Then, the keyword sets of all multilingual documents are merged into a global keyword set;

[0021] And each keyword is encoded into a dense vector using a pre-trained text encoder, and a similarity matrix between keywords is constructed by calculating the cosine similarity, which is used for subsequent semantic modeling.

[0022] Further, the Step2 specific steps include:

[0023] Step2.1, using a prompt-based large language model LLM to extract keywords from each multilingual document d i ,…} in the multilingual document set D={d i ; these keywords are a semantic compact representation of the content of the multilingual document;

[0024] Step2.2, uniformly extract a fixed number of m keywords from each multilingual document to obtain the keyword set K i of the multilingual document, and the extraction process is represented as:

[0025]

[0026] Wherein, represents the mth keyword in the multilingual document d i ;

[0027] The keyword extraction process is standardized across all multilingual documents;

[0028] Merge the keyword sets K i of all multilingual documents into a global keyword set K;

[0029]

[0030] Wherein, the global keyword set K contains non-repeating keywords, a total of n=|K|, denoted as N is the number of multilingual documents;

[0031] Step2.4, using a pre-trained text encoder to encode each keyword into a dense vector v i ∈R d , calculate the cosine similarity between any two keywords, and construct a similarity matrix S∈R n×n between keywords according to the cosine similarity score; the cosine similarity S i , v j of the dense vector v ij is calculated as follows:

[0032]

[0033] where d is the dimension of the dense vector.

[0034] Further, the Step3 comprises: setting a similarity threshold according to the similarity matrix, clustering the pairs of keywords whose similarity scores are greater than the similarity threshold, and separately clustering the keywords whose similarity scores are less than the similarity threshold, using an atomic ID to represent each cluster, and assigning a document identifier to the multi-language document according to the cluster in which the keyword is located.

[0035] Further, the Step3 comprises the following specific steps:

[0036] Step3.1, according to the similarity matrix S between the keywords constructed in Step2, when there is a path between any two keywords, and the similarity of all adjacent keywords on the path satisfies S ij ≥ θ, these keywords are clustered into a cluster, and the similarity threshold θ ∈ [0, 1]; if the similarity of a keyword with any other keyword is lower than the similarity threshold, an independent cluster is formed;

[0037] Step3.2, for the clustering results of all keywords, a unique atomic identifier, i.e. atomic ID, is assigned to each cluster, and an atomic set A = {a1, a2, …, a C} is obtained; an atom represents a group of semantically similar keywords, and the number of clusters C satisfies C < N, i.e. the number of clusters C is less than the number of documents N;

[0038] Step3.3, since each keyword belongs to a cluster, each keyword is replaced by the atomic ID of the cluster to which it belongs;

[0039] For the keyword set i of the multi-language document d , each keyword of it is mapped to the corresponding atomic ID in the atomic ID set A for representation;

[0040] Step3.4, since each multi-language document is represented by its keywords, the document identifier DocID is represented as the atomic ID set to which the keywords of the multi-language document belong, i.e.

[0041]

[0042] Each multi-language document obtains an equal-length atomic ID sequence, and forms a consistent document representation in the shared semantic space.

[0043] Further, the Step4 comprises the following steps:

[0044] Step4.1、In the model inference stage, the document identifier of a fixed length m is decoded, the document is represented as an atomic ID sequence, the decoding space is compressed into selecting in C atoms, and the size of the overall decoding space is reduced;

[0045] Step4.2、In the decoding process, the query q is input, and the model generates the target document d q The process of generating the document identifier of the document d

[0046] Wherein, a t represents the atom generated at the t step, the conditional probability P(a t |a<t,q) represents that the generation of the current step depends on the atom sequence a<t and the query q, and P(DocID|q) represents that the document identifier DocID of the target document d q is generated by taking the query q as input.

[0047] Step4.3、In each decoding step, a dynamic multi-step constraint decoding strategy is used, that is, according to the decoding result of the previous step, the decoding range of the current step is limited.

[0048] Step4.4、Through the above steps, m iterations are performed to generate the complete final document identifier.

[0049] The application also provides a multilingual generative retrieval system based on cross-language semantic compression, which comprises a module for executing the multilingual generative retrieval method based on cross-language semantic compression.

[0050] The application has the following beneficial effects:

[0051] 1、The application first constructs a multilingual document retrieval data set; then extracts document keywords using a keyword extraction model, clusters the keywords with high similarity scores by calculating the semantic similarity of all multilingual document keywords, encodes atomic IDs for each cluster, and assigns atomic IDs to multilingual documents as document identifiers; finally, in the inference stage, after inputting the query, a dynamic multi-step constraint decoding is used to gradually narrow the decoding range to obtain the final document identifier;

[0052] 2、The application aligns the semantics across languages, enabling queries and documents to be generated and matched in a unified semantic space, significantly improving the adaptability and generalization effect of the generative retrieval model in a multilingual environment;

[0053] 3、The application uses different baseline models to conduct experiments on the constructed data set, and the results show that the multilingual retrieval ability of the application is significantly improved compared with other models, and the implementation results verify the effectiveness of the application method in the multilingual retrieval task. BRIEF DESCRIPTION OF DRAWINGS

[0054] Fig. 1 For the multilingual generative retrieval model framework based on cross-language semantic compression in the present application;

[0055] Fig. 2 For the influence of the number of keywords on the performance of the model in the embodiment of the present application;

[0056] Fig. 3 For the influence of the semantic similarity threshold on the performance of the model in the embodiment of the present application. DETAILED DESCRIPTION

[0057] Embodiment 1: As shown in the multilingual generative retrieval method based on cross-language semantic compression, the method comprises: Figs. 1-3

[0058] Step 1, constructing a multilingual document retrieval data set;

[0059] Further, in Step 1, the multilingual document retrieval data set is derived from Wikipedia data and related web page data, by extracting the title and multiple parts of content in the web page, and filtering invalid information, forming a title-content pair, thereby obtaining a multilingual document retrieval data set.

[0060] Further, Step 1 comprises:

[0061] Step 1.1, automatically obtaining Wikipedia web page content through crawler technology, and obtaining documents in different languages by modifying the language field of the Wikipedia URL;

[0062] Specifically, first, take a Wikipedia page in a certain language as a basic template, modify the language field in the Wikipedia URL through a program, access the page content of the same topic under different language versions, and thus construct a multilingual content alignment corpus; the language field includes "fr", "sv", "mk";

[0063] Step 1.2, in order to improve the success rate and stability of web page requests, use code to construct and send disguised web page requests, inject user agent (User-Agent), access key and other key information in the request header, and thus successfully establish a secure connection with the target website server;

[0064] ​Specifically, an automated script written in Python is used to construct and send web page requests simulating real user behavior; specifically including: injecting customized fields in the request header (HTTP Headers), such as User-Agent (simulating different browser types), Referer (source page address), Accept-Language (accepted language range) and Cookies and access keys, etc. information, simulating legal client access scenarios;

[0065] Step1.3, after obtaining the web page source code, the lxml.etree module in Python is used to perform DOM parsing on the HTML document structure;

[0066] Subsequently, the label hierarchy relationship of the web page is analyzed, and the anchor point position of the required text content is located through XPath path extraction technology, such as title, paragraph, hyperlink and other information fields, which are used to accurately collect the target text content in the page;

[0067] The lxml.etree module has high-efficiency structured data extraction capability and is suitable for deep information extraction of complex pages.

[0068] In order to ensure the consistency of multi-language page information extraction, a general and robust XPath extraction logic needs to be designed according to the actual page structure to ensure the extraction accuracy of the target text in different language versions.

[0070] Step1.4, finally, the title-content information pair of the obtained target text content is constructed to obtain the multi-language document retrieval data set.

[0071] Each language document data set includes title, abstract and content when constructed;

[0072] Step2, extract keywords from multi-language documents from multiple angles through a keyword extraction model, and calculate the extracted keywords using semantic similarity to construct a similarity matrix;

[0073] Further, the Step2 includes:

[0074] A large language model (LLM) based on prompts is used as a keyword extraction model to extract a fixed number of keywords from each multi-language document in the multi-language document set as a semantic compact representation of the multi-language document content;

[0075] The extracted keyword set is in a unified format in all languages;

[0076] Then, the keyword sets of all multi-language documents are combined into a global keyword set;

[0077] and use a pre-trained text encoder to encode each keyword into a dense vector, calculate the cosine similarity to construct the similarity matrix between keywords, which is used for subsequent semantic modeling.

[0078] Further, the Step2 specific steps include:

[0079] Step2.1, using a prompt-based large language model LLM to extract keywords from a multi-lingual document set D = {d1, d2, …, d i , …} to capture its key semantic information. These keywords are a semantic compact representation of the content of the multi-lingual document; i

[0080] Step2.2, uniformly extract a fixed number of m keywords from each multi-lingual document to obtain a keyword set K i for the multi-lingual document. The extraction process is represented as:

[0081]

[0082] wherein, represents the mth keyword in the multi-lingual document d i ; the semantic difference between keywords is large to enhance the ability to express multi-angle semantics;

[0083] The keyword extraction process is standardized across all multi-lingual documents to ensure consistency in semantic representation;

[0084] Combine the keyword sets K i from all multi-lingual documents into a global keyword set K;

[0085]

[0086] wherein the global keyword set K contains no duplicate keywords, a total of n = |K|, denoted as N is the number of multi-lingual documents;

[0087] Step2.4, use a pre-trained text encoder to encode each keyword into a dense vector v i ∈R d , calculate the cosine similarity between any two keywords, and construct a similarity matrix S ∈ R n×n between keywords according to the cosine similarity score; the cosine similarity S i between dense vectors v j , v ij is calculated as follows:

[0088]

[0089] wherein d is the dimension of the dense vector.​

[0090] Step3、According to the similarity matrix, semantic clustering is performed, and the clustering clusters are represented by atomic IDs, and then the document identifiers are assigned to each multi-language document according to the clustering clusters where the keywords are located;

[0091] Further, the Step3 includes: according to the similarity matrix, a similarity threshold is set, pairs of keywords with a similarity score greater than the similarity threshold are clustered, and those less than the similarity threshold are individually clustered, each cluster is represented by an atomic ID, and a document identifier is assigned to a multi-language document according to the cluster where the keyword is located.

[0092] Specifically, based on the similarity matrix S between keywords, keywords with a cosine similarity equal to or higher than a similarity threshold θ are clustered into the same cluster through a connected path, and those that do not meet the condition are individually clustered to form semantic clustering.

[0093] Each clustering cluster is assigned a unique atomic identifier (Atom ID), i.e., an atomic ID, to form an atomic set A = {a1, a2, …, a C}, as a semantic unit, where a C is an atom, and C is the number of clusters.

[0094] Subsequently, each keyword in the multi-language document is mapped to its corresponding atomic ID representation, so that each multi-language document is represented by a sequence of atomic IDs corresponding to its keywords as is the atomic ID corresponding to the mth keyword of the ith multi-language document. This process realizes the equal-length consistent representation of cross-language documents in a unified semantic space.

[0095] Further, the specific steps of Step3 include:

[0096] Step3.1, according to the similarity matrix S between keywords constructed in Step2, when there is a path between any two keywords, and the similarity of all adjacent keywords on the path satisfies S ij ≥ θ, these keywords are clustered into a cluster, and the similarity threshold θ ∈ [0, 1]; if the similarity of a keyword with any other keyword is lower than the similarity threshold, then an independent cluster is formed.

[0097] Step3.2, for the clustering results of all keywords, a unique atomic identifier, i.e., an atomic ID, is assigned to each clustering cluster, and an atomic set A = {a1, a2, …, a C} is obtained; an atom represents a group of semantically similar keywords, and the number of clusters C satisfies C < N, i.e., the number of clusters C is less than the number of documents N.

[0098] Step3.3、Since each keyword belongs to a cluster, each keyword is replaced by the atomic ID of its belonging cluster;

[0099] For each multi-lingual document d i , its keyword set K(d ) is mapped to the corresponding atomic ID set A(d ) by mapping each keyword k to the atomic ID of its belonging cluster;

[0100] Step3.4、Since each multi-lingual document is represented by its keywords, the document identifier DocID is represented by the atomic ID set that the keywords of the multi-lingual document belong to, i.e.

[0101]

[0102] is the atomic ID corresponding to the m-th keyword of the i-th multi-lingual document,

[0103] Each multi-lingual document obtains an equal-length atomic ID sequence, forming a consistent document representation in the shared semantic space.

[0104] Step4、In the inference stage, after inputting the query, a dynamic multi-restriction decoding method is adopted to gradually narrow down the decoding range of the current document identifier, so as to obtain the final document identifier.

[0105] Further, the Step4 includes the following steps:

[0106] Step4.1、In the model inference stage, the retrieval process of the document is modeled as a multi-step decoding process with length m. The traditional method needs to select from all N documents at each step, and the search space is O(N m ); while our method represents the document as an atomic ID sequence, and compresses the decoding space to select from C semantic atoms, so that the overall decoding space is reduced to O(C m );

[0107] Step4.2、In the decoding process, the query q is input, and the process of generating the document identifier DocID of the target document d q is described as follows:

[0108] where a t represents the atomic generated at the t-th step, the conditional probability P(a t |a<t,q) represents that the generation at the current step depends on the previously generated atomic sequence a <t and the query q, and P(DocID|q) represents the generation of the target document dq a document identifier DocID;

[0109] Step4.3, in each decoding step, a dynamic multi-step constraint decoding strategy is used, that is, according to the decoding result of the previous step, the decoding range of the current step is limited;

[0110] Specifically, the decoding of each step t adopts a dynamic multi-step constraint mechanism, that is, according to the generated prefix, that is, the document identifier DocID=[a1, a2, …, a t-1 ] dynamically adjust the current available atomic ID set A t :

[0111]

[0112] Wherein, a t-1 represents the atom generated in the t-1 step, Constraint(K i ) represents the keyword set that meets the condition under the current prefix restriction, which corresponds to the current available atomic ID set A t ;

[0113] The optimal solution of the current step is:

[0114]

[0115] Wherein, is the probability of predicting the decoding result of the current step based on the query and the decoding result of the previous step.

[0116] Step4.4, through the above steps Step4.1-Step4.3, m iterations are performed to generate a complete final document identifier:

[0117] DocID=[a1, a2, …, a m ].

[0118] Table 1 is a dynamic multi-step constraint decoding pseudo code

[0119]

[0120] The application also provides a multilingual generative retrieval system based on cross-language semantic compression, which comprises:

[0121] A first construction module is used for constructing a multilingual document retrieval data set;

[0122] A second construction module is used for extracting keywords of the multilingual document from multiple angles through a keyword extraction model, and calculating the extracted keywords using semantic similarity to construct a similarity matrix;

[0123] The distribution module is configured to perform semantic clustering according to the similarity matrix, and represent the clustering clusters by using atomic IDs, and then assign a document identifier to each multilingual document according to the clustering cluster in which the keyword is located.

[0124] The obtaining module is configured to, in the reasoning stage, input a query, and then adopt a dynamic multi-constraint decoding manner to gradually narrow the decoding range of the current document identifier according to the decoding result of the previous step, so as to obtain a final document identifier.

[0125] The above method proposed in the application is experimented on a multilingual document retrieval data set, and the number of documents in each language is divided as follows:

[0126] Table 2 is the language distribution of the multilingual document retrieval data set

[0127]

[0128] In the application, all methods are based on the Transformer architecture and are reproduced on the mT5-base model, and the training data of the application adopts the strategy of representing documents using multilingual queries. Specifically, the model used in all pseudo query generation tasks is Llama3.1-8B, and the temperature during generation is set to 0.7. Each document sample generates 10 multilingual pseudo queries through the model, which are used to construct training data. In the keyword generation task, the Llama3.1-8B model is still used, but the temperature is set to 0 to ensure the determinacy and consistency of the generated content. In the semantic similarity calculation, the paraphrase-multilingual-MiniLM-L12-v2 model is used, which can effectively handle multilingual semantic matching tasks.

[0129] In terms of training settings, the application uses the PyTorch and Transformers framework for implementation. For the multilingual document retrieval data set, the learning rate is set to 5×10 -4 , the batch size is 128, the training rounds are 50 rounds, and m=3 keywords are selected. The training objective function adopts a cross-entropy loss function. All experiments are performed on eight NVIDIA A40 graphics cards, each with 46GB of video memory, ensuring efficient execution of large-scale training tasks.

[0130] In terms of evaluation metrics, the present application follows the settings of existing research and evaluates model performance on the validation set of the multilingual document retrieval dataset. The main evaluation metrics are Recall@1 and Recall@10, which represent the proportion of relevant documents in the top 1 and top 10 search results, respectively. Considering that the candidate document corpus consists of multiple languages, we report the search results for each query language separately during the evaluation process to more comprehensively reflect the performance of the model in a multilingual search environment.

[0131] To verify the effectiveness of the method proposed in the present application, the multilingual retrieval method related to the present application is selected as the baseline model, which is BM25, mDPR, DSI, DSI-QG, and SE-DSI, as follows:

[0132] BM25 is a classic sparse retrieval method based on inverted index structure, which relies on exact matching of keywords for retrieval operation and is widely used in traditional information retrieval systems.

[0133] mDPR is a dense retrieval method that uses a dual-encoder architecture to encode queries and documents into vector representations and perform similarity matching in vector space. In the present application, LaBSE is used as its encoder to enhance multilingual representation capabilities.

[0134] DSI is a generative retrieval method that takes the document itself as training input and constructs document IDs (DocIDs) through hierarchical clustering to achieve a retrieval process without explicit indexing.

[0135] DSI-QG is an improved generative retrieval method that trains a query generation model to convert original documents into short pseudo queries for modeling, while using random numbers to represent corresponding document IDs to enhance the generalization ability of the model.

[0136] SE-DSI also belongs to the generative retrieval paradigm, and its characteristic is to use a string containing semantic information as the document ID. In the present application, the title of the multilingual document is used as the DocID to improve the semantic richness and cross-language consistency of document representation.

[0137] Tables 3 and 4 show that on the validation set, BM25 and DSI have relatively limited retrieval capabilities; while mDPR maintains a moderate level of performance, it is quite sensitive to language variations. DSI-QG and SE-DSI show improvements in cross-linguistic recall balance compared to previous methods, but their performance is still affected by the richness of language resources. In contrast, this invention achieves stable and high recall rates across all languages ​​in the multilingual document retrieval dataset, showing a significant improvement, especially in languages ​​with low to medium resource availability. Recall@1 is improved by 2.79% compared to other methods, and recall@10 is improved by 5.01%. This demonstrates that the invention exhibits good generalization ability across language families with similar and significantly different language types.

[0138] Table 3 shows the Recall@1 performance comparison between the present invention and the baseline model.

[0139]

[0140]

[0141] Table 4 shows the Recall@10 performance comparison between the present invention and the baseline model.

[0142]

[0143] To further validate the multilingual generative retrieval method proposed in this invention, hyperparameter experiments were designed to investigate the impact of semantic similarity threshold θ and the number of keywords m per document on model performance, as shown below:

[0144] To investigate the impact of the number of keywords *m* on retrieval performance and decoding time, we conducted experiments on a multilingual document retrieval dataset. For example... Fig. 2 As shown, increasing the number of keywords can improve semantic representation capabilities, but it also increases the length of the DocID, leading to slower decoding speed and a decrease in overall performance. The optimal Recall@10 is 48.06%, corresponding to the use of three keywords. Fewer keywords (such as two) limit semantic coverage; while more keywords (such as five or six) reduce performance due to increased decoding complexity. Decoding time is positively correlated with DocID length. When the number of keywords increases from two to six, the decoding time approximately doubles. In comparison, SE-DSI has a Recall@10 of 29.96, and its latency is comparable to our method when using six keywords, further illustrating the strong correlation between sequence length and inference time.

[0145] Furthermore, to evaluate the sensitivity of keyword clustering to the semantic similarity threshold θ, we assessed retrieval performance on a multilingual document retrieval dataset with a threshold range of 0.5 to 0.9. Fig. 3As shown, when the threshold is raised from 0.5 to 0.8, both Recall@1 and Recall@10 steadily increase, indicating that more refined clustering helps improve retrieval accuracy. The best performance is achieved when the threshold is 0.8, with Recall@1 of 24.48% and Recall@10 of 47.89%. However, when the threshold is further raised to 0.9, the performance decreases, with Recall@1 dropping to 22.02% (a decrease of 2.46%) and Recall@10 dropping to 44.71% (a decrease of 3.18%). This indicates that an excessively high threshold can cause the clustering to be too refined, dividing semantically related keywords into different groups, thereby reducing the semantic generalization ability and limiting the model's retrieval coverage.

[0146] The specific embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.

Claims

1. A multilingual generative retrieval method based on cross-lingual semantic compression, characterized in that: The method includes: Step 1: Construct a multilingual document retrieval dataset; Step 2: Extract keywords from multilingual documents from multiple perspectives using a keyword extraction model, and use semantic similarity to calculate the extracted keywords and construct a similarity matrix; Step 3: Perform semantic clustering based on the similarity matrix, and represent the clusters using atomic IDs. Then, assign a document identifier to each multilingual document based on the cluster where the keyword belongs. Step 4: In the reasoning stage, after inputting the query, a dynamic multi-complement constraint decoding method is adopted. Based on the decoding results of the previous steps, the decoding range of the document identifier in the current step is gradually narrowed to obtain the final document identifier.

2. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: In Step 1, the multilingual document retrieval dataset is derived from Wikipedia data and related web page data. By extracting titles and multiple parts of content from the web pages and filtering out invalid information, title-content pairs are formed, thereby obtaining the multilingual document retrieval dataset.

3. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: Step 1 includes: Step 1.1: Automatically obtain Wikipedia webpage content using web crawling technology; obtain documents in different languages ​​by modifying the language field of the Wikipedia URL. Step 1.2: Use code to construct and send a fake web page request, injecting key information including user agent and access key into the request header, thereby successfully establishing a secure connection with the target website server; Step 1.3: After obtaining the webpage source code, use the lxml.etree module in Python to parse the page structure; then, use XPath path extraction technology to locate the anchor points of the required information for accurate collection of target text content in the page. Step 1.4: Finally, construct title-content information pairs from the obtained target text content to obtain a multilingual document retrieval dataset.

4. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: Step 2 includes: Using a prompt-based large language model (LLM) as a keyword extraction model, a fixed number of keywords are extracted from each multilingual document in the multilingual document collection as a semantically compact representation of the multilingual document content. The extracted keyword set has a unified format across all languages; Then, the keyword sets of all multilingual documents are merged into a global keyword set; Each keyword is encoded into a dense vector using a pre-trained text encoder, and cosine similarity is calculated to construct a similarity matrix between keywords for subsequent semantic modeling.

5. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: The specific steps of Step 2 include: Step 2.1: Use a prompt-based Large Language Model (LLM) to process a collection of multilingual documents D = {d1, d2, ..., d...} i Extract each multilingual document d from ,…} i Keywords; these keywords serve as a semantically compact representation of the content of multilingual documents; Step 2.2: Extract a fixed number of m keywords from each multilingual document to obtain the keyword set K of the multilingual document. i The extraction process is represented as follows: in, Indicates a multilingual document d i The m-th keyword in; The keyword extraction process is performed in a standardized manner across all multilingual documents; The set of keywords K for all multilingual documents i Merge into a global keyword set K; The global keyword set K contains n = |K| unique keywords, denoted as K. N represents the number of multilingual documents; Step 2.4: Use a pre-trained text encoder to encode each keyword into a dense vector v. i ∈R d Calculate the cosine similarity between any two keywords, and construct a keyword similarity matrix S∈R based on the cosine similarity scores. n×n Dense vector v i v j cosine similarity S ij The calculation formula is as follows: Where d is the dimension of the dense vector.

6. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: Step 3 includes: setting a similarity threshold based on the similarity matrix; clustering pairs of keywords with similarity scores greater than the similarity threshold; grouping keywords with similarity scores less than the similarity threshold into a single cluster; using atomic IDs to represent each cluster; and assigning document identifiers to multilingual documents based on the cluster to which the keywords belong.

7. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: The specific steps in Step 3 include: Step 3.1: Based on the keyword similarity matrix S constructed in Step 2, when there is a path between any two keywords, the similarity of all adjacent keywords on that path satisfies S. ij When the similarity threshold is ≥θ, these keywords are clustered into a cluster, with the similarity threshold θ∈[0,1]; if the similarity of a keyword to any other keyword is lower than the similarity threshold, then an independent cluster is formed. Step3.

2. Assign a unique atomic identifier, i.e., atomic ID, to the clustering results of all keywords to obtain an atomic set A = {a1, a2, …, a C}; An atom represents a group of keywords with similar semantics, and the number of clusters C satisfies C < N, that is, the number of clusters C is less than the number of documents N; Step 3.3: Since each keyword belongs to a cluster, each keyword... Replace it with the atomic ID of its cluster; For multilingual documents d i Keyword set Each keyword Map the corresponding atomic ID representation in the atomic ID set A; Step 3.4: Since each multilingual document is represented by its keywords, the document identifier DocID represents the set of atomic IDs to which the keywords of the multilingual document belong, i.e. For the m-th keyword in the i-th multilingual document, Each multilingual document receives a sequence of atomic IDs of equal length, forming a consistent document representation in a shared semantic space.

8. The multilingual generative retrieval method based on cross-lingual semantic compression according to claim 1, characterized in that: Step 4 includes the following steps: Step 4.1: In the model inference stage, decode the document identifier with a fixed length of m, represent the document as an atomic ID sequence, compress the decoding space into selection from C atoms, and reduce the overall size of the decoding space. Step 4.2: During the decoding process, input query q, and the model generates target document d. q The process of document identifiers is described as follows: where, a t represents the atom generated at the t-th step, and the conditional probability P(a t |a<t,q) indicates that the generation at the current step depends on the previously generated atomic sequence a <t and the query q, and P(DocID|q) represents generating the target document d q with the query q as the input and the document identifier DocID of the document; Step 4.3: In each decoding step, a dynamic multi-step constraint decoding strategy is used, that is, the decoding range of the current step is limited according to the decoding result of the previous step. Step 4.4: Perform m iterations through the above steps to generate the complete final document identifier.

9. A multilingual generative retrieval system based on cross-lingual semantic compression, characterized in that, The system includes a module for executing the multilingual generative retrieval method based on cross-lingual semantic compression as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cross-language question answering system construction method and device based on generative multi-language model

    CN115795009A

  • Monolingual or multilingual document detection method

    CN116089632A