A dynamic alignment anchor point based multilingual embedded semantic retrieval method
Patent Information
- Application Number
- CN202610960536.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
传统的关键词匹配或基于规则的机器翻译检索方法,存在显著局限性:一方面,不同语言表达同一语义时词汇与句式差异巨大,导致关键词召回率低;另一方面,先翻译后检索的两阶段方案不仅引入翻译误差,且难以处理一词多义、文化语境差异等复杂语义问题,严重影响检索结果的准确性与相关性
[0009]本公开实施例的有益效果在于:通过动态语义锚点机制,在检索阶段动态生成符合当前检索意图和检索领域的锚点,实时反映当前查询的局部语义分布,实现检索阶段的领域与意图自适应对齐,显著提升跨语言语义检索的准确率与相关性;并且通过基于领域标签确定自适应对齐强度并融合语义引导向量,能够在领域置信度较高时增强锚点对齐效果,在置信度不足时自动降低对齐强度,有效避免语义过度偏移与失真,从向量几何层面保证检索结果的稳定性与稳健性;最后基于对齐查询向量召回候选文档并进行语义校正排序,能够进一步缓解跨语言检索中的歧义问题,使最终检索结果在语义层面与查询真实意图高度一致。
Smart Images

Figure CN122817433A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of information retrieval technology, and in particular to a multilingual embedded semantic retrieval method, storage medium, and electronic device based on dynamic alignment anchors. Background Technology
[0002] With the acceleration of globalization and the continuous expansion of enterprises' international business, the demand for cross-language search (CLIR) is growing daily. Enterprises often need to quickly and accurately obtain semantically relevant content from massive amounts of multilingual documents (such as patent documents, technical standards, contract texts, customer service tickets, product manuals, etc.). Traditional keyword matching or rule-based machine translation retrieval methods have significant limitations: on the one hand, different languages express the same meaning with huge differences in vocabulary and sentence structure, resulting in low keyword recall rates; on the other hand, the two-stage approach of translating first and then retrieving not only introduces translation errors but also struggles to handle complex semantic issues such as polysemy and cultural context differences, seriously affecting the accuracy and relevance of search results.
[0003] In recent years, many studies have proposed multilingual embedding models, such as Qwen3-Embedding and jina-embeddings-v4, attempting to construct a unified cross-lingual semantic space so that semantically equivalent texts in different languages are close to each other in the vector space. However, existing technologies still have the following problems: (1) Semantic alignment staticization: Existing cross-language embedding models usually achieve cross-language alignment through pre-training, which constructs a fixed and universal semantic space. The language distribution covered depends on the data input to the model during pre-training, and there is still the problem of language clustering, that is, the same language is still close to each other in the same semantic space. (2) Lack of adaptive capability during retrieval: Existing solutions only perform "encoding-comparison" operations during the reasoning stage, and cannot make real-time corrections based on the differences between the query and candidate documents (such as cultural metaphors and terminological ambiguities), thus failing to effectively alleviate the language and semantic gap. (3) High model update cost: When new language or domain data is added, the entire model often needs to be retrained or fine-tuned, resulting in high maintenance costs and long iteration cycles, which makes it difficult to meet the needs of enterprises to respond quickly to business changes.
[0004] Therefore, there is an urgent need for a technical solution that can ensure the accuracy of cross-language semantic retrieval while also taking into account system efficiency, scalability, and adaptability, in order to solve the bottleneck problems of existing technologies in practical applications. Summary of the Invention
[0005] The purpose of this disclosure is to provide a multilingual embedded semantic retrieval method, storage medium, and electronic device based on dynamic alignment anchors, in order to solve the problems existing in the prior art.
[0006] The embodiments of this disclosure adopt the following technical solution: a multilingual embedded semantic retrieval method based on dynamic alignment anchors, comprising: preprocessing the original query text input by the user to output standardized query text, the target language type of the original query text, query domain tags, and query keywords; determining a target adapter according to the target language type, merging the target adapter with the backbone network to form an embedding model, vectorizing the standardized query text based on the embedding model to obtain an original query vector; dynamically generating a dynamic anchor set for the current session round based on the original query vector and the query domain tags; calculating the attention weight between the original query vector and each dynamic anchor in the dynamic anchor set, weighting and aggregating the dynamic anchors according to the attention weight to obtain a semantic guidance vector, and determining an alignment strength coefficient according to the query domain tags; determining an aligned query vector based on the alignment strength coefficient, the semantic guidance vector, and the original query vector; recalling a candidate document vector set based on the aligned query vector, and sorting all candidate document vectors in the candidate document vector set based on semantic correction to obtain the query results.
[0007] This disclosure also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described multilingual embedded semantic retrieval method based on dynamic alignment anchors.
[0008] This disclosure also provides an electronic device, including at least a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the above-described multilingual embedded semantic retrieval method based on dynamic alignment anchors.
[0009] The beneficial effects of this disclosure are as follows: Through a dynamic semantic anchor mechanism, anchors that conform to the current search intent and search domain are dynamically generated during the retrieval phase, reflecting the local semantic distribution of the current query in real time. This achieves adaptive alignment between the domain and intent during the retrieval phase, significantly improving the accuracy and relevance of cross-language semantic retrieval. Furthermore, by determining the adaptive alignment strength based on domain labels and fusing semantic guidance vectors, the anchor alignment effect can be enhanced when the domain confidence is high, and the alignment strength can be automatically reduced when the confidence is insufficient, effectively avoiding excessive semantic shift and distortion, and ensuring the stability and robustness of the retrieval results from a vector geometry perspective. Finally, by recalling candidate documents based on the aligned query vector and performing semantic correction and ranking, the ambiguity problem in cross-language retrieval can be further alleviated, ensuring that the final retrieval results are highly consistent with the true query intent at the semantic level. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of the multilingual embedded semantic retrieval method based on dynamic alignment anchors in the first embodiment of this disclosure. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0013] To address the problems existing in the prior art, the first embodiment of this disclosure provides a multilingual embedded semantic retrieval method based on dynamically aligned anchor points. This method can be implemented based on a retrieval system or retrieval platform with a user interface. Figure 1 The flowchart of the multilingual embedded semantic retrieval method of this embodiment is shown, which mainly includes steps S10 to S50: S10: Preprocess the original query text input by the user and output the standardized query text, the target language type of the original query text, the query domain label, and the query keywords.
[0014] Users input their original query text in natural language to represent their search purpose. The original query text can be presented in any language. After receiving the original query text, the system standardizes and structures it to minimize the impact of different languages, different expression habits, and noise on subsequent semantic modeling.
[0015] Specifically, the preprocessing implemented in this embodiment first performs language recognition processing on the original query text to determine the target language type. Language recognition can employ methods based on character distribution and statistical features, determining the language type by statistically analyzing the distribution ratio of different character sets; alternatively, it can use methods based on lightweight language recognition models, such as FastText, n-gram language models, or small Transformer classification models, to classify the query text into languages. For mixed language queries, the primary language type and auxiliary language tagging information are output. The language recognition results are used to subsequently select the corresponding target adapter and provide a basis for word segmentation and semantic analysis.
[0016] After language recognition, the original query text undergoes text normalization and cleaning, including at least removing special symbols, standardizing character formats, standardizing numerical and date units of measurement, and filtering stop words to obtain standardized query text. Specifically, this may include removing or replacing special symbols, control characters, HTML tags, and abnormally encoded characters that have no semantic value; standardizing the representation of full-width and half-width characters and regulating uppercase and lowercase letter forms; standardizing the representation of numbers, dates, and units of measurement using rules or dictionaries; and marking or filtering redundant words or stop words that have no semantic contribution. The stop word list can be configured separately for each language. Through text normalization, vector bias caused by the same semantic meaning in different written forms is reduced.
[0017] This embodiment also includes word segmentation and sub-word segmentation of the standardized query text. Appropriate word segmentation strategies are implemented for different language characteristics. For languages without explicit delimiters, such as Chinese and Japanese, a word segmentation method combining dictionary and statistical models is used, or a sub-word segmentation algorithm consistent with the embedding model, such as BPE or SentencePiece, is directly employed. For English and other languages separated by spaces, regular word segmentation is used, further combined with sub-word segmentation to handle out-of-vocabulary words. The word segmentation operation in this embodiment must ensure that the segmentation method used is consistent with the tokenizer of the subsequent embedding encoding model to avoid semantic inconsistencies. The output of this step can simultaneously retain word-level and sub-word-level information, providing support for subsequent semantic enhancement.
[0018] After basic text processing, preliminary semantic intent and domain identification are performed on the query statements to determine which business scenario the query is more inclined towards, and the corresponding query domain label is output. Domain identification can be achieved by using keyword and phrase matching methods to map the query to a predefined set of domains such as patents, law, medicine, and engineering technology. Alternatively, it can be achieved by using similarity calculation methods between query vectors and domain prototype vectors, selecting the domain with the highest similarity as the query domain label, which serves as important prior information for subsequent dynamic semantic anchor point selection.
[0019] Finally, key query terms are extracted from the standardized query text. Specifically, term frequency-inverse document frequency (TF-IDF) can be used to filter out distinguishable terms that are high-frequency in the current query and low-frequency in the corpus; or NER entity recognition can be used to identify structured entities such as technical terms and personal names in the query text through a domain-pre-trained model, or a pre-built professional terminology database can be used for matching and extraction.
[0020] After completing the above processing steps, a structured query representation result is output. This result includes at least the standardized query text, target language type, preliminarily identified query domain labels, and query keywords. This structured query result will be passed as input to the embedding model and dynamic anchor generation step for subsequent embedding representation and semantic alignment processing.
[0021] S20: Determine the target adapter based on the target language type, merge the target adapter with the backbone network to form an embedding model, and vectorize the standardized query text based on the embedding model to obtain the original query vector.
[0022] To reduce resource consumption and enable flexible expansion, this embodiment employs a frozen backbone combined with a pluggable adapter architecture for the embedding model used for vectorization. The backbone network is a pre-trained multilingual Transformer model with its parameters frozen. The adapter is a lightweight adapter trained using LoRA fine-tuning techniques. Its base model can be a vectorized model such as Qwen3-Embedding or multilingual-e5. By combining training samples of different target language types, the parameters are fine-tuned and optimized to form adapters corresponding to different languages.
[0023] Specifically, in the process of fine-tuning the adapter, the training sample data is prepared first. High-resource languages such as Chinese and English are selected, and a labeled dataset is constructed in the form of query statement-label pairs, with labels that fit the core semantic intent of the business scenario. Then, the labeled dataset is split into dataset N1 for building the retrieval index and dataset N2 for generating comparative learning samples. N1 is constructed by stratified sampling of the labeled dataset, ensuring that the label distribution in N1 is consistent with the label distribution in the original labeled dataset. Then, positive and negative sample pairs are constructed. Positive samples are generated by translating the query statement of N2 and matching queries with the same label in N1, while negative samples are generated by combining random negative examples (i.e., the same query statement but different labels) and weighted sampling of difficult negative examples (i.e., similar labels but different query statements). Finally, a large language model is used to generate semantically consistent positive examples and semantically different negative examples for the target language query, covering diverse expressions. After selecting the base model of the adapter, the prepared positive and negative sample pairs are used, with InfoNCE Loss as the objective function, to minimize the positive sample similarity loss and maximize the positive and negative sample difference, allowing the model to learn a cross-language semantically consistent vector representation.
[0024] In the actual retrieval process, for the target language type of the current query text output in step S10, the system can load only the adapter for the corresponding language through the model service framework according to the target language type of the current query, so that only the adapter for the active language exists in memory at the same time, which significantly reduces the peak memory usage; and when adding a new target language, the parameters can be trained separately for the adapter without training the backbone network, which significantly reduces the training cost and facilitates expansion.
[0025] Finally, the adapter corresponding to the current target language is combined with the backbone network to form an embedding model. Specifically, the merging refers to superimposing the low-rank weight matrix obtained by training the target adapter into the original pre-trained weight matrix of the backbone network during the inference stage, so as to achieve linear addition of weights. This results in an embedding model with specific language adaptation capabilities without increasing the number of model inference parameters and computational latency. The standardized query text is then vectorized using this embedding model to obtain the original query vector.
[0026] S30: Based on the original query vector and query domain labels, dynamically generate a set of dynamic anchor points for the current session round.
[0027] Subsequently, in the retrieval phase, this embodiment utilizes dynamic anchor generation for dynamic correction at the domain and semantic intent levels. Dynamic anchors are language-independent semantic anchors, which are dynamically adjustable semantic reference centers constructed in a unified vector space, thereby guiding the query vector to align with the target domain or target intent distribution. In this embodiment, dynamic anchors are high-dimensional vectors in a shared semantic vector space. Unlike synonyms or keyword sets, semantic anchors are not discrete terms, but rather continuous vector representations generated by the embedding model in step S20, inherently possessing language independence and contextual abstraction capabilities.
[0028] During the retrieval phase, a dynamic anchor set for the current session round is dynamically generated based on the original query vector and query domain tags to improve adaptability to the specific semantic context of the current query. First, an initial retrieval is performed in the vector database based on the original query vector to obtain a set of candidate document vectors. Then, the candidate document vector set is weighted and aggregated to generate the first dynamic semantic anchor. as follows:
[0029] Among them, weight Indicates the first The relevance of each candidate document to the query text is specifically obtained in this embodiment by calculating the similarity between the original query vector and the candidate document vector, and represented by softmax normalization. The first dynamic semantic anchor point obtained based on the above formula (1) can reflect the local center of the current query in the actual candidate semantic distribution.
[0030] Subsequently, based on the query domain tags, a set of descriptive text corresponding to the query domain tags is selected. The source of this descriptive text set can be pre-built structured business knowledge base content such as domain encyclopedias, terminology explanations, and technical white papers; it can also be expert-annotated texts such as descriptions of typical query intents and scenario descriptions written by domain experts; it can be high-confidence search logs such as frequently clicked and long-stayed documents from historical searches; or it can be descriptive text generated based on domain tags using a large language model and manually reviewed. Inputting the descriptive text set into the embedding model yields a set of descriptive text vectors. Averaging and aggregating these descriptive text vectors yields the second dynamic semantic anchor point based on the domain descriptive text. This is used to represent the center of the current semantic query in the semantics of a clearly defined business domain.
[0031] After determining the first and second dynamic semantic anchors, they can be integrated to form the dynamic anchor set for the current session round. In some embodiments, the system can pre-build a language-independent semantic anchor library to store semantic anchor vectors corresponding to different domains or different search intentions. The construction methods include: building anchors by business domain, such as anchors for patent technology fields, legal clauses, and medical knowledge; and building anchors by search intention, such as anchors for technical solution search intentions, background technology search intentions, and risk analysis search intentions. The anchor library supports a maintenance method combining offline construction and online updates: in the offline stage, initial anchors are built based on high-quality corpora; in the online stage, anchor vectors are smoothly updated or their weights adjusted based on user feedback. At this point, the similarity of each semantic anchor vector in the original query vector can be calculated in the anchor library. One or more semantic anchors with the highest similarity are selected as dynamic anchors in the dynamic anchor set of the current session round. This allows for precise location of the semantic anchor closest to the current query semantics. Combined with the first and second dynamic semantic anchors, this achieves differentiated fusion of multiple anchors, avoiding the one-sidedness of a single anchor from affecting the accuracy and robustness of subsequent semantic guidance.
[0032] In some embodiments, corresponding to multi-round session operations, a composite semantic anchor can be generated by combining query vectors from historical session rounds and query vectors from the current session round. Assume the set of historical query vectors is: Then the compound semantic anchor point It can be represented as:
[0033] in, This represents the time decay weight, whose value is inversely proportional to the value of t. It is used to enhance the impact of the query semantics of the current session round on anchor point generation. After determining the composite semantic anchor point, it can also be added to the dynamic anchor point set as the basis for subsequent guidance vector generation.
[0034] It is important to note that the dynamically generated semantic anchors in this embodiment are only valid within the current retrieval period and are discarded after the current retrieval ends. In actual use, if the dynamic semantic anchors meet the stability conditions, they can be written into the anchor library for long-term storage. The stability conditions include, but are not limited to, the following: being selected repeatedly in multiple independent query samples; producing a continuous and positive effect on the ranking of search results; and having a similarity with existing anchors that does not exceed a preset merging threshold. Through the above update constraints, frequent fluctuations or semantic drift in the anchor library are prevented, ensuring the long-term stable evolution of the anchor system.
[0035] S40, calculate the attention weight between the original query vector and each dynamic anchor in the dynamic anchor set, perform weighted aggregation on the dynamic anchors according to the attention weight to obtain the semantic guidance vector, determine the alignment strength coefficient according to the query domain label, and determine the alignment query vector according to the alignment strength coefficient, the semantic guidance vector and the original query vector.
[0036] Determining the set of dynamic anchor points Then, anchor-based alignment can be performed on the original query vector. First, the original query vector is calculated. Attention weights between each dynamic semantic anchor :
[0037] in, The vector similarity calculation function is represented, preferably cosine similarity. Then, the semantic anchors are weighted and aggregated to obtain the semantically guided vector. :
[0038] In this embodiment, the semantic guidance vector is a weighted aggregation representation of multiple dynamic anchors. It centrally reflects the semantic center direction of the current query at the domain and intent levels, while controlling the direction and degree of the original query vector's offset towards the target semantic space during the subsequent alignment process, and enhancing the adaptability to complex queries.
[0039] To avoid semantic jumps caused by low-quality or noisy anchors, this embodiment introduces a reliability screening mechanism for semantic anchors participating in alignment, assuming the original query vector... With dynamic semantic anchors Similarity between :
[0040] Only when At that time, the anchor point Only then were they included in the calculation of the semantic guidance vector, among which, This is the anchor point similarity threshold, used to filter anchor points with insufficient semantic relevance.
[0041] Based on the calculated semantic guidance vector, the original query vector is linearly combined with the semantic guidance vector to obtain the aligned query vector. as follows:
[0042] in, This indicates the alignment of the query vector. Represents the original query vector. Indicates the alignment strength coefficient. This represents the semantic guidance vector; the alignment strength coefficient controls the degree of influence of the semantic anchor on the query vector, and is specifically determined based on the query domain labels. Its calculation formula is as follows:
[0043] This embodiment is limited. , ,in The preset maximum alignment strength coefficient, for example , This represents the domain identification confidence level corresponding to the query domain label. It can be output synchronously by the recognition model when recognizing the domain tags of the query text. The confidence threshold is used to automatically reduce the alignment strength and use the original query vector when the confidence of domain recognition does not exceed the confidence threshold. Only when the domain judgment is clear and the confidence is high will the anchor alignment effect be enhanced, thereby achieving a balance between semantic guidance capability and robustness.
[0044] Then, the offset between the aligned query vector and the original query vector is calculated. And detect the offset. Does it exceed the offset threshold? When the offset Offset threshold not exceeded In this case, output the currently calculated alignment query vector at the current offset. Exceeding the offset threshold In such cases, the system triggers a constraint mechanism to prune the anchor point alignment results, for example, by reducing... Take the value and use the reduced value. Regenerate the alignment query vector until the offset does not exceed the offset threshold. It is important to note that... The reduction can be proportional to the degree of offset, and can be preset based on historical experience values. In actual implementation, it can also be directly rolled back to use the original query vector. This is to ensure the stability of query semantics.
[0045] In some embodiments, for a multi-round session process, when performing the current session round, if the current session round has adjacent session rounds, before generating the alignment query vector for the current round, the dynamic anchor set of the adjacent session rounds is obtained, and the dynamic anchor intersection ratio between the dynamic anchor set of the adjacent session rounds and the dynamic anchor set of the current session round is calculated. If the dynamic anchor intersection ratio is greater than or equal to the ratio threshold, it indicates that the semantic space change of the dynamic anchors in consecutive rounds is small. At this time, the alignment query vector can be generated based on the current alignment strength coefficient. If the dynamic anchor intersection ratio is less than the ratio threshold, it indicates that the semantic space of the dynamic anchors in adjacent rounds may change abruptly. At this time, the alignment strength coefficient is reduced to avoid the semantic space change abruptly affecting the actual query semantics.
[0046] S50, based on the aligned query vector, recall the candidate document vector set, sort all the candidate document vectors in the candidate document vector set based on semantic correction, and obtain the query result.
[0047] After obtaining the alignment query vector, the system performs a nearest neighbor search operation on the vector database based on the alignment query vector, and recalls the top-N candidate document vectors with the highest semantic similarity to the alignment query vector. The vector similarity can be calculated using the inverse function of cosine similarity, dot product similarity, or Euclidean distance.
[0048] The candidate document vector set obtained based on similarity is relatively close to the aligned query vector in the overall semantic space. However, due to issues such as differences in cross-language expression, polysemy, or domain ambiguity, some candidate results may not be completely consistent with the actual retrieval intent. Therefore, this embodiment performs semantic correction-based sorting on the initially recalled candidate document vector set to identify potential semantic biases or ambiguities. Specifically, firstly, the document keywords and domain tags of the candidate document corresponding to each candidate document vector are determined. The specific determination method can be consistent with the determination method of the query keywords and query domain tags of the original query text. Then, the semantic consistency score between the document keywords and query keywords is determined based on a cross-language term alignment table. Specifically, the cross-language term alignment table in this embodiment is a pre-constructed multilingual domain knowledge graph or bilingual term mapping dictionary, which records term pairs with the same semantics in different languages and their corresponding standard semantic vectors. The semantic consistency score is calculated as follows: First, key terms in documents that have a cross-language mapping relationship with the query key terms are retrieved from the cross-language term alignment table. If a precise cross-language mapping relationship exists, a basic matching score is assigned. Second, standard semantic vectors corresponding to the query key terms and document key terms in the cross-language term alignment table are extracted, and the cosine similarity between the two is calculated as the vector matching score. Finally, the basic matching score and the vector matching score are weighted and summed by combining the inverse document frequency (IDF) weights of the terms in the overall corpus to obtain the semantic consistency score. This method can effectively eliminate semantic biases caused by polysemy or inconsistent term translation in cross-language expressions.
[0049] Subsequently, the first similarity score between the domain label of each document and the query domain label, the second similarity score between each candidate document vector and the aligned query vector, and the context consistency score of each candidate document are determined based on the co-occurrence statistical features corresponding to the domain label of the document. The context consistency score in this embodiment can be calculated based on the following steps: extracting local context windows containing query keywords in the candidate documents; counting the frequency of occurrence of domain feature words related to the domain label of the document within the local context window; calculating the point mutual information (PMI) value based on the co-occurrence frequency of query keywords and domain feature words within the local context window, as a co-occurrence statistical feature; simultaneously, inputting the text within the local context window into the embedding model, calculating the similarity between its context vector and the domain prototype vector corresponding to the domain label of the document; and weightedly fusing the point mutual information value and the similarity to obtain the context consistency score. This score can quantify the degree of fit between the local context of the candidate document and the query intent and domain, effectively filtering out noisy documents that only contain isolated keywords but whose context does not match. Finally, the semantic consistency score, first similarity score, second similarity score, and contextual consistency score of each candidate document are weighted and summed to determine the semantic correction score of each candidate document vector, thus quantifying the degree of matching between the candidate document's semantics and the true intent of the query. It is important to note that the weights of each score can be configured according to empirical rules for different business domains, or adaptively adjusted based on historical retrieval feedback data.
[0050] After obtaining the semantic correction scores for candidate documents, the system re-ranks the preliminary recall results based on these scores. All candidate documents are sorted in descending order of their semantic correction scores, and the query results are presented to the user. In practice, if the semantic correction score of a candidate document vector is below a preset threshold, the document is downweighted or directly filtered. In multilingual mixed retrieval scenarios, candidate documents with higher semantic consistency are prioritized. Through this re-ranking mechanism, the final output retrieval results are not only semantically close to the query at the vector space level, but also highly consistent with the query in terms of terminology, context, and domain intent.
[0051] In actual implementation, the system also acquires explicit and implicit feedback from users during actual use, including but not limited to click behavior, dwell time, and manual annotation results. Based on the above feedback data, it constructs optimization samples and drives the adapter corresponding to the current language to make fine adjustments based on the above feedback data when the retrieval fails, without updating the backbone network parameters, thereby achieving low-cost online learning and continuous optimization.
[0052] This embodiment utilizes a dynamic semantic anchor mechanism to dynamically generate anchors that align with the current search intent and domain during the retrieval phase. This reflects the local semantic distribution of the current query in real time, achieving adaptive alignment between the domain and intent during the retrieval phase. This significantly improves the accuracy and relevance of cross-language semantic retrieval. Furthermore, by determining the adaptive alignment strength based on domain labels and fusing semantic guidance vectors, the anchor alignment effect is enhanced when the domain confidence is high, and the alignment strength is automatically reduced when the confidence is insufficient. This effectively avoids excessive semantic shift and distortion, ensuring the stability and robustness of the retrieval results from a vector geometry perspective. Finally, by recalling candidate documents based on the aligned query vectors and performing semantic correction and ranking, the ambiguity problem in cross-language retrieval can be further alleviated, ensuring that the final retrieval results are highly consistent with the true query intent at the semantic level.
[0053] Based on the same inventive concept, the second embodiment of this disclosure provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the multilingual embedded semantic retrieval method based on dynamic alignment anchors described in the first embodiment of this disclosure.
[0054] Based on the same inventive concept, the third embodiment of this disclosure provides an electronic device, including at least a memory and a processor. The memory stores a computer program, and when the processor executes the computer program in the memory, it implements the steps of the multilingual embedded semantic retrieval method based on dynamic alignment anchors described in the first embodiment of this disclosure.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A multilingual embedded semantic retrieval method based on dynamically aligned anchor points, characterized in that, include: The original query text input by the user is preprocessed, and the standardized query text, the target language type of the original query text, the query domain tags, and the query keywords are output. The target adapter is determined based on the target language type, and the target adapter is merged with the backbone network to form an embedding model. The standardized query text is vectorized based on the embedding model to obtain the original query vector. Based on the original query vector and the query domain label, dynamically generate a set of dynamic anchor points for the current session round; Calculate the attention weight between the original query vector and each dynamic anchor in the set of dynamic anchors, perform weighted aggregation on the dynamic anchors according to the attention weight to obtain the semantic guidance vector, determine the alignment strength coefficient according to the query domain label, and determine the alignment query vector according to the alignment strength coefficient, the semantic guidance vector and the original query vector; Based on the alignment query vector, a set of candidate document vectors is recalled, and all candidate document vectors in the set are sorted based on semantic correction to obtain the query results.
2. The multilingual embedded semantic retrieval method according to claim 1, characterized in that, The preprocessing of the user-inputted raw query text includes: The original query text is subjected to language recognition processing to determine the target language type; The original query text is subjected to text normalization and cleaning, including at least: removing special symbols, unifying character format, standardizing numerical and date units of measurement, and filtering stop words to obtain standardized query text; Perform domain identification on the standardized query text and output the query domain label; Extract the query keywords from the standardized query text.
3. The multilingual embedded semantic retrieval method according to claim 1, characterized in that, The step of merging the target adapter with the backbone network to form an embedded model includes: The backbone network is a pre-trained multilingual Transformer model, and the parameters of the multilingual Transformer model are kept frozen. The target adapter is a lightweight adapter trained based on LoRA fine-tuning technology. The parameters of each lightweight adapter are trained based on training sample pairs of its corresponding target language type. The parameters of the target adapter are fused with the parameters of the backbone network to form the embedding model.
4. The multilingual embedded semantic retrieval method according to claim 1, characterized in that, The step of dynamically generating a set of dynamic anchor points for this query based on the original query vector and the query domain tags includes: Based on the original query vector, a preliminary recall is performed in the vector database to obtain a set of candidate document vectors; The similarity between the original query vector and each candidate document vector in the candidate document vector set is determined and used as the weight of the candidate document vector; All candidate document vectors are weighted and aggregated according to their weights to generate a first dynamic semantic anchor point based on the distribution of candidate documents. Based on the query domain tags, select the descriptive text set corresponding to the query domain tags; The descriptive text set is input into the embedding model for vectorization processing to obtain a descriptive text vector set; All descriptive text vectors in the descriptive text vector set are averaged and aggregated to generate a second dynamic semantic anchor point based on domain description text; The dynamic anchor set is formed based on the first dynamic semantic anchor and the second dynamic semantic anchor.
5. The multilingual embedded semantic retrieval method according to claim 1, characterized in that, The step of determining the alignment strength coefficient based on the query domain label includes: The alignment strength coefficient is calculated based on the following formula: in, This represents the alignment strength coefficient. , The preset maximum alignment strength coefficient, This represents the domain identification confidence level corresponding to the query domain label. This is the confidence threshold.
6. The multilingual embedded semantic retrieval method according to claim 1, characterized in that, The step of determining the alignment query vector based on the alignment strength coefficient, the semantic guidance vector, and the original query vector includes: The alignment query vector is calculated based on the following formula: in, This indicates the alignment of the query vector. Represents the original query vector. Indicates the alignment strength coefficient. Represents a semantically guided vector; Calculate whether the offset between the alignment query vector and the original query vector exceeds an offset threshold. If the offset does not exceed the offset threshold, output the alignment query vector. If the offset exceeds the offset threshold, reduce the alignment strength coefficient and recalculate the alignment query vector until the offset does not exceed the offset threshold.
7. The multilingual embedded semantic retrieval method according to claim 6, characterized in that, Before determining the alignment query vector based on the alignment strength coefficient, the semantic guidance vector, and the original query vector, the method further includes: If the current session round has adjacent session rounds, obtain the dynamic anchor point set of the adjacent session rounds, and calculate the dynamic anchor point intersection ratio between the dynamic anchor point set of the adjacent session rounds and the dynamic anchor point set of the current session round. If the intersection ratio of the dynamic anchor points is less than the ratio threshold, the alignment strength coefficient is reduced.
8. The multilingual embedded semantic retrieval method according to any one of claims 1 to 7, characterized in that, The step of sorting all candidate document vectors in the candidate document vector set based on semantic correction includes: Determine the document keywords and domain tags of the candidate document corresponding to each candidate document vector; The semantic consistency score between the document's key terms and the query's key terms is determined based on a cross-language term alignment table. Determine the first similarity score between the domain tag to which each document belongs and the query domain tag; Determine a second similarity score between each candidate document vector and the aligned query vector; Based on the co-occurrence statistical features corresponding to the domain tags of the document, determine the context consistency score of each candidate document; The semantic consistency score, the first similarity score, the second similarity score, and the context consistency score of each candidate document are weighted and summed to determine the semantic correction score of each candidate document vector. All candidate documents are sorted in descending order of their semantic correction scores.
9. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multilingual embedded semantic retrieval method based on dynamic alignment anchors as described in any one of claims 1 to 7.
10. An electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program on the memory, it implements the steps of the multilingual embedded semantic retrieval method based on dynamic alignment anchors as described in any one of claims 1 to 7.