Method and device for enhancing embedded vector features based on semantic weight model

By introducing a semantic weight model and dynamically adjusting the vocabulary weight, the problem of insufficient similarity calculation accuracy caused by the inability to distinguish the importance of vocabulary in RAG technology is solved, and more accurate semantic matching and correlation improvement of generated content is achieved.

CN120106085AInactive Publication Date: 2025-06-06CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510602023.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing RAG technology, when the embedding model converts text into dense vector representations, it is impossible to effectively distinguish the importance of vocabulary, resulting in insufficient accuracy of similarity calculations and bias in semantic matching.

Method used

The semantic weight model is introduced, and the weight of vocabulary is dynamically adjusted through corpus preparation, semantic weight label generation, model architecture selection and training, and the Embedding vector characteristics are enhanced.

Benefits of technology

It improves the ability to capture the core semantic components of the sentence, reduces the impact of unimportant vocabulary on similarity calculation, and significantly improves the accuracy and relevance of the generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106085A_ABST
    Figure CN120106085A_ABST
Patent Text Reader

Abstract

The invention relates to a method and a device for enhancing embedded vector features based on a semantic weight model, which can dynamically adjust the weight of each vocabulary in the feature extraction process by introducing the semantic weight model. According to the method, core semantic components in sentences can be captured more accurately, and similarity misjudgment caused by unimportant modification components is reduced. Through the improvement of the precision, the RAG system can better match user query and knowledge base content in the retrieval and generation process, and the accuracy and correlation of the generated content are remarkably improved. In practical application, a text generally contains complex semantic structures and hierarchies, and the existing feature extraction method is often poor in performance when facing the complex semantic structures and hierarchies. The semantic weight model can better understand and match a complex sentence structure by performing fine weighting processing on semantic elements. The method is particularly suitable for processing queries involving a plurality of entities, relationships and semantic levels, and the capability of processing complex problems of the system is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application field of large models generated based on retrieval enhancement, and in particular to a method and device for enhancing embedded vector features based on a semantic weight model. Background Art

[0002] In the current field of natural language processing, generative pre-trained models (such as GPT and BERT) are widely used in various text generation tasks. These models are trained on a large amount of public data and have strong language understanding and generation capabilities. However, due to the limitations of training data, these models often fail to provide accurate answers when faced with some specific fields or real-time problems. In addition, the content generated by large models sometimes has the so-called "hallucination problem", that is, the content generated by the model may be reasonable in form, but may be completely wrong in practical terms.

[0003] To make up for these shortcomings, RAG technology came into being. RAG stands for Retrieval-Augmented Generation. RAG combines the retrieval function of external knowledge bases with the generation capability of large models to generate more accurate and contextually relevant content. The main processes of RAG include: Retrieval: Retrieve relevant information from an external knowledge base based on the user’s query content, usually using an embedding model to convert the query into a vector and matching relevant content through similarity search.

[0004] Enhancement: Integrate the user’s query content with the retrieved information to better support subsequent content generation.

[0005] Generate: Pass the augmented input to the generative model and output the final result.

[0006] In actual use, the implementation of RAG technology usually relies on two key components: embedding model and vector database. The embedding model is responsible for mapping text data into a dense vector space, so that texts with similar semantics are closer in the vector space. The vector database is responsible for storing these embedded vectors and supporting efficient similarity retrieval. In actual applications, RAG converts the user's query into a vector, retrieves the most similar results in the vector database, integrates these results with the query content, and inputs them into the generation model for content generation.

[0007] Among them, the Embedding vector is the embedding vector. The insufficient accuracy of similarity calculation is a major technical problem in the existing RAG technology.

[0008] In the RAG architecture, the Embedding model converts text into dense vector representations, which are usually mapped into a high-dimensional space to retrieve relevant information through similarity search. In practical applications, the similarity between vectors is usually calculated using cosine similarity or Euclidean distance. However, these similarity measurement methods have limitations, mainly reflected in the inability to effectively distinguish key semantic information in the text.

[0009] 1. The assumptions about the importance of each component are inaccurate: The traditional Embedding model is only trained by vectorization through a large number of synonymous text sentences. Although the semantics it has learned has a certain degree of differentiation, it does not have strong differentiation of key information. The difference can be seen by comparing the search performance of general text vector search and keyword search engines. In natural language processing, different words in a sentence have different semantic weights. For example, the subject and predicate usually carry the core information of the sentence, while modifiers such as adjectives or adverbs, although they supplement the meaning of the sentence, are usually less important than the subject and predicate. Existing Embedding feature extraction methods cannot distinguish these differences more finely, which may cause words that are not important in actual semantics to have too much influence on the calculation of similarity.

[0010] 2. Key semantic information matching deviation: Since the existing Embedding model cannot effectively distinguish the importance of words when calculating similarity, it is easy to cause deviations in semantic matching. For example, if the key words in two sentences are not exactly the same, but their modifiers are highly similar, the existing similarity calculation may overestimate their similarity. This situation is particularly obvious when processing complex queries, resulting in a large deviation between the retrieved content and user needs, affecting the accuracy of the final generated results. Summary of the invention

[0011] The present invention provides a method and device for enhancing Embedding vector features based on a semantic weight model. The method for enhancing Embedding vector features based on a semantic weight model includes a semantic weight model training part and a semantic weight similarity calculation part in a retrieval enhancement generation process; In the semantic weight model training part, the following steps are included: S1. Corpus preparation and text preprocessing; S2. Semantic weight label generation and dataset construction; S3. Semantic weight model architecture selection, training and evaluation; S4. Obtaining a semantic weight model; The semantic weight similarity calculation part in the retrieval enhancement generation process includes the following steps: A1. Text segmentation and storage, Embedding weight vector generation; A2. Generate user query retrieval results; A3. Enhancement of search results and output of results to users.

[0012] In particular, in said S1, data in the target application field is collected and aggregated into a corpus; Perform word segmentation, part-of-speech tagging and stop word removal on the text, use named entity recognition to extract entities, and obtain entity extraction results; Rewrite the text sentences according to grammatical rules, extract the core information of the sentences and obtain the key elements extraction results.

[0013] In particular, in S2, based on the entity extraction results and the key element extraction results, semantic weights are assigned to the words in each sentence, and the text composed of the processed sentences is paired with the semantic weight labels to form a training data set. The data set contains the embedding vector of the text and the corresponding semantic weight information, and is divided into training set, validation set, and test set.

[0014] In particular, in S3, a model architecture for semantic weight calculation is selected, the input is an Embedding vector, the output is an adjusted weighted vector, and a similarity score is performed.

[0015] In particular, in A1, the text and / or fragments are imported into the Embedding model under the RAG framework, converted into vectors, and the semantic weight model obtained in S4 is used to enhance the component weights of the features processed by the Embedding model, and the enhanced feature vectors are stored in the vector database.

[0016] In particular, the Embedding model is adapted to the semantic weight model and consistency is ensured.

[0017] In particular, in A2, the vector data obtained after the text of the user's query is processed by the Embedding model and the semantic weight model is matched with the data in the vector database, and the top K results with the highest similarity in the matching results are selected as the retrieval results.

[0018] In particular, in A3, the text of the user's query is integrated with the search results to form enhanced prompt words, and the generated prompt words are used to generate output results through a large language model and fed back to the user.

[0019] In particular, the large language model is GPT-4.

[0020] The present invention also provides a device for enhancing embedded vector features based on a semantic weight model, comprising a processor, the processor being coupled to a memory, the memory being used to store a computer program, the processor being used to execute the computer program stored in the memory, so that the device for enhancing embedded vector features based on a semantic weight model executes a method for enhancing embedded vector features based on a semantic weight model.

[0021] The beneficial effects of the present invention include at least one of the following: (1) By introducing a semantic weight model, the weight of each word can be dynamically adjusted during the feature extraction process. Compared with the traditional embedding model, this method can more accurately capture the core semantic components in a sentence and reduce similarity misjudgments caused by unimportant modifying components. This improvement in accuracy enables the RAG system to better match user queries with knowledge base content during retrieval and generation, significantly improving the accuracy and relevance of generated content.

[0022] (2) In practical applications, texts usually contain complex semantic structures and levels, and traditional embedding feature extraction methods often perform poorly in the face of these complexities. The semantic weight model can better understand and match complex sentence structures by finely weighting semantic elements. This method is particularly suitable for processing queries involving multiple entities, relationships, and semantic levels, greatly improving the system's ability to handle complex problems.

[0023] (3) The semantic weight model can reduce the interference of irrelevant information through more accurate semantic feature calculation, thereby improving the relevance and quality of retrieval results. This method can reduce mismatches and improve the overall efficiency of the system. Compared with the traditional embedding method, although a new calculation model is introduced, since the text with high relevance is obtained during the retrieval comparison process, the time for the ranking model to do relevance matching in the traditional method can be saved, and the overall resource consumption and response time may be further optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the embodiments of the present application, constitute a part of the present application, and do not constitute a limitation on the embodiments of the present invention.

[0025] Figure 1 Flowchart generated for the RAG search. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0027] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the various embodiments can be referred to each other.

[0029] like Figure 1 As shown, in this embodiment, a method for enhancing Embedding vector features based on a semantic weight model includes a semantic weight model training part and a semantic weight similarity calculation part in a retrieval enhancement generation process; In the semantic weight model training part, the following steps are included: S1. Corpus preparation and text preprocessing; S2. Semantic weight label generation and dataset construction; S3. Semantic weight model architecture selection, training and evaluation; S4. Obtaining a semantic weight model; The semantic weight similarity calculation part in the retrieval enhancement generation process includes the following steps: A1. Text segmentation and storage, Embedding weight vector generation; A2. Generate user query retrieval results; A3. Enhancement of search results and output of results to users.

[0030] The purpose of this design is to dynamically adjust the weight of each word during the feature extraction process by introducing a semantic weight model. Compared with the traditional Embedding model, this method can more accurately capture the core semantic components in the sentence and reduce similarity misjudgments caused by unimportant modifying components. This improvement in accuracy can enable the RAG system to better match user queries with knowledge base content during retrieval and generation, significantly improving the accuracy and relevance of generated content.

[0031] In this embodiment, in S1, data in the target application field is collected and aggregated into a corpus; Perform word segmentation, part-of-speech tagging and stop word removal on the text, use named entity recognition to extract entities, and obtain entity extraction results; Rewrite the text sentences according to grammatical rules, extract the core information of the sentences and obtain the key elements extraction results.

[0032] At the same time, in S2, based on the entity extraction results and key element extraction results, semantic weights are assigned to the words in each sentence, and the text composed of the processed sentences is paired with the semantic weight labels to form a training data set. The data set contains the embedding vector of the text and the corresponding semantic weight information, and is divided into training set, validation set, and test set.

[0033] Furthermore, in S3, the model architecture for semantic weight calculation is selected, the input is the Embedding vector, and the output is the adjusted weighted vector for similarity scoring.

[0034] The purpose of this design is to propose a solution to enhance the embedding vector based on the semantic weight model. Through the preparation and processing of a large number of corpora, combined with traditional natural language processing (NLP) technology and large model technology, the entity extraction, sentence rewriting and key element extraction of text sentences are completed to form a data set with semantic weights. On this basis, a model that can dynamically adjust the semantic weight is trained to enhance the vector without semantic weight extracted by the traditional Embedding model. When searching, it does not simply rely on traditional text feature extraction, but calls the semantic weight model to extract features and then infer the similarity, so as to more accurately reflect the semantic relationship between texts.

[0035] At the same time, in this embodiment, in A1, the text and / or fragments are imported into the Embedding model under the RAG framework, converted into vectors, and the semantic weight model obtained in S4 is used to enhance the component weights of the features processed by the Embedding model, and the enhanced feature vectors are stored in the vector database.

[0036] At the same time, the Embedding model is adapted to the semantic weight model and consistency is guaranteed.

[0037] Furthermore, in A2, the vector data obtained after the text of the user's query is processed by the Embedding model and the semantic weight model is matched with the data in the vector database, and the top K results with the highest similarity in the matching results are selected as the retrieval results.

[0038] At the same time, in A3, the text of the user's query is integrated with the search results to form enhanced prompt words, and the generated prompt words are used to generate output results through a large language model and fed back to the user.

[0039] Finally, the large language model can be GPT-4, etc.

[0040] In the specific implementation, in order to quantify the improvement of retrieval accuracy after Embedding, standard information retrieval evaluation indicators, such as MRR (Mean Reciprocal Rank), NDCG (Normalized Discounted Cumulative Gain), Recall@K, etc., are used to compare the retrieval effects of the traditional Embedding method and the enhanced Embedding method that introduces semantic weights.

[0041] First, a quantitative experiment design is conducted and the MS MARCO (Microsoft MAchine Reading Comprehension) dataset is selected. This dataset contains real users' query and paragraph matching information and is widely used in information retrieval tasks.

[0042] Then set two groups of data, one group is the control group using the traditional Embedding method, such as BERTEmbedding + cosine similarity; The other group, the experimental group, used the semantic weight enhanced Embedding method, which added additional semantic weighting based on BERT.

[0043] Then set the evaluation indicators and select the following indicators for evaluation: MRR@10, used to assess higher ranking importance; Recall@5 / Recall@10, used to evaluate whether the correct answer is contained in the first 5 / 10 retrieved documents; NDCG@10, which is used to evaluate the importance of considering document ranking.

[0044] The results are shown in Table 1:

[0045] The experimental results show that the experimental group has significant improvements in multiple indicators compared with the control group, especially in MRR@10 (+35%) and Recall@5 (+22%), indicating that it can find relevant documents more accurately and improve the accuracy of the top 5 candidate results.

[0046] In the example task comparison, taking the legal document retrieval scenario as an example, the user queries, "cases involving contract breach": The workflow for the control group is to return: “Contract Definition and Types” “How to sign an effective contract” "Contract breach case: A company was sued for delayed delivery" (correct) The workflow for the experimental group is to return: "Contract breach case: A company was sued for delayed delivery" (correct) "Contractual Breach of Contract Compensation Standards and Legal Basis" (Related) "How to Deal with Contract Breach: A Legal Guide for Businesses" (Related).

[0047] From the above examples, it can be concluded that the control group, i.e., the existing method, may match some non-core semantic content due to the high weight of the word "contract", while the experimental group, i.e., the semantic weight method, can better match "breach of contract" related cases and improve relevance.

[0048] In this way, the semantic weight model provides a more fine-grained semantic understanding capability, while the existing Embedding model only relies on the semantic vector feature extraction of text and sentences, and cannot accurately distinguish the importance of different components in the text. By introducing the semantic weight model, this method effectively overcomes the shortcomings of traditional methods, making the RAG system more intelligent and accurate when processing complex queries and generation tasks. Ultimately, through this improvement, RAG technology can better serve application scenarios that require high-precision text matching, such as enterprise knowledge management, intelligent question-answering systems, and advanced search engines.

[0049] When the S3 step is implemented, taking the Euclidean distance as an example, without changing the Euclidean distance formula, the weights can be introduced k i , d i To adjust the contribution of each vector component to the final distance calculation. The standard formula for Euclidean distance is: ; Here A = [A1, A2, ....., An] and B = [B1, B2, .....Bn] represent two vectors. ki and d i When acting on each component of vector A and B respectively, the formula of Euclidean distance can be expressed as: ; Here k i is the first i The weight of the component, d i is the first i In this way, the components of vectors A and B will have different weight adjustments, which can further refine the control of the semantic weight of each component.

[0050] That is, if k i Compare d i large, indicating that the first i The weight of the component in the distance calculation is higher, reflecting the semantic importance of the component; if d i Compare k i is large, it means that the first i The quantity is more important.

[0051] This way of introducing weights can handle the calculation of semantic similarity more flexibly, especially when different weights need to be assigned to different vectors. It helps to solve the deficiency of the traditional Euclidean distance formula of embedding vectors that does not distinguish the importance of each component.

[0052] Ultimately, the improvement of this formula can more accurately calculate the similarity between text features, thereby improving the accuracy of the retrieval stage in the RAG process.

[0053] Such texts usually contain complex semantic structures and levels, and traditional embedding feature extraction methods often perform poorly in the face of these complexities. The semantic weight model can better understand and match complex sentence structures by finely weighting semantic elements. This method is particularly suitable for processing queries involving multiple entities, relationships, and semantic levels, greatly improving the system's ability to handle complex problems. The semantic weight model can reduce the interference of irrelevant information through more accurate semantic feature calculations, thereby improving the relevance and quality of retrieval results. This method can reduce mismatches and improve the overall efficiency of the system. Compared with the traditional embedding method, although a new calculation model is introduced, since the text with higher relevance is obtained during the retrieval comparison process, the time for the ranking model to do relevance matching in the traditional method can be saved, and the overall resource consumption and response time may be further optimized.

[0054] In this embodiment, a device for enhancing embedded vector features based on a semantic weight model is also provided, including a processor, the processor is coupled to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the device for enhancing embedded vector features based on a semantic weight model executes a method for enhancing embedded vector features based on a semantic weight model.

[0055] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for enhancing embedding vector features based on a semantic weight model, characterized in that: It includes the semantic weight model training part and the semantic weight similarity calculation part in the retrieval enhancement generation process; In the semantic weight model training part, the following steps are included: S1. Corpus preparation and text preprocessing; S2. Semantic weight label generation and dataset construction; S3. Semantic weight model architecture selection, training and evaluation; S4. Obtaining a semantic weight model; The semantic weight similarity calculation part in the retrieval enhancement generation process includes the following steps: A1. Text segmentation and storage, Embedding weight vector generation; A2. Generate user query retrieval results; A3. Enhancement of search results and output of results to users.

2. The method for enhancing embedding vector features based on a semantic weight model according to claim 1, characterized in that: In S1, data in the target application field is collected and aggregated into a corpus; Perform word segmentation, part-of-speech tagging and stop word removal on the text, and use named entity recognition to extract entities to obtain entity extraction results; Rewrite the text sentences according to grammatical rules, extract the core information of the sentences and obtain the key elements extraction results.

3. The method for enhancing the features of embedded vectors based on a semantic weight model according to claim 2, characterized in that: In S2, based on the entity extraction results and the key element extraction results, semantic weights are assigned to the words in each sentence, and the text composed of the processed sentences is paired with the semantic weight labels to form a training data set. The data set contains the embedding vector of the text and the corresponding semantic weight information, and is divided into training set, validation set, and test set.

4. The method for enhancing the features of embedded vectors based on a semantic weight model according to claim 3, characterized in that: In S3, a model architecture for semantic weight calculation is selected, the input is an Embedding vector, the output is an adjusted weighted vector, and a similarity score is performed.

5. The method for enhancing embedding vector features based on a semantic weight model according to claim 1, characterized in that: In A1, the text and / or fragments are imported into the Embedding model under the RAG framework, converted into vectors, and the semantic weight model obtained in S4 is used to enhance the component weights of the features processed by the Embedding model, and the enhanced feature vectors are stored in the vector database.

6. The method for enhancing the features of embedded vectors based on a semantic weight model according to claim 5, characterized in that: The Embedding model is compatible with the semantic weight model and consistency is guaranteed.

7. The method for enhancing embedding vector features based on a semantic weight model according to claim 5, characterized in that: In A2, the vector data obtained after the text of the user's query is processed by the Embedding model and the semantic weight model is matched with the data in the vector database, and the top K results with the highest similarity in the matching results are selected as the retrieval results.

8. The method for enhancing embedding vector features based on a semantic weight model according to claim 7, characterized in that: In A3, the text of the user's query is integrated with the search results to form enhanced prompt words, and the generated prompt words are outputted through a large language model and fed back to the user.

9. The method for enhancing embedding vector features based on a semantic weight model according to claim 8, characterized in that: The large language model is GPT-4.

10. A device for enhancing embedded vector features based on a semantic weight model, characterized in that: The invention comprises a processor coupled to a memory, the memory being used to store a computer program, and the processor being used to execute the computer program stored in the memory, so that the device for enhancing the features of an embedded vector based on a semantic weight model performs the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Retrieval enhancement generation-based retrieval method, product, equipment and medium

    CN119003795A

  • Intelligent question answering system based on large language model

    CN119623646A