Vector data consistency quality management method and system
Semantic triple extraction and positive negative matching rules are performed through semantic analyzer, and vector database index is dynamically corrected, which solves the shortcomings of vector similarity search mechanism in large language models, improves retrieval efficiency and accuracy, and enhances the reliability and transparency of the model.
Patent Information
- Application Number
- CN202510614280.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-22
AI Technical Summary
The existing search mechanism based on vector similarity has problems such as lack of semantic conflict perception, insufficient data quality optimization and solidification of index structure in large language models, which affects the retrieval efficiency and accuracy.
By constructing a semantic analyzer, semantic triple extraction is performed, positive and negative matching rules are used to classify contexts, integrate matching knowledge and dynamically correct vector database indexes to optimize the quality of the knowledge base.
It improves the generation accuracy and data consistency of the search enhanced large language model, enhances the reliability and transparency of the model, and reduces operation and maintenance costs.
Smart Images

Figure CN120523926A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and database management, and in particular to a vector data consistency quality management method and system. Background Art
[0002] The breakthrough development of large language models (LLMs) has triggered a paradigm shift in the field of natural language processing. LLMs, based on deep learning architectures, have demonstrated unprecedented performance in core tasks such as language understanding, generation, and reasoning. However, despite their remarkable achievements in general-purpose domains, LLMs are increasingly limited in knowledge-intensive tasks. This is particularly true when processing queries that exceed the timeliness of training data or involve specialized domains, where they are prone to "hallucinations" of facts. Retrieval-Augmented Language Models (RALMs) organically integrate information retrieval mechanisms with language generation capabilities to construct a dynamic knowledge update system. RALMs can effectively reduce the occurrence of hallucinations, enable real-time updates of knowledge bases at marginal cost, and provide a traceable reference for generated content.
[0003] Retrieval-enhanced language models (RALMs) effectively alleviate the limitations of large language models (LLMs) in terms of data timeliness and domain adaptability by integrating external knowledge bases. Existing retrieval mechanisms based on vector similarity have the following drawbacks:
[0004] Lack of semantic conflict perception: The retrieval mechanism based on vector similarity has difficulty capturing logical contradictions between concepts, resulting in retrieval noise interfering with model reasoning.
[0005] Insufficient data quality optimization: Existing methods mainly rely on noise detection and filtering, ignoring potential learning signals in inconsistent data.
[0006] Solidified index structure: Vector databases lack a dynamic correction mechanism, and incorrect semantic associations persist for a long time, affecting retrieval efficiency.
[0007] To address the above issues, a systematic solution is urgently needed to improve the robustness of the retrieval enhancement system from the perspective of data quality management. Summary of the Invention
[0008] The purpose of the present invention is to provide a vector data consistency quality management method and system to solve the technical problems of defects in the existing retrieval mechanism based on vector similarity, and to improve the generation accuracy and data consistency of RALMs by dynamically optimizing the knowledge base index structure and classifying and utilizing positive / negative matching contexts.
[0009] In order to achieve the above object, the technical solution of the present invention is as follows:
[0010] A vector data consistency quality management method comprises the following steps:
[0011] S1. Extract semantic triples from the input question and retrieval context using the Large Language Model (LLM).
[0012] S2. Based on the CCMDs rule, the search context is classified into positive match, negative match and fuzzy match;
[0013] S3. Integrate positive and negative matching contexts to construct enhanced prompt words, and input them into LLM to generate answers;
[0014] S4. Dynamically modify the vector database index based on the negative matching results to optimize the knowledge base quality.
[0015] Furthermore, the S1 includes the following steps:
[0016] S11. Build a semantic analyzer based on LLM and write prompt words to help the model implement traditional natural language processing (NLP) semantic analysis technology;
[0017] S12. Identify the sentence triple structure {subject, predicate, object} centered on the predicate, and return the set of semantic triples contained in the question and context.
[0018] Furthermore, the S2 includes the following steps:
[0019] S21. Migrate the matching dependency MDs and negation constraints DCs in the database field to the natural language processing scenario and propose a context-specific matching dependency CCMDs rule system;
[0020] S22.CCMDs rules adopt a dual-path decision mechanism: positive matching constraints PMCs verify knowledge consistency, and negative matching constraints NMCs detect semantic conflicts;
[0021] S23. According to the designed hints, the retrieval context is dynamically divided into three categories: positive matching context that satisfies PMCs, i.e., positive matching, negative matching context that triggers NMCs, i.e., negative matching, and fuzzy context that does not trigger the constraints, i.e., fuzzy matching.
[0022] Further: S23 includes the following steps:
[0023] S231. The designed prompts include instructions and rules for semantic extraction, semantic element comparison, and context type judgment;
[0024] S232. In the instruction section, guide the LLM to play the role of a semantic analyzer when extracting sentence elements, let it master some specific techniques in the field of natural language processing, and explain to it using natural language the pattern of matching semantic elements between two sentences;
[0025] S233. In the rule section, explain the logic of CCMDs to the LLM and instruct it to compare the sentence components in the triple pairs to determine whether the context meets or violates the constraints.
[0026] Furthermore, the S3 includes the following steps:
[0027] S31. The positive matching and negative matching contexts are marked separately and combined with the question to form the final prompt word;
[0028] S32. Positive matching knowledge provides supplementary information that the correct LLM lacks, while negative matching knowledge defines the boundaries of the LLM answer and avoids typical errors.
[0029] S33.LLM integrates its internal knowledge with the input search context to generate accurate query answers;
[0030] S34. When searching a low-quality or mismatched knowledge base, if there is no "positive matching" knowledge, directly input the top-k contexts into the LLM as a reference.
[0031] Furthermore, the S4 includes the following steps:
[0032] S41. Capture potential causal relationships through loose screening, then design a falsification context identification mechanism based on the causal graph structure, use interference variable blocking analysis to construct the falsification logic of causal pairs, and transform the graphical constraints of d-separation theory into retrieval filtering rules based on interference variable detection;
[0033] S42. Based on the “eyewitness theorem”, a dynamic index correction strategy is proposed to gradually optimize the index structure of the knowledge base by detecting conflicts between positive or negative examples in multiple searches, cutting off incorrect semantic similarity associations in the vector space.
[0034] Furthermore, the S41 includes the following steps:
[0035] S411. Use a model to pre-filter all contexts, guide the model through natural language instructions to identify potential causal relationships, set loose filtering conditions, and determine whether there is a potential causal relationship between the mentioned events;
[0036] S412. Develop a new binary matching mechanism, design a falsification context recognition mechanism based on the causal graph structure, construct the falsification logic of causal pairs using interference variable blocking analysis, and transform the graphical constraints of d-separation theory into retrieval and filtering rules based on interference variable detection.
[0037] S413. Integrate the confirmatory context and the falsification context with the question and input them into the big model to obtain the expected model reasoning answer.
[0038] Furthermore, the S42 includes the following steps:
[0039] S421. An index pruning algorithm based on a voting mechanism identifies and removes incorrect index connections in the search data by comparing and judging multiple data points, ensuring that semantically inconsistent contexts and their corresponding vectors are no longer incorrectly associated.
[0040] S422. By detecting conflicts between positive or negative examples in multiple searches, the incorrect semantic similarity associations in the vector space are cut off, and the index structure of the knowledge base is gradually optimized.
[0041] A vector data consistency quality management system, comprising:
[0042] Semantic analysis module, used for triple extraction and CCMDs rule determination; and
[0043] Hint enhancement module, which integrates positive or negative matching knowledge to generate enhanced input; and
[0044] The index correction module dynamically optimizes the vector database based on negative example learning.
[0045] By adopting the above technical solution, the present invention has the following advantages:
[0046] This paper provides a vector data consistency quality management method and system, optimizing the use of vector databases in RALMs. Based on data quality rule theory, it explores vector data quality management methods, enabling advanced data processing and knowledge extraction capabilities, and providing theoretical guarantees for vector data quality for the practical application of large language models. By designing data management methods to provide high-precision retrieval methods, contextual consistency is guaranteed, thereby enhancing the reliability of large models. Simultaneously, interpretable rule design enables non-technical users to understand the model's decisions, increasing model transparency and trust, supporting dynamic optimization of large-scale vector databases, and reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Flowchart of the vector data consistency quality management method of the present invention;
[0048] Figure 2Performance comparison of the proposed method with the original RALM and CDIT on Bloomz-7b1;
[0049] Figure 3 The performance comparison of the method of the present invention and the original RALM and CDIT on Llama2-7b is shown. DETAILED DESCRIPTION
[0050] The technical solution of the present invention is described in detail below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as first and second, etc., are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include," "comprise," or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or apparatus.
[0051] A vector data consistency quality management system includes: a semantic analysis module for triple extraction and CCMDs rule determination; a prompt enhancement module for integrating positive or negative matching knowledge to generate enhanced input; and an index correction module for dynamically optimizing the vector database based on negative example learning.
[0052] A vector data consistency quality management method is as follows: Figure 1 As shown, applying the above-mentioned vector data consistency quality management system includes the following steps:
[0053] S1. Extract semantic triples from the input question and retrieval context using the Large Language Model (LLM).
[0054] S1 includes the following specific steps:
[0055] S11. Build a semantic analyzer based on LLM and write prompt words to help the model implement traditional natural language processing (NLP) semantic analysis technology;
[0056] S12. Identify the sentence triple structure {subject, predicate, object} centered on the predicate, and return the set of semantic triples contained in the question and context.
[0057] Natural language processing (NLP) employs three analysis techniques: lexical analysis, syntactic analysis, and semantic analysis. These techniques enable entity recognition and relation extraction, further analyzing the dependencies between lexical elements within a sentence and annotating semantic role structures such as predicates and arguments. We have designed a set of prompts to enable the LLM semantic analyzer to implement these analysis methods and accurately identify and extract the triple structure of a sentence.
[0058] Definition 1 (LLM semantic analyzer): Given a sentence s and an empty aligned triple T. We introduce the function F to define the LLM semantic analyzer:
[0059] F(s)=T=[t1,t2,…,t n ]
[0060] Here, t = {sub, pre, obj} is a triple containing the predicate and the core parameters we selected—subject and object, referred to as pre, sub, and obj, respectively. It can be seen that depending on the number of predicates, a single sentence may produce more than one semantic triple after passing through the semantic analyzer. Focusing on the predicate, the sentence's triple structure ({subject, predicate, object}) is identified, and the set of semantic triples contained in the question and context is returned.
[0061] S2. Based on the CCMDs rule, the search context is classified into positive match, negative match and fuzzy match;
[0062] S2 includes the following specific steps:
[0063] S21. Migrate the matching dependency MDs and negation constraints DCs in the database field to the natural language processing scenario and propose a context-specific matching dependency CCMDs rule system;
[0064] S22.CCMDs rules adopt a dual-path decision mechanism: positive matching constraints PMCs verify knowledge consistency, and negative matching constraints NMCs detect semantic conflicts;
[0065] S23. According to the designed hints, the retrieval context is dynamically divided into three categories: positive matching context that satisfies PMCs, i.e., positive matching, negative matching context that triggers NMCs, i.e., negative matching, and fuzzy context that does not trigger the constraints, i.e., fuzzy matching.
[0066] Based on matching dependencies (MDs) and negation constraints (DCs), we propose contextually explicit matching dependencies (CCMDs). This model identifies "explicit contexts" consisting of both positive and negative aspects by establishing positive and negative matching rules. Similar to a tuple in a relational database, which consists of multiple attributes, a natural language sentence can also be represented using semantic elements such as {predicate, argument}. During the construction process, we draw on the idea of entity integrity constraints in relational databases and directly use the triple data model t = {sub, pre, obj} as attributes. Under the guidance of the constraint rules, we determine whether two sentences satisfy a precisely aligned question-answer relationship.
[0067] Specifically, semantic id (sid) is defined as the semantic representation of a sentence. Similar to the "id" as the primary key in a relational database, sid represents the uniqueness of a sentence in a high-dimensional semantic space. If the semantics of sentences s1 and s2 are completely similar, then their semantic ids have positive consistency, represented as s1[sid]~s2[sid]. On the contrary, if the negative matching constraint is met, it is marked as Indicates that their semantic ids are negatively matched. This dual identification mechanism provides a quantitative basis for the establishment of subsequent constraint rules. CCMDs consists of two core components: positive matching constraints (PMCs) based on MDs and negative matching constraints (NMCs) based on DCs.
[0068] In a specific embodiment, for the positive matching constraint
[0069] Definition 2 (PMCs): For a sentence s α and s β , they satisfy the positive matching constraint in the following form:
[0070]
[0071] Among them, T α and T β The sentences are α and s β The semantic triple set of t α and t β It means s α ,s β The semantic triples in , can be used to call different semantic elements in {sub, pred, obj} through t[i]. ≈ indicates that the corresponding sentence components are matched. In addition, and They represent the inconsistency of sentence components and the semantic dissimilarity between two sentences respectively.
[0072] For PMCs, our rules are based on MDs and are used to determine whether the context is a positive match for the query. We define this as: matching the query and the context semantic triples in the set of triples one by one. When there is a triple pair in which all elements match, the two sentences are considered a positive match, meaning that the context can provide correct knowledge support for the query.
[0073] For negative matching constraints:
[0074] Definition 3 (NMCs): For a sentence s α and s β , they satisfy the negative matching constraints in the following form:
[0075]
[0076] in, Refers to exclusive-or logic.
[0077] For NMCs, we define negative matching constraints based on DCs. Through observation and analysis of natural language, we find that when all three elements in the semantic triples of the question and context do not match, the context is in an ambiguous state, unable to determine whether it supports, misleads, or has no effect on the question; however, when two elements of a triple pair of two sentences match while the other does not, it can be determined that the context provides misleading information to the query. Using this rule, we define NMCs: when the triples in the query and context semantic triple sets match one-to-one, and there is a triple pair with two elements matching and the other element not matching, the two sentences are negatively matched, which means that the context is providing misleading information to the query.
[0078] During the semantic triple comparison process, special consideration is given to the antonym relationship between active and passive voice conversion. Specifically, for the triple {sub, pred, obj} extracted from the original sentence, we generate the corresponding antonym triple {obj, pred*, sub} through active and passive voice conversion and add it to the triple set. Here, pred* represents the passive voice form of the original predicate pred. Furthermore, when determining PMCs and NMCs, PMCs have higher priority. If a triple contains a tuple that matches both PMCs and NMCs, the context is considered a "positive matching context." If no triple pair between the two sentences satisfies either PMCs or NMCs, the context is considered a worthless fuzzy context. Using CCMDs, we perform positive and negative matching of the relevant context with the question, dividing it into explicit context (positive and negative matching context) and fuzzy context. This classification enables us to provide high-quality, valuable knowledge support for large language models.
[0079] CCMDs rules employ a dual-path decision mechanism: positive matching constraints (PMCs) verify knowledge consistency, and negative matching constraints (NMCs) detect semantic conflicts. Based on this, the retrieval context is dynamically divided into three categories: positive matching context (providing reliable knowledge) that satisfies the PMCs, i.e., positive matching; negative matching context (revealing misleading information) that triggers the NMCs, i.e., negative matching; and fuzzy context (having no decision value) that does not trigger the constraints, i.e., fuzzy matching. This hierarchical processing mechanism overcomes the "one-size-fits-all" discarding of inconsistent data in traditional methods. In particular, the design of generating antonymous triples through active and passive voice conversion enhances the ability to capture semantic conflicts.
[0080] S3. Integrate positive and negative matching context to construct enhanced prompt words, which are then fed into the LLM to generate answers. Although theoretically, the LLM's inherent attention allocation mechanism may result in insufficient contextual attention, causing the LLM to select words that are inconsistent with the context when generating sentences, thereby generating hallucinations and leading to incorrect answers, some studies have found that the documents with the highest scores in the retriever are not directly relevant to the question (e.g., they do not contain the answer), indicating that the retrieval context itself is inaccurate. Furthermore, they found that adding random documents to the prompt can improve the LLM's accuracy by 35%. Therefore, providing seemingly correct but actually incorrect examples can help the LLM avoid typical hallucination errors.
[0081] Among them, S3 includes the following steps:
[0082] S31. The positive matching and negative matching contexts are marked separately and combined with the question to form the final prompt word;
[0083] S32. Positive matching knowledge provides supplementary information that the correct LLM lacks, while negative matching knowledge defines the boundaries of the LLM answer and avoids typical errors.
[0084] S33.LLM integrates its internal knowledge with the input search context to generate accurate query answers;
[0085] S34. When searching low-quality or mismatched knowledge bases, it is common to encounter situations where there is no "positive match" knowledge. In this case, the top-k contexts are directly input into the LLM as reference when there is no "positive match" knowledge.
[0086] S4. Dynamically modify the vector database index based on the negative matching results to optimize the knowledge base quality.
[0087] S4 includes the following steps:
[0088] S41. Capture potential causal relationships through loose screening, then design a falsification context identification mechanism based on the causal graph structure, use interference variable blocking analysis to construct the falsification logic of causal pairs, and transform the graphical constraints of d-separation theory into retrieval filtering rules based on interference variable detection;
[0089] S41 includes the following steps:
[0090] S411. Use a model to pre-filter all contexts, guide the model through natural language instructions to identify potential causal relationships, set loose filtering conditions, and determine whether there is a potential causal relationship between the mentioned events;
[0091] S412. Develop a new binary matching mechanism, design a falsification context recognition mechanism based on the causal graph structure, construct the falsification logic of causal pairs using interference variable blocking analysis, and transform the graphical constraints of d-separation theory into retrieval and filtering rules based on interference variable detection.
[0092] S413. Integrate the confirmatory context and the falsification context with the question and input them into the big model to obtain the expected model reasoning answer.
[0093] Causal discovery, which aims to identify causal relationships from data, requires large amounts of data and complex computation, capabilities that large models typically lack. Furthermore, causal discovery involves exploratory analysis, and mainstream causal discovery tasks can be divided into two categories: multiple-choice inference based on premise events (identifying the correct outcome from a set of hypotheses) and binary truth-or-falsification of causal relationships.
[0094] The structure of the causal graph implies a graphical constraint called d-separation, which specifies the conditional association between variables. For any given variable pair v i , v j ∈V, unless otherwise specified, we generally use V′ to represent V / {v i , v j}, which is likely to affect v i and v j The set of other variables that have a relationship between them.
[0095] Here we define d-separation symbolically: a set of variables V′ can block a path l if (a) l contains at least one arrow-emitting variable belonging to V′, or (b) l contains at least one collider variable (if variable v i is a collider variable), and the collider does not belong to V′ and has no descendants that belong to V′. If V′ blocks all paths from vi to vj, then V′ is said to d-separate vi and vj.
[0096] Let α(ij|V′)∈{0,1} be the conditional association between variables vi,vj∈V, with the variable set V′ as the condition. α(ij|V′)=0 means that according to the joint distribution P,v i and v j Conditioned on V′, they are independent, and α(ij|V′) = 1 indicates correlation. When V′ = φ, we denote it as α(ij).
[0097] Then, according to the Markov assumption and the fidelity assumption and the definition of d-separation, for v i , v j ∈V, we can derive:
[0098] 1. V′d-separation v i and
[0099] 2.α(ij)=1 and
[0100] 3.
[0101] Based on the above theoretical framework, we designed new causal context matching rules to provide accurate and effective background information support for LLM. We first designed a method for the situation where, given a causal event as a premise, the LLM needs to determine which hypothesis is the correct outcome under that premise. This set of judgment rules is divided into two stages. The first stage extracts all contexts containing causal relationships from the search results. In the second stage, we establish a rule similar to negated dependency to determine whether the context containing causal relationships matches the causal relationship in the question.
[0102] In the first stage, we use a model to pre-filter all context. Although large models lack the ability to judge causal relationships, they are more accurate in judging correlations. We use natural language instructions to guide the model to identify potential causal relationships (including explicit and implicit associations), and set a loose screening condition: "Determine whether there is a potential causal relationship between the mentioned events. Even if this relationship is not explicitly stated or only loosely implied, consider whether there is a reasonable connection." This ensures that some valid but less obvious causal relationships are not screened out at this step.
[0103] In the second stage, a new bigram matching mechanism is built to extract from each context using the following conditions.
[0104] We define a directed acyclic graph G as the causal graph extracted from this context, where each node v represents a variable or event. Pa(x) represents the parent node (cause) of node x, and Ch(x) represents the child node (result) of node x.
[0105] 1. Analyze the potential causal graph G in the context, extract all parent nodes, or all direct / indirect causes, and obtain the set
[0106] 2. Analyze the potential causal graph G in the context and establish a set of child nodes (results) for each parent node
[0107] (a) First, the cause set elements in the context are matched one-to-one with the premises. (b) If the match is successful, the hypothesis is matched one-to-one with the child node set elements of the parent node. (c) Next, define the interference variable i. If there is an interference variable i that exists in both the parent node set and the child node set of the matching parent event and child event (blocking variables of fork, chain, and collision structures) and is controllable, then the causal relationship contained in this context has a falsifying effect on the causal relationship between the premise and hypothesis in the problem (falsifying context). This is because d-separation proves that the matching parent event and child event are not a direct, correct causal relationship, but a causal path blocked by a controllable interference variable. When the controllable interference variable i does not exist, regardless of whether the first two rules (a, b) are met, they can provide support for the subsequent judgment of the large model. Therefore, this type of context can be defined as a (confirmatory context). The rest (where all three rules are not met) are interference items with no value to the large model (fuzzy context). The above judgment rules can be written as the following negation constraints.
[0108] Here the semantic consistency symbol sid continues to use the previous definition. Indicates that the causal relationship in the context is completely irrelevant to the problem, and the context is fuzzy. Indicates that the context at this time has a misleading causal relationship that is very likely to be correct and matches the problem but is actually wrong. This is a falsification context. Use → to indicate that the context at this time has a correct causal relationship that matches the original problem. This is a confirmation context.
[0109]
[0110] Conversely, if the premise is a certain result, we need to determine which hypothesis is the correct cause. We design rules using the same approach. We organize the confirmatory and falsification contexts along with the question and input them into the final large model to generate the judgment.
[0111] Example: Suppose there is a causal discovery multiple choice case as follows:
[0112] Prerequisite: Taking drug X
[0113] Question: Please choose the result that will occur for the premise event from the following three hypotheses.
[0114] Hypothesis 1: Lower blood pressure
[0115] Hypothesis 2: Headache
[0116] Hypothesis 3: Increased appetite
[0117] Based on the input question, we retrieve the following relevant context:
[0118] Context 1: "Drug X needs to bind to a co-enzyme, Y, to lower blood pressure. If the patient lacks enzyme Y, drug X will not be effective."
[0119] Context 2: "Drug X directly stimulates nerve receptors, triggering headaches."
[0120] Context 3: "Side effects of drug X include appetite suppression, but appetite returns after drug discontinuation."
[0121] Phase 1: Causal Relationship Extraction
[0122] All contexts, including potential causal relationships, are retained through relaxed filtering conditions.
[0123] Phase 2: Bigram Matching and Classification
[0124] Context 1 analysis (causal graph: drug X → enzyme Y → blood pressure drop): Matching premise (drug X) with hypothesis 1 (blood pressure drop), p[sid] ~ c[sid], h1[sid] ~ e[sid]; there is an interfering variable: enzyme Y is the direct cause of the blood pressure drop, but the effect of drug X depends on the existence of enzyme Y, that is, enzyme Y belongs to the child node set of drug X and is the parent node of the blood pressure drop. According to the above formula, if there is an interfering variable (i∈Pa(e), i∈Pa(e)), the causal relationship is denied. Category: Falsification context, the causal relationship of hypothesis 1 depends on enzyme Y, which is not a direct causal relationship.
[0125] Context 2 Analysis (Causal Diagram: Drug X → Neuroreceptor Activation → Headache): Matching the premise (Drug X) with Hypothesis 2 (Headache), p[sid] ~ c[sid], h2[sid] ~ e[sid]; no interfering variables block the path (neuroreceptor activation is an uncontrollable intermediate variable, so no additional conditions are introduced). According to the otherwise condition in Equation (4.1), c → e. Classification: Confirmatory Context, capable of supporting Hypothesis 2.
[0126] Context 3 analysis (causal diagram: drug discontinuation → appetite recovery): The premise (drug X) does not match the parent node (drug discontinuation). h1[sid]~e[sid]. Category: fuzzy context, irrelevant to the question.
[0127] Combine the falsifying context (Context 1) and the confirming context (Context 2) and input them into the model: "Drug X requires enzyme Y to lower blood pressure (falsification). Drug X directly causes headaches (confirmation). If the premise is taking drug X, choose the outcome of the premise event from the following three hypotheses."
[0128] Expected model reasoning: Hypothesis 1 (decreased blood pressure) is weakened by its dependence on the confounding variable (enzyme Y). Hypothesis 2 (headache) is supported by the confirmatory context. Hypothesis 3 (increased appetite) has no relevant evidence. Final answer: Hypothesis 2 (headache).
[0129] S42. Based on the “eyewitness theorem”, a dynamic index correction strategy is proposed to gradually optimize the index structure of the knowledge base by detecting conflicts between positive or negative examples in multiple searches, cutting off incorrect semantic similarity associations in the vector space.
[0130] In the process of building and optimizing a knowledge base, index structure plays a crucial role in the efficiency and quality of information retrieval. The Eyewitness Theorem is an index pruning algorithm based on a voting mechanism. Its core idea is to identify and remove incorrect index connections in the retrieval data by comparing and judging multiple data points, thereby ensuring that semantically inconsistent contexts and their corresponding vectors are no longer incorrectly associated.
[0131] Definition 4 (Eyewitness Theorem): Consider a question q and two sentences s1 and s2. If q[sid] ~ s1[sid], and Then the question q is regarded as a witness of the pruning process between sentences s1 and s2.
[0132] In essence, if there is a disagreement in the SID consistency judgment between q and s1 and between q and s2, that is, the existence of q reveals the difference between s1 and s2, then in the next retrieval, it should be suggested that these two contexts should not be retrieved at the same time.
[0133] After accumulating a sufficient number of such witnesses, we have reason to believe that s1 and s2 are indeed different. Based on this judgment, we adjust the vector index by severing the false similarity links between s1 and s2. This approach is applicable to indexing methods such as hierarchical navigable small worlds. The eyewitness theorem can identify and eliminate connections between contexts in the vector database that are mistakenly considered similar, allowing for efficient and accurate acquisition of more valuable context in the next search.
[0134] Wherein, S42 includes the following steps:
[0135] S421. An index pruning algorithm based on a voting mechanism identifies and removes incorrect index connections in the search data by comparing and judging multiple data points, ensuring that semantically inconsistent contexts and their corresponding vectors are no longer incorrectly associated.
[0136] S422. By detecting conflicts between positive or negative examples in multiple searches, the incorrect semantic similarity associations in the vector space are cut off, and the index structure of the knowledge base is gradually optimized.
[0137] Systematic experiments validated the effectiveness of our method in enhancing query generation models. The experiments conducted a multi-dimensional evaluation on four knowledge-sensitive datasets, including PopQA and NQ. The benchmark included six mainstream open-source models and two comparison methods (original RALM and CDIT), using average accuracy as the core evaluation metric.
[0138] Table 1 lists the comparison results of our method with CDIT and the original RALM, which are evaluated on different processing language models and tasks.
[0139] Table 1: Experimental results of the proposed method, CDIT, and original RALM (original) on different datasets and language models
[0140]
[0141] Experimental results show that the proposed method exhibits significant performance advantages: compared with the original RALM method, it achieves an average accuracy improvement of 2.79%-11.28% on five baselines such as Llama2-7b, especially in the NQ task, the Falcon-7b model improves by 17.2%; compared with the CDIT method, the average improvement is 3.7%, and Bloomz-7b1 improves by 21.09% in the NQ task.
[0142] In order to analyze the impact of the number of contexts returned by the retriever on the performance of NDIC, we used Bloomz-7b1 and Llama2-7b as generators and the NQ dataset as the task. We changed the number of retrieval contexts returned by the retriever (increasing Top-k from 6 to 20 / 30) and conducted experiments. The specific experimental results are shown in the figure. Figure 2 、 Figure 3 shown.
[0143] from Figure 2 、 Figure 3 It can be seen that the method of the present invention has different degrees of improvement on the original RALM as the Top-k value changes, and the improvement is more obvious when the Top-k value is larger. It can be seen that as the Top-k number increases, when the number of contexts reaches 20, Bloomz-7b1 completely loses its reasoning ability in the original RALM framework, while NDIC still maintains good performance. When Top-k is 30, Llama2-7b also retains the ability to answer correctly under the method of the present invention. The main reason is that: the larger the Top-k value, the more useless information is returned, thereby reducing the performance of LLM. The method of the present invention can filter out this useless information. On the contrary, a smaller Top-k value may provide a relatively small pruning space, resulting in limited performance improvement.
[0144] Finally, it should be pointed out that although the present invention has been described with reference to the current specific embodiments, ordinary technicians in this technical field should realize that the above embodiments are only used to illustrate the present invention and are not used to limit the present invention. Various equivalent changes or substitutions can be made without departing from the concept of the present invention. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the essential spirit of the present invention, they will fall within the scope of the claims of the present invention.
Claims
1. A vector data consistency quality management method, characterized in that: The following steps are involved: S1. Extract semantic triples from the input question and retrieval context using the Large Language Model (LLM). S2. Based on the CCMDs rule, the search context is classified into positive match, negative match and fuzzy match; S3. Integrate positive and negative matching contexts to construct enhanced prompt words, and input them into LLM to generate answers; S4. Dynamically modify the vector database index based on the negative matching results to optimize the knowledge base quality.
2. A vector data consistency quality management method according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Build a semantic analyzer based on LLM and write prompt words to help the model implement traditional natural language processing (NLP) semantic analysis technology; S12. Identify the sentence triple structure {subject, predicate, object} centered on the predicate, and return the set of semantic triples contained in the question and context.
3. A vector data consistency quality management method according to claim 1, characterized in that: The S2 includes the following steps: S21. Migrate the matching dependency MDs and negation constraints DCs in the database field to the natural language processing scenario and propose a context-specific matching dependency CCMDs rule system; S22.CCMDs rules adopt a dual-path decision mechanism: positive matching constraints PMCs verify knowledge consistency, and negative matching constraints NMCs detect semantic conflicts; S23. According to the designed hints, the retrieval context is dynamically divided into three categories: positive matching context that satisfies PMCs, i.e., positive matching, negative matching context that triggers NMCs, i.e., negative matching, and fuzzy context that does not trigger the constraints, i.e., fuzzy matching.
4. A vector data consistency quality management method according to claim 3, characterized in that: The S2 includes the following steps: The S23 includes the following steps: S231. The designed prompts include instructions and rules for semantic extraction, semantic element comparison, and context type judgment; S232. In the instruction section, guide the LLM to play the role of a semantic analyzer when extracting sentence elements, let it master some specific techniques in the field of natural language processing, and explain to it using natural language the pattern of matching semantic elements between two sentences; S233. In the rule section, explain the logic of CCMDs to the LLM and instruct it to compare the sentence components in the triple pairs to determine whether the context meets or violates the constraints.
5. A vector data consistency quality management method according to claim 1, characterized in that: The S3 includes the following steps: S31. The positive matching and negative matching contexts are marked separately and combined with the question to form the final prompt word; S32. Positive matching knowledge provides supplementary information that the correct LLM lacks, while negative matching knowledge defines the boundaries of the LLM answer and avoids typical errors. S33.LLM integrates its internal knowledge with the input search context to generate accurate query answers; S34. When searching a low-quality or mismatched knowledge base, if there is no "positive matching" knowledge, directly input the top-k contexts into the LLM as a reference.
6. A vector data consistency quality management method according to claim 5, characterized in that: The S4 includes the following steps: S41. Capture potential causal relationships through loose screening, then design a falsification context identification mechanism based on the causal graph structure, use interference variable blocking analysis to construct the falsification logic of causal pairs, and transform the graphical constraints of d-separation theory into retrieval filtering rules based on interference variable detection; S42. Based on the "eyewitness theorem", a dynamic index correction strategy is proposed. By detecting conflicts between positive or negative examples in multiple searches, incorrect semantic similarity associations in the vector space are cut off, and the index structure of the knowledge base is gradually optimized.
7. A vector data consistency quality management method according to claim 6, characterized in that: The S4 includes the following steps: The S41 includes the following steps: S411. Use a model to pre-filter all contexts, guide the model through natural language instructions to identify potential causal relationships, set loose filtering conditions, and determine whether there is a potential causal relationship between the mentioned events; S412. Develop a new binary matching mechanism, design a falsification context recognition mechanism based on the causal graph structure, construct the falsification logic of causal pairs using interference variable blocking analysis, and transform the graphical constraints of d-separation theory into retrieval and filtering rules based on interference variable detection. S413. Integrate the confirmatory context and the falsification context with the question and input them into the big model to obtain the expected model reasoning answer.
8. A vector data consistency quality management method according to claim 6, characterized in that: The S42 includes the following steps: S421. An index pruning algorithm based on a voting mechanism identifies and removes incorrect index connections in the search data by comparing and judging multiple data points, ensuring that semantically inconsistent contexts and their corresponding vectors are no longer incorrectly associated. S422. By detecting conflicts between positive or negative examples in multiple searches, the incorrect semantic similarity associations in the vector space are cut off, and the index structure of the knowledge base is gradually optimized.
9. A vector data consistency quality management system, used to implement the method according to claims 1-8, characterized in that: include: Semantic analysis module, used for triple extraction and CCMDs rule determination; and The prompt enhancement module integrates positive or negative matching knowledge to generate enhanced input; and The index correction module dynamically optimizes the vector database based on negative example learning.
Citation Information
Cited By
Method for performing control logic decision in complex dialogue based on large model fine tuning and dynamic sample
CN121071100A