A semantic predicate-oriented auditable approximate query processing method and system

CN122817435APending Publication Date: 2026-09-25BEIJING ELECTRONICS SCI & TECH INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610987158.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

依托大语言模型实现的语义判定算子完全颠覆该类前置假设,一方面单次语义判定需发起模型推理调用,会产生持续递增的算力与服务调用成本;另一方面模型输出具备内在随机性,受提示文本、解码配置、模型迭代更新等因素影响,判定精度易出现无预警衰减

Benefits of technology

(1)在统计推断层面,实现了有限总体下语义查询计数的精确区间估计。本发明摒弃了依赖渐进正态近似的经典Wald区间构造方法,通过采用Clopper-Pearson精确二项区间并引入有限总体校正因子,同时结合超几何分布的精确反演,使得在大语言模型预算约束所限的小样本条件下,所构造的置信区间依然具有精确的覆盖率保证,从根本上克服了传统方法在小样本情形下覆盖率严重不足的理论缺陷,确保了语义查询结果的可量化统计可信度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817435A_ABST
    Figure CN122817435A_ABST
Patent Text Reader

Abstract

The application discloses a kind of semantic predicate-oriented auditable approximate query processing method and system, it is related to database query processing and artificial intelligence technical field, wherein the method includes: obtaining semantic predicate, confidence level and sampling constraint, determine document finite population and obtain basic sample without replacement uniform sampling;Adopt cheap big model to complete sample semantic determination, construct counting confidence interval in combination with accurate binomial interval and finite population correction factor;Extract audit subsample is re-evaluated by reference model, calculate model consistency lower bound and compare safety threshold to obtain credible determination;Finally output query point estimation, confidence interval and safety conclusion, complete semantic retrieval statistical estimation and model audit whole process.The application realizes model credible quantitative discrimination by double-layer check audit framework, solves the problem that cheap model evaluation is not verifiable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of database query processing and artificial intelligence technology, and in particular to an auditable approximate query processing method and system oriented towards semantic predicates. Background Technology

[0002] Large language models have gradually become the mainstream execution carrier for various semantic predicate tasks. Semantic predicates cover a variety of text analysis operators, such as text event recognition, topic classification, and risk level discrimination. When these semantic operators are embedded into the structured analysis query chain, approximate query execution schemes have significant application value. However, the underlying cost adaptation logic of traditional approximate query processing technology is fundamentally different from the semantic computing scenario driven by large language models.

[0003] Traditional approximate query processing systems rely on two inherent fundamental assumptions: semantic predicate determination operations have no computational overhead and the output is unique and stable; the system performance bottleneck stems solely from storage read / write interactions. Semantic determination operators implemented using large language models completely overturn these presuppositions. On one hand, each semantic determination requires initiating a model inference call, resulting in continuously increasing computational power and service call costs. On the other hand, the model output possesses inherent randomness, influenced by factors such as prompt text, decoding configuration, and model iteration updates, making it prone to unannounced decay in determination accuracy.

[0004] Changes in underlying costs and output characteristics have led to three unavoidable technical flaws in traditional approximate query frameworks. First, confidence intervals built on the theory of asymptotic normality only have reliable coverage in scenarios with large sample sizes. However, the inference cost of large language models constrains the sampling scale, and the statistical reliability of intervals drops significantly in scenarios with small samples. Second, traditional solutions rely on offline pre-built data summary structures, and the dynamic changes in model output can cause pre-computed summaries to become invalid. Third, existing approximate query mechanisms lack the ability to identify the reliability of semantic operators themselves and cannot perceive the silent degradation of model accuracy.

[0005] Current semantic query tools for large language models, while supporting the embedding of custom semantic operators, can only output aggregated statistical single-point estimates and cannot provide theoretically guaranteed confidence intervals for aggregated results. Some optimization schemes that introduce surrogate models for assisted sampling rely on the ideal premise of low-cost scoring of the entire text by the surrogate model; if the surrogate carrier is still a large language model, this precondition cannot be met. All existing semantic query architectures lack verification and auditing mechanisms for the reliability of semantic operator judgments, making it difficult to quantify and identify model output biases.

[0006] In summary, traditional approximate query processing frameworks cannot adapt to the business characteristics of high inference costs, random fluctuations in output, and silent decay of accuracy of semantic operators in large language models. Existing semantic query tools also cannot simultaneously achieve reliable interval estimation for small samples and operator credibility verification. There is an urgent need for auditable approximate query processing methods that are adapted to the scenarios of semantic operators in large language models. Summary of the Invention

[0007] The main objective of this invention is to provide an auditable approximate query processing method oriented towards semantic predicates.

[0008] Another objective of this invention is to propose an auditable approximate query processing system oriented towards semantic predicates.

[0009] The third objective of this invention is to provide an electronic device.

[0010] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0011] To achieve the above objectives, a first aspect of the present invention proposes an auditable approximate query processing method oriented towards semantic predicates, comprising:

[0012] Obtain the semantic query predicate, confidence level parameters, and sampling workload constraints specified by the user, determine the finite population consisting of all documents to be retrieved, and perform uniform random sampling without replacement on the finite population to obtain the basic sample set. A cheap large language model is used to perform semantic predicate judgment on each document in the basic sample set to generate an evaluation record with initial evaluation tags. Based on the evaluation record, the exact binomial interval method is used in combination with a finite population correction factor to construct a confidence interval for the population conformity number using the obtained confidence level parameters. Randomly select audit subsamples from the basic sample, use the reference large language model for secondary judgment, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate the model credibility judgment conclusion. Based on the evaluation records, the point estimates of the query results are calculated, and the point estimates, the confidence intervals, and the credibility judgment conclusions are summarized and output in a unified manner to complete the entire process of statistical estimation, model bias verification, and credibility judgment for large-scale document semantic retrieval.

[0013] Optionally, uniform random sampling without replacement is performed on the finite population to obtain a basic sample set, including: The size of the basic sample set is determined based on the constraints of the sampling workload. With the principle that each document in the finite population has an equal probability of being selected, documents are extracted one by one from the finite population. After each extraction, the document is removed from the set to be extracted until the number of extracted documents reaches the determined size. All the extracted documents are then summarized into the basic sample set.

[0014] Optionally, the process of constructing confidence intervals for the population conformity number includes: The number of documents in the basic sample set that are determined to conform to the semantic predicate is taken as the number of positive examples, and the ratio of the number of positive examples to the size of the basic sample set is calculated as the sample positive example ratio. The inverse function of the regularized incomplete beta function is used to calculate the lower and upper limits of the binomial proportion interval at a given significance level. Based on the total number of documents in the finite population and the size of the basic sample set, calculate the finite population correction factor, and apply the correction factor to the lower and upper limits of the binomial proportion interval respectively, so that the binomial proportion interval shrinks according to the sampling ratio. Multiply the lower and upper limits of the contracted proportional interval by the total number of documents in the finite population, and truncate to the valid range of document counts to obtain the confidence interval in terms of document counts.

[0015] Optionally, the calculation method for the finite population correction factor is as follows: Using the total number of documents in the finite population and the size of the basic sample set as input, calculate the ratio of the size of the basic sample set to the total number of documents in the finite population, and use this ratio as the sampling proportion; The value of the finite population correction factor is determined based on the sampling ratio, such that the value of the finite population correction factor decreases as the sampling ratio increases.

[0016] Optionally, the process of calculating the lower bound of consistency between the inexpensive model and the reference model includes: The number of documents in the audit subsample whose initial evaluation markers of the cheap large language model are consistent with the secondary judgment results of the reference large language model is counted as the number of consistent positive examples. The inverse function of the regularized incomplete beta function is used as input to construct a one-sided confidence lower bound of the consistency rate, with the number of consistent positive examples, the number of documents judged to be inconsistent in the audit subsample plus one, and a given significance level. The one-sided confidence lower bound is used as the consistency lower bound between the cheap large language model and the reference large language model.

[0017] Optionally, the process of estimating the point estimate of the query result based on the evaluation record includes: The number of documents in the complete sample evaluation record that are determined to conform to the semantic predicate is used as the sample conformity number; The proportion of the number of samples that match the semantic predicates is calculated as the proportion of the number of documents that match the semantic predicates in the finite population. Multiplying the estimated value by the total number of documents in the finite population yields a point estimate of the total number of documents in the finite population that conform to the semantic predicate.

[0018] To achieve the above objectives, a second aspect of the present invention proposes an auditable approximate query processing system oriented towards semantic predicates, comprising: The overall sampling module is used to obtain the semantic query predicate, confidence level parameters and sampling workload constraints specified by the user, thereby determining the finite population consisting of all documents to be retrieved, and performing uniform random sampling without replacement on the finite population to obtain the basic sample set. The interval estimation module is used to perform semantic predicate determination on each document in the basic sample set using an inexpensive large language model, generate evaluation records with initial evaluation tags, and construct confidence intervals for the total number of conformities based on the obtained confidence level parameters by using the exact binomial interval method and combining it with a finite population correction factor. The audit verification module is used to randomly select audit sub-samples from the basic sample, perform secondary judgment using the reference large language model, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate a model credibility judgment conclusion. The results aggregation module is used to calculate the point estimate of the query results based on the evaluation records, and to aggregate and output the point estimate, the confidence interval, and the credibility judgment conclusion in a unified manner, thereby completing the entire process of statistical estimation, model bias verification, and credibility judgment for large-scale document semantic retrieval.

[0019] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0020] To achieve the above objectives, a third aspect of this application provides an electronic device, including a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, for implementing an auditable approximate query processing method oriented towards semantic predicates as described in the first aspect embodiment.

[0021] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements an auditable approximate query processing method oriented towards semantic predicates as described in the first aspect embodiment.

[0022] The embodiments of the present invention have the following beneficial effects: (1) At the statistical inference level, accurate interval estimation of semantic query counts under finite population is achieved. This invention abandons the classic Wald interval construction method that relies on asymptotic normal approximation. By adopting the Clopper-Pearson accurate binomial interval and introducing a finite population correction factor, and combining it with the accurate inversion of the hypergeometric distribution, the constructed confidence interval still has accurate coverage under the small sample condition limited by the budget constraint of the large language model. This fundamentally overcomes the theoretical defect of the traditional method in the case of insufficient coverage under small sample conditions, and ensures the quantifiable statistical credibility of the semantic query results.

[0023] (2) At the model credibility level, a self-checking system for the reliability of inexpensive model evaluation results based on a verification layer mechanism is constructed. This invention introduces a reference large language model to perform independent secondary judgment on audited subsamples, and calculates a strict one-sided confidence lower bound for the consistency of the judgment results between the two based on the Clopper-Pearson method, thereby transforming the silent performance degradation that may occur in the semantic understanding process of the large language model into a clear signal of safety or insecurity. This mechanism has a provable theoretical upper bound on the false negative rate, providing an auditable and verifiable reliability pledge for the semantic evaluation process that relies on a black-box large language model.

[0024] (3) In terms of resource efficiency, an adaptive optimal balance between statistical accuracy and computational cost is achieved. The adaptive sampling sizing mechanism set up in this invention can dynamically estimate the minimum sample size required to reach the user-specified confidence interval width target based on the semantic conformity ratio observed in the trial samples. This avoids the loss of estimation accuracy due to insufficient sample size, and also eliminates the overhead of large language model application interface calls caused by oversampling. In practical applications, this mechanism can save about 80% to 96% of the call cost, significantly reducing economic expenditure in large-scale document semantic retrieval scenarios.

[0025] (4) In terms of scope of application, a complete solution covering multiple semantic query operators under a unified framework has been constructed. This invention designs dedicated confidence interval construction methods for three core structured semantic retrieval operators: semantic predicate filtering, semantic category grouping, and semantic Top-k ranking extraction. This enables the auditable and verifiable characteristics of statistical inference to be seamlessly extended to multiple semantic query modes, meeting the unified framework deployment requirements of diverse query needs in actual information retrieval systems. Attached Figure Description

[0026] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an auditable approximate query processing method oriented towards semantic predicates, provided as an embodiment of the present invention; Figure 2 A flowchart illustrating the layered execution architecture of an auditable approximate query processing method oriented towards semantic predicates, provided in an embodiment of the present invention. Figure 3 This is a comparative diagram of two execution modes provided in an embodiment of the present invention: query without audit and cross-verification with audit. Figure 4 This is a structural diagram of an auditable approximate query processing system oriented towards semantic predicates, provided in an embodiment of the present invention. Detailed Implementation

[0027] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] The following describes, with reference to the accompanying drawings, an auditable approximate query processing method and system based on semantic predicates, according to an embodiment of the present invention.

[0030] Example 1 This invention provides an auditable approximate query processing method oriented towards semantic predicates. Figure 1 This is a flowchart illustrating an auditable approximate query processing method based on semantic predicates, provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: Step S1: Obtain the semantic query predicate, confidence level parameters, and sampling workload constraints specified by the user. Based on this, determine the finite population consisting of all documents to be retrieved, and perform uniform random sampling without replacement on the finite population to obtain the basic sample set.

[0031] In this embodiment, the initial collection of basic data and business parameters is first completed. The complete documents to be processed are summarized into a finite population D, which contains N independent text documents, and can be uniformly denoted as D. It synchronously receives three types of core constraint parameters from the user's front end, namely semantic query predicates. Confidence level Sampling budget Semantic query predicate Used to define semantic judgment criteria and confidence levels for document filtering, grouping, and sorting. Determine the statistical coverage guarantee for all subsequent confidence intervals, and the sampling budget. The total number of documents in the basic sample set obtained from this sampling is limited.

[0032] When implementing the complete execution logic of uniform random sampling without replacement, the sampling operation is carried out in accordance with the uniform sampling rule of a finite population. This sampling method corresponds to... Figure 2 The sampling module within the statistical layer performs a uniform random sampling process without replacement and ultimately outputs a subset of samples. This refers to the basic sample set obtained in this step. The specific extraction rule is based on the core principle that all documents in the finite population have an equal probability of being selected. Each document is extracted sequentially from the finite population document pool. After each extraction, the document is removed from the remaining document set to avoid duplicate inclusion. This extraction process continues until the cumulative number of extracted documents exceeds the sampling budget. Exact match, integrating and summarizing all extracted documents to form a scale of basic sample set .

[0033] To address the need for dynamic control of interval estimation accuracy in certain business scenarios where sample size is uncertain, this application embodiment also adds an adaptive sampling preprocessing procedure. This procedure is also executed in the S1 stage, before the formal sampling without replacement operation, to adapt to use cases where the interval width has clear accuracy constraints.

[0034] The complete execution logic of adaptive sampling is to first select a smaller initial trial sample. After completing the exploratory sampling operation, a low-cost large language model is used to perform semantic predicate evaluation on all documents within the exploratory sample. The proportion of positive documents in the exploratory sample is then calculated and denoted as the exploratory positive rate. The user synchronously inputs the relative width target ε as the interval precision constraint, and iteratively calculates based on the interval width calculation formula, as follows:

[0035] in, The interval width, This represents the standard normal quantile at the corresponding confidence level. Represents the total number of documents in a finite population. The sample size to be solved is... To test the positive rate.

[0036] Through multiple iterations, the interval width that can be satisfied is selected. Minimum sample size for constraints Subsequent reuse of data extracted in the earlier stages An additional sample was drawn as part of the base sample. The documents, after being merged, will have a size of [size missing]. The basic sample set is used to replace the fixed sampling budget mode to complete the sampling, thereby reducing unnecessary large model call overhead.

[0037] This application's implementation examples include a set of adaptive sampling examples, where the total number of documents in a finite population is set to N=1000, and the user provides a relative width target. First, an initial trial sample is drawn. After the evaluation is completed, the trial positive rate is calculated. Substituting the values ​​into the interval width formula and solving iteratively, we obtain the minimum sample size that satisfies the accuracy constraint. An additional 72 documents were subsequently extracted. These 50 trial samples were combined with the 72 newly added documents to form a base sample of 122. This was compared with the traditional fixed sampling strategy for the worst-case positive rate. With a preset requirement of 278 samples, the adaptive sampling mode can reduce the number of calls to large language models by 56%, significantly reducing computational resource consumption.

[0038] Step S2: Use a cheap large language model to perform semantic predicate determination on each document in the basic sample set to generate an evaluation record with initial evaluation tags. Based on the evaluation record, use the exact binomial interval method and combine it with a finite population correction factor to construct a confidence interval for the population conformity number using the obtained confidence level parameters.

[0039] This step corresponds completely. Figure 2 The semantic evaluation module within the statistical layer consists of three parallel semantic operator processing branches. The semantic evaluation module employs an inexpensive large language model. The predicate evaluation is completed for all sample documents. The parallel branches are semantic filtering operator, semantic grouping operator, and semantic Top-k rank interval estimation operator. The three types of operators independently complete the interval solution for their respective business scenarios, and finally output a unified set of upper and lower limits of confidence intervals that converge to the same set. ,correspond Figure 2 The confidence interval output node at the end of the statistical layer, in the no-audit mode, relies solely on the confidence interval output in this step as the final query statistical result. Figure 3 (a) shows the execution process without auditing.

[0040] First, the evaluation and execution process of the inexpensive large language model will be explained. In the embodiments of this application, a lightweight and inexpensive large language model with lower inference costs will be invoked. Traverse the basic sample set All internal documents Match semantic query predicates to each individual document. Perform semantic predicate evaluation to generate unique initial evaluation tags. These tags are categorized into three types based on the operator type: binary labels, classification labels, and semantic scores. All documents are then integrated with their corresponding initial evaluation tags to generate a complete evaluation record. Figure 3 Table (a) in the table visually demonstrates the format of assessment records in a no-audit scenario. The table includes document number, The evaluation results are displayed in two columns. Checkmarks in the table indicate that the document matches the semantic predicate, while crosses indicate that the document does not match the semantic predicate. The confidence interval is calculated directly based on this set of evaluation records, and there is no subsequent cross-validation process using a reference model.

[0041] The evaluation process is illustrated using a semantic filtering operator implementation case. The total number of documents in the finite population is set to N=5000, the semantic query predicate q is "mention of bill disputes", the sampling budget is n=200, and the confidence level is... Right now We used deepseek-chat as an inexpensive large language model, and fixed the inference temperature parameter to 0 to ensure that the semantic evaluation output was deterministic. After evaluating each of the 200 sample documents, we counted the number of positive documents matching the predicate within the sample, X=60.

[0042] The core method for constructing confidence intervals based on evaluation records is the Clopper-Pearson exact binomial interval superposition with finite population correction factors. This method is suitable for semantic filtering operator scenarios, and the complete calculation process is broken down step by step as follows: The first step is to count the number of documents in the basic sample set that match the semantic predicate, defined as the number of positive samples X, and then calculate the positive sample ratio. Based on this ratio, the estimated value of the query result points is directly derived. Substituting numerical values ​​into the implementation case of the semantic filtering operator yields the following results: Point estimate .

[0043] The second step is to determine the upper and lower boundaries of the Clopper-Pearson binomial ratio interval. The calculation formula is as follows:

[0044]

[0045] In the formula The function representing the inverse of the incompletely regularized beta function, with parameters... The number of positive examples in the sample. Based on the total size of the basic sample, Give the user a significance level, For the overall confidence level; This represents the lower limit of the uncorrected ratio range for the first two terms; This represents the upper limit of the uncorrected ratio range for the first two terms.

[0046] The semantic filtering operator implementation case, by substituting the parameters for calculation, yields... .

[0047] The third step is to calculate the finite population correction factor and complete the interval shrinkage. The formula for calculating the correction factor is as follows: , within the formula For a finite total number of documents, Based on the base sample size, the correction factor value will decrease synchronously as the sampling ratio increases, achieving a reasonable contraction of the loose binomial interval. After contraction, the boundary of the ratio interval is... ;in, This is the lower limit of the proportional range after finite population correction; This represents the upper limit of the proportion range after finite population correction. For finite population correction factors.

[0048] Semantic filtering operator implementation case: numerical solution for correction factor Further calculate the boundary after contraction .

[0049] The fourth step is to convert the proportion interval into a confidence interval for the document count dimension. The conversion formula is as follows: ,in, This represents the lower bound of the confidence interval for the total number of matching documents. To determine the upper limit of the overall matching document count confidence interval, the max and min functions are used to constrain the interval boundary within the effective document count range of 0 to N. The semantic filtering operator implementation case calculates the complete count confidence interval [1196, 1836], which is the final confidence interval output by the statistical layer. .

[0050] In addition to the basic Clopper-Pearson correction interval scheme, the semantic filtering operator can also use the exact hypergeometric distribution inversion method as an alternative authentication estimator, which is adapted to the statistical property that sampling without replacement naturally follows a hypergeometric distribution, and the number of positive samples X follows a hypergeometric distribution. The distribution parameter N is the total number of documents, K is the number of documents in the total population that truly match the predicate, and n is the sample size. The hypergeometric distribution test threshold is obtained by inversely solving for K, ensuring that K has a 1 / 2 or higher threshold. The logic for solving the precise confidence interval guaranteed by α coverage is as follows:

[0051]

[0052] in, This represents the lower confidence limit for the total number of true matching documents. This represents the upper confidence limit for the total number of true matching documents. α represents the probability calculation function; X is the number of positive examples obtained from the observation; α is the significance level.

[0053] This interval remains unaffected by the overall size, sample size, or true matching ratio, and maintains accurate statistical coverage.

[0054] For multi-class semantic grouping business scenarios, the statistical layer semantic grouping branch introduces the Bonferroni correction mechanism to achieve multi-class synchronous confidence coverage guarantee. Figure 2 Semantic grouping processing branch function. In this embodiment, all category tags are integrated into a category set. For each independent category g within the set, assign category tags to individual documents. Convert to binary indicator variable For each category, an independent confidence interval is constructed; to ensure that the joint confidence level for all category intervals is simultaneously valid is not less than 1. α, is used to scale the single-class significance level, resulting in a single-class significance level of . | For each category, the aforementioned Clopper-Pearson interval solution formula with finite population correction factors is applied independently. Based on the Bonferroni joint bound theorem, the sum of the interval failure probabilities of all categories does not exceed α, thereby achieving multi-group synchronous statistical constraints.

[0055] The implementation case presents a product sentiment classification scenario. The total number of documents is N=1000, and the sentiment classification set includes five categories: one-star, two-star, three-star, four-star, and five-star documents. Confidence level Original significance level The adjusted single-class significance level is 0.01, the upper limit of the failure probability of each of the five class intervals is 0.01, the upper limit of the total probability of failure of all classes simultaneously is 0.05, and the joint confidence level is maintained at 0.95.

[0056] For semantic Top-k ranking retrieval scenarios, the statistical layer semantic Top-k branch uses the rank interval estimation method to solve for the overall confidence range corresponding to the document score ranking. Figure 2Semantic Top-k rank interval estimation processing branch. In this embodiment, the inexpensive large language model first outputs semantic relevance scores in the interval of 0 to 1 for all documents in the base sample. Then, all sample documents are sorted in descending order according to the score values, and the target prefix length of the sample is calculated. In the formula, k represents the total number of documents to be retrieved as specified by the user. For sample size, Given the total number of documents, select the document with a rank equal to [the specified value]. Document score based on location as a threshold ; The semantic score within the statistical baseline sample is greater than or equal to the threshold Total number of documents ,based on Applying the Clopper-Pearson interval formula to calculate the sample proportion interval, and then scaling it to a finite population size N, we obtain the upper and lower limits of the population rank confidence interval. The calculation formula is as follows:

[0057]

[0058] in, The lower confidence limit for the number of documents with an overall semantic score higher than the threshold τ; For overall semantic scores above a threshold The upper confidence limit for the number of documents. This rank interval can accurately identify a finite population whose semantic score exceeds the threshold. The total number of documents, if the lower limit of the interval This allows us to determine that there are at least k highly relevant documents within the total, thus providing a Top-k search set that fully covers the user's needs.

[0059] The implementation case provides retrieval scenario parameters: a finite population N=1000, a user target retrieval quantity k=10, and a sampling budget n=200. The calculations are as follows: The score of the second-ranked sample was selected as the threshold. The number of documents with a score of at least τ within the statistical sample. The formula for the rank interval of the descendant is used to solve the statistical verification of the Top-k results.

[0060] Step S3: Randomly select audit sub-samples from the basic sample, perform secondary judgment using the reference large language model, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate a model credibility judgment conclusion.

[0061] This step corresponds completely. Figure 2The entire execution process of the validation layer is as follows: the input source of the validation layer is the sample subset S and confidence interval output by the statistical layer. Internally, the process sequentially completes the selection of audit subsets, reference model evaluation, consistency calculation, and security threshold comparison and determination. The complete audit verification logic is intuitively displayed on [the platform / information]. Figure 3 (b) shows the audit execution diagram, and Figure 3 The (a) no-audit model in the text forms a clear contrast.

[0062] First, perform the audit subset selection operation, corresponding to... Figure 2 The verification layer audit subset selection module function obtains the complete basic sample set from step S2. Internally, m documents are randomly selected and integrated into an audit subset A, which is the audit subsample; the size m of the audit subset is not randomly set, but is subject to a minimum audit budget constraint. ,in, The total number of documents contained in the audit subset; α is the global significance level; A pre-defined model safety threshold is set for the user; this constraint ensures that the low-cost model maintains a high degree of consistency with the reference model. When the system incorrectly determines the result as safe, the probability of this happening will not exceed α, thus avoiding the risk of audit failure from a statistical perspective.

[0063] After selecting the audit subset, the reference large language model, which has higher inference accuracy and higher computational cost, is invoked. Re-execute the semantic predicate secondary judgment operation on all documents within the audit subset A, outputting a new secondary judgment result, and compare this result with the low-cost model in step S2. The generated initial evaluation tags are compared one-to-one, and the number of documents that completely match the two sets of evaluation results is counted and denoted as the number of consistent documents 'a'. Figure 3 (b) Consistency column in table (b): The table displays both document number and document ID. Preliminary evaluation results The audit results are divided into four categories: secondary evaluation results, consistency matching indicators, etc. An equal sign indicates that the two evaluation results are consistent, an inequality sign indicates that the evaluation results are different, and documents that were not audited have no reference model evaluation records and are not included in the consistency statistics. The semantic filtering operator audit implementation case provides specific audit data. The audit subset size is m=30. After completing two rounds of model evaluation comparison, the number of consistent documents is a=27.

[0064] The lower bound of consistency of the evaluation results of the two models is calculated based on the number of consistent documents *a* and the total size of the audit subset *m*. The solution is obtained using the Clopper-Pearson one-sided binomial lower bound, and the calculation formula is as follows:

[0065] in, This is a one-sided confidence lower bound for the agreement rate between the low-cost model and the reference model in the evaluation results. The number of documents with identical labels in both rounds of model evaluation within the audit subset; This refers to the total size of the audit subset documents.

[0066] Substituting parameters into the semantic filtering operator audit case, a consistency lower bound is obtained. .

[0067] Obtaining a consistent lower bound Then, compare this value with the user-preset security threshold. Perform numerical comparisons to complete the security determination logic, corresponding to Figure 2 The decision branch at the end of the verification layer Figure 3 (b) Right-hand side judgment process, if satisfied The numerical relationship and the branch output of the safety judgment conclusion indicate that the overall evaluation bias of the low-cost large language model is controllable and the evaluation result has reliable validity; if The branch outputting an unsafe conclusion indicates that the evaluation bias of the cheap model exceeds the allowable range, and the confidence interval output by the statistical layer is at risk of distortion. The semantic filtering operator case sets a safety threshold. Calculations yielded If the value is less than the safety threshold, an unsafe result is generated.

[0068] In this embodiment, when the security determination is secure, an additional information interval expansion correction step is added, based on a consistency lower bound. Original confidence intervals for statistical layer Boundary widening is performed to offset the evaluation bias between the cheap model and the reference model, ensuring that the expanded interval fully covers the true answer evaluated by the reference model across all documents. The interval expansion calculation formula is as follows:

[0069]

[0070] in, This represents the lower limit of the confidence interval for the overall matched document count after dilation correction. This represents the upper limit of the confidence interval for the total number of matching documents after dilation correction; The boundary of the uncorrected original confidence interval for the statistical layer; N is the total number of documents in the finite population; The lower bound of the model consistency rate obtained in step S3; The function is a truncation function. After constraint inflation, the interval boundary is within the range of valid document counts from 0 to N.

[0071] Step S4: Calculate the point estimate of the query result based on the evaluation record, summarize and output the point estimate, the confidence interval, and the credibility judgment conclusion in a unified manner, and complete the entire process of statistical estimation, model bias verification, and credibility judgment for large-scale document semantic retrieval.

[0072] The embodiments of this application correspond to Figure 2 The output node at the bottom of the architecture eventually outputs three core types of content: query result point estimation, confidence interval obtained from the statistical layer, and security judgment result generated by the verification layer.

[0073] The complete logic for deriving the estimated value of the query result points relies on the full set of evaluation records generated in step S2. First, the number of samples matching the semantic predicate in all documents within the basic sample set, X, is counted. Then, the ratio of the number of samples matching X to the total size n of the basic sample set is calculated. This ratio is considered an unbiased estimate of the proportion of documents matching semantic predicates within a finite population. This ratio estimate is then multiplied by the total number of documents N in the finite population to obtain a point estimate of the total number of documents matching semantic predicates within the finite population. In the semantic filtering operator case, the point estimate is 1500.

[0074] When integrating all the information to be output, two output scenarios are distinguished: in the no-audit scenario, only the point estimate and the original confidence interval are output. Figure 3 (a) Output content at the end of the process; In scenarios where the audit verification layer is enabled, the point estimate, the confidence interval after inflation correction or the original uncorrected confidence interval, and the two types of credibility judgment conclusions (safe / unsafe) are output synchronously. Figure 3 (b) The complete audit process output results, the semantic filtering operator case finally outputs three items: point estimate 1500, original confidence interval [1196,1836], and insecurity determination.

[0075] Once the entire method is executed, it will only generate n calls to the inexpensive large language model and m calls to the reference large language model. It does not require performing large model inference on all N documents within a finite population. In the semantic filtering operator case, with a total number of documents N=5000, only 200 calls to the inexpensive model and 30 calls to the reference model are made. Compared with the full document traversal inference scheme, it saves 95.4% of the computational cost of large model inference. In the business scenario of semantic approximate query of massive text databases, it can significantly reduce the computational cost of semantic retrieval under the premise of strict statistical coverage and auditable security verification capabilities.

[0076] Example 2 This invention provides an auditable approximate query processing system oriented towards semantic predicates. Figure 4This is a schematic diagram illustrating the structure of an auditable approximate query processing system based on semantic predicates, as provided in an embodiment of the present invention. Figure 4 As shown, the system includes: The overall sampling module 100 is used to obtain the semantic query predicate, confidence level parameters and sampling workload constraints specified by the user, thereby determining the finite population consisting of all documents to be retrieved, and performing uniform random sampling without replacement on the finite population to obtain the basic sample set. Interval estimation module 200 is used to perform semantic predicate determination on each document in the basic sample set using a cheap large language model, generate evaluation records with initial evaluation tags, and construct confidence intervals for the total number of conformities based on the obtained confidence level parameters by using the exact binomial interval method and combining it with a finite population correction factor. The audit verification module 300 is used to randomly select audit sub-samples from the basic sample, perform secondary judgment using the reference large language model, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate a model credibility judgment conclusion. The result aggregation module 400 is used to calculate the point estimate of the query result based on the evaluation record, and to aggregate and output the point estimate, the confidence interval and the credibility judgment conclusion in a unified manner, thereby completing the entire process of statistical estimation, model bias verification and credibility judgment for large-scale document semantic retrieval.

[0077] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0078] Example 3 To implement the methods of the above embodiments, the present invention also provides an electronic device, which includes a memory and a processor; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.

[0079] Example 4 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.

[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0081] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0082] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. An auditable approximate query processing method oriented towards semantic predicates, characterized in that, include: Obtain the semantic query predicate, confidence level parameters, and sampling workload constraints specified by the user, determine the finite population consisting of all documents to be retrieved, and perform uniform random sampling without replacement on the finite population to obtain the basic sample set. A cheap large language model is used to perform semantic predicate judgment on each document in the basic sample set to generate an evaluation record with initial evaluation tags. Based on the evaluation record, the exact binomial interval method is used in combination with a finite population correction factor to construct a confidence interval for the population conformity number using the obtained confidence level parameters. Randomly select audit subsamples from the basic sample, use the reference large language model for secondary judgment, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate the model credibility judgment conclusion. Based on the evaluation records, the point estimates of the query results are calculated, and the point estimates, the confidence intervals, and the credibility judgment conclusions are summarized and output in a unified manner to complete the entire process of statistical estimation, model bias verification, and credibility judgment for large-scale document semantic retrieval.

2. The method according to claim 1, characterized in that, A uniform random sampling without replacement is performed on the finite population to obtain the basic sample set, including: The size of the basic sample set is determined based on the constraints of the sampling workload. With the principle that each document in the finite population has an equal probability of being selected, documents are extracted one by one from the finite population. After each extraction, the document is removed from the set to be extracted until the number of extracted documents reaches the determined size. All the extracted documents are then summarized into the basic sample set.

3. The method according to claim 2, characterized in that, The process of constructing confidence intervals for the population conformity count includes: The number of documents in the basic sample set that are determined to conform to the semantic predicate is taken as the number of positive examples, and the ratio of the number of positive examples to the size of the basic sample set is calculated as the sample positive example ratio. The inverse function of the regularized incomplete beta function is used to calculate the lower and upper limits of the binomial proportion interval at a given significance level. Based on the total number of documents in the finite population and the size of the basic sample set, calculate the finite population correction factor, and apply the correction factor to the lower and upper limits of the binomial proportion interval respectively, so that the binomial proportion interval shrinks according to the sampling ratio. Multiply the lower and upper limits of the contracted proportional interval by the total number of documents in the finite population, and truncate to the valid range of document counts to obtain the confidence interval in terms of document counts.

4. The method according to claim 3, characterized in that, The specific method for calculating the finite population correction factor is as follows: Using the total number of documents in the finite population and the size of the basic sample set as input, calculate the ratio of the size of the basic sample set to the total number of documents in the finite population, and use this ratio as the sampling proportion; The value of the finite population correction factor is determined based on the sampling ratio, such that the value of the finite population correction factor decreases as the sampling ratio increases.

5. The method according to claim 4, characterized in that, The process of calculating the lower bound of consistency between the inexpensive model and the reference model includes: The number of documents in the audit subsample whose initial evaluation markers of the cheap large language model are consistent with the secondary judgment results of the reference large language model is counted as the number of consistent positive examples. The inverse function of the regularized incomplete beta function is used as input to construct a one-sided confidence lower bound of the consistency rate, with the number of consistent positive examples, the number of documents judged to be inconsistent in the audit subsample plus one, and a given significance level. The one-sided confidence lower bound is used as the consistency lower bound between the cheap large language model and the reference large language model.

6. The method according to claim 5, characterized in that, The process of estimating the point estimate of the query result based on the evaluation record includes: The number of documents in the complete sample evaluation record that are determined to conform to the semantic predicate is used as the sample conformity number; The proportion of the number of samples that match the semantic predicates is calculated as the proportion of the number of documents that match the semantic predicates in the finite population. Multiplying the estimated value by the total number of documents in the finite population yields a point estimate of the total number of documents in the finite population that conform to the semantic predicate.

7. An auditable approximate query processing system oriented towards semantic predicates, characterized in that, include: The overall sampling module is used to obtain the semantic query predicate, confidence level parameters and sampling workload constraints specified by the user, thereby determining the finite population consisting of all documents to be retrieved, and performing uniform random sampling without replacement on the finite population to obtain the basic sample set. The interval estimation module is used to perform semantic predicate determination on each document in the basic sample set using an inexpensive large language model, generate evaluation records with initial evaluation tags, and construct confidence intervals for the total number of conformities based on the obtained confidence level parameters by using the exact binomial interval method and combining it with a finite population correction factor. The audit verification module is used to randomly select audit sub-samples from the basic sample, perform secondary judgment using the reference large language model, compare the secondary judgment results with the initial evaluation marks item by item, calculate the lower bound of consistency between the low-cost model and the reference model, and compare the lower bound of consistency with the preset security threshold to generate a model credibility judgment conclusion. The results aggregation module is used to calculate the point estimate of the query results based on the evaluation records, and to aggregate and output the point estimate, the confidence interval, and the credibility judgment conclusion in a unified manner, thereby completing the entire process of statistical estimation, model bias verification, and credibility judgment for large-scale document semantic retrieval.

8. The system according to claim 7, characterized in that, The overall sampling module is also used for: The size of the basic sample set is determined based on the constraints of the sampling workload. With the principle that each document in the finite population has an equal probability of being selected, documents are extracted one by one from the finite population. After each extraction, the document is removed from the set to be extracted until the number of extracted documents reaches the determined size. All the extracted documents are then summarized into the basic sample set.

9. An electronic device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.