Knowledge retrieval candidate library generation method and system based on incremental pre-training optimization

By analyzing the feature vectors of user retrieval behavior, constructing a mapping function between normalized co-occurrence frequency and incremental pre-training level, and dynamically adjusting the knowledge retrieval candidate library generation process, the problems of insufficient semantic coverage and resource waste in traditional methods are solved, and efficient and accurate knowledge retrieval candidate library generation is achieved.

CN120670565AActive Publication Date: 2025-09-19SHANGHAI ANSHUO ENTERPRISE CREDIT SERVICE CO LTD

Patent Information

Application Number
CN202511140546.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-19
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Traditional knowledge retrieval candidate library generation methods are unable to dynamically adjust the incremental pre-training level based on user retrieval behavior characteristics, resulting in insufficient semantic coverage, waste of computing resources and inaccurate generated candidate libraries.

Method used

By collecting the current user's search sentences, analyzing and obtaining the search behavior feature vector, calculating the normalized co-occurrence frequency, constructing its mapping function relationship with the incremental pre-training level, and using the matching incremental pre-training level to perform multi-level expansion on the sentences, a knowledge retrieval candidate library is generated.

Benefits of technology

It achieves precise expansion of retrieval semantics and efficient generation of candidate libraries, improves the accuracy and efficiency of knowledge retrieval, and optimizes computing resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670565A_ABST
    Figure CN120670565A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge retrieval candidate library generation method and system based on incremental pre-training optimization, and relates to the technical field of information retrieval, the method comprises the following steps: collecting a first retrieval statement and analyzing to obtain a retrieval behavior feature vector; calculating to obtain a normalized co-occurrence frequency; constructing a mapping function relationship to obtain a matching increment pre-training stage number; and executing multi-stage extension of the first retrieval statement by using the statement increment optimization model to output an extended retrieval statement, and further generating a knowledge retrieval candidate library. The technical problems that in electric digital data processing, a traditional knowledge retrieval candidate library generation method cannot dynamically adjust the increment pre-training series based on user retrieval behavior characteristics, so that semantic coverage is insufficient, computing resources are wasted, and candidate library generation is inaccurate are solved; the technical effects of semantic coverage range expansion, computing resource optimization configuration and candidate library accurate generation are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a method and system for generating a knowledge retrieval candidate library based on incremental pre-training optimization. Background Art

[0002] In the process of processing electronic digital data, the generation of candidate libraries for knowledge retrieval is crucial for accurate information acquisition. Existing technologies mainly generate candidate libraries by expanding search statements through fixed pre-trained models, which has played a role in stable retrieval scenarios. However, with the diversification of retrieval requirements and the improvement of accuracy, traditional methods have shown limitations in the application of electronic digital data processing. Because traditional methods cannot dynamically adjust the number of incremental pre-training levels based on user search behavior characteristics, in the electronic digital data processing process, the candidate library has insufficient semantic coverage, computing resources are wasted, and the generated candidate library is inaccurate, making it difficult to meet the requirements of accurate knowledge retrieval evaluation and efficient application. Summary of the Invention

[0003] The present application provides a method and system for generating a knowledge retrieval candidate library based on incremental pre-training optimization, which is used to solve the technical problems in electronic digital data processing that the traditional knowledge retrieval candidate library generation method cannot dynamically adjust the incremental pre-training level based on the user's retrieval behavior characteristics, resulting in insufficient semantic coverage, waste of computing resources and inaccurate candidate library generation.

[0004] The first aspect of the present application provides a method for generating a knowledge retrieval candidate library based on incremental pre-training optimization, the method comprising: collecting the first retrieval statement of the current user, calling the historical retrieval statement behavior library to analyze the first retrieval statement, and obtaining a retrieval behavior feature vector; calculating the retrieval behavior feature vector to obtain a normalized co-occurrence frequency; constructing a mapping function relationship between the normalized co-occurrence frequency sample and the incremental pre-training level sample, and using the mapping function relationship to obtain the matching incremental pre-training level corresponding to the normalized co-occurrence frequency; pre-training a statement incremental optimization model, the statement incremental optimization model using the matching incremental pre-training level to perform multi-level expansion on the first retrieval statement, outputting an expanded retrieval statement, and generating a knowledge retrieval candidate library based on the expanded retrieval statement.

[0005] The second aspect of the present application provides a knowledge retrieval candidate library generation system based on incremental pre-training optimization, the system including: a retrieval behavior feature vector acquisition module, used to collect the first retrieval statement of the current user, call the historical retrieval statement behavior library to analyze the first retrieval statement, and obtain the retrieval behavior feature vector; a normalized co-occurrence frequency acquisition module, used to calculate the retrieval behavior feature vector and obtain the normalized co-occurrence frequency; a mapping function relationship construction module, used to construct a mapping function relationship between the normalized co-occurrence frequency sample and the incremental pre-training level sample, and use the mapping function relationship to obtain the matching incremental pre-training level corresponding to the normalized co-occurrence frequency; a knowledge retrieval candidate library construction module, used to pre-train a statement incremental optimization model, the statement incremental optimization model uses the matching incremental pre-training level to perform multi-level expansion on the first retrieval statement, outputs the expanded retrieval statement, and generates a knowledge retrieval candidate library based on the expanded retrieval statement.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application collects the first search statement of the current user, calls the historical search statement behavior library to analyze and obtain the search behavior feature vector, calculates the normalized co-occurrence frequency and constructs a mapping function relationship between it and the incremental pre-training level, uses the matching incremental pre-training level to perform multi-level expansion of the statement incremental optimization model, outputs the expanded search statement and generates a knowledge retrieval candidate library, thereby realizing the precise expansion of the search semantics and the efficient generation of the candidate library, improving the accuracy and efficiency of knowledge retrieval, and achieving the technical effects of expanding the semantic coverage, optimizing the configuration of computing resources and accurately generating the candidate library. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0009] Figure 1 It is a flow chart of the method for generating a knowledge retrieval candidate library based on incremental pre-training optimization provided in an embodiment of the present application.

[0010] Figure 2 It is a structural diagram of a knowledge retrieval candidate library generation system based on incremental pre-training optimization provided in an embodiment of the present application.

[0011] Explanation of the accompanying symbols: retrieval behavior feature vector acquisition module 1, normalized co-occurrence frequency acquisition module 2, mapping function relationship construction module 3, knowledge retrieval candidate library construction module 4. DETAILED DESCRIPTION

[0012] The present application provides a method and system for generating a knowledge retrieval candidate library based on incremental pre-training optimization, which is used to solve the technical problems in electronic digital data processing that the traditional knowledge retrieval candidate library generation method cannot dynamically adjust the incremental pre-training level based on the user's retrieval behavior characteristics, resulting in insufficient semantic coverage, waste of computing resources and inaccurate candidate library generation.

[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0015] Example 1, as Figure 1 As shown, a method for generating a knowledge retrieval candidate library based on incremental pre-training optimization, wherein the method includes:

[0016] Step A100: Collect the first search statement of the current user, call the historical search statement behavior library to analyze the first search statement, and obtain a search behavior feature vector.

[0017] In this embodiment of the present application, the first search statement is the pending search statement entered by the current user and is used to trigger the process of generating a knowledge retrieval candidate library. The historical search statement behavior library is a database that stores historical user search behavior data, including valid historical search statements and their related features within a preset period, and is used to analyze the search behavior characteristics of the first search statement.

[0018] Specifically, the original sentence currently entered by the user to trigger knowledge retrieval is first obtained. This sentence serves as the starting point for the entire knowledge retrieval candidate library generation process. It is actively entered by the user through an interactive interface (such as a search box) and can include keywords, phrases, or complete sentences, such as the application of artificial intelligence in science and technology. This sentence is captured in real time through the interface, providing the raw data foundation for subsequent calls to the historical search behavior library for feature analysis, generation of expanded search sentences, and construction of the candidate library.

[0019] Next, a preset period is obtained and valid historical search statement behaviors within the period are extracted from the historical search statement behavior library, and the first search statement is analyzed to obtain a search behavior feature vector. The specific steps are described in detail in A110-A120.

[0020] By collecting the user's first search statement and deeply analyzing the characteristics of historical search behavior, we can accurately characterize the user's search needs, laying a data foundation for subsequent dynamic matching of incremental pre-training levels and efficient generation of knowledge retrieval candidate libraries.

[0021] Step A200: Calculate the search behavior feature vector to obtain a normalized co-occurrence frequency.

[0022] In the embodiment of the present application, the normalized co-occurrence frequency is a quantitative value obtained by mapping the feature fusion index to the target interval.

[0023] Optionally, a weight calculation is performed on the retrieval behavior feature vector including the retrieval frequency feature vector representing the historical retrieval frequency, the generation time feature vector representing the time required to generate the retrieval candidate library, and the retrieval feedback index vector representing the number of feedbacks of the retrieval candidate library, to obtain the feature fusion index and map it to the target interval, thereby obtaining the normalized co-occurrence frequency. The specific steps are described in detail in A210-A220.

[0024] Step A300: constructing a mapping function relationship between normalized co-occurrence frequency samples and incremental pre-training level samples, and using the mapping function relationship to obtain a matching incremental pre-training level corresponding to the normalized co-occurrence frequency.

[0025] In an embodiment of the present application, the incremental pre-training level sample is a pre-set set of different levels, which is used to establish a mapping relationship with the normalized co-occurrence frequency sample, and the optimal level solution corresponding to each co-occurrence frequency is obtained by calculating the scoring value, thereby realizing dynamic matching of the optimal incremental pre-training level according to the user's retrieval behavior characteristics.

[0026] In one embodiment of the present application, first, the scoring values ​​of the normalized co-occurrence frequency samples and the incremental pre-training level samples are calculated (the information entropy mean for generating quality and computing resource scores for the candidate library), the incremental pre-training level optimal solution of each co-occurrence frequency sample is obtained and the functional relationship is fitted, thereby constructing a mapping relationship between the two. The specific steps are described in detail in A310-A330.

[0027] The steps for obtaining the matching incremental pre-training levels using the mapping function relationship are as follows: First, based on the constructed normalized co-occurrence frequency samples, such as {0.1, 0.3, 0.5, 0.7, 0.9}, and the incremental pre-training level samples, such as {levels 1-5}, the score value of each frequency and level combination is calculated, that is, the information entropy mean of the candidate library generation quality and computing resource score, to determine the optimal level solution corresponding to each frequency.

[0028] Secondly, we use the constructed exponential decreasing function model (such as ) fits the mapping relationship between frequency and level, optimizing parameters using the least squares method to converge the mean squared error between the model output and the optimal solution to below 0.1. When a user's normalized co-occurrence frequency of 0.6 is input, the model calculates y = 5·e^(-2.3×0.6)+0.5≈5×0.257+0.5≈1.785, rounded to 2 levels. This reduces computational time compared to the traditional fixed 5-level approach, while maintaining a semantic query hit rate of over 85%.

[0029] By constructing a dynamic mapping function based on user retrieval behavior characteristics, we achieve accurate matching of normalized co-occurrence frequency and incremental pre-training levels, solving the problems of insufficient semantic coverage and resource waste caused by fixed levels in existing technologies, and providing a quantitative level matching solution for the efficient generation of knowledge retrieval candidate libraries.

[0030] Step A400: pre-training a statement incremental optimization model, wherein the statement incremental optimization model uses the matching incremental pre-training level to perform multi-level expansion on the first search statement, outputs the expanded search statement, and generates a knowledge search candidate library based on the expanded search statement.

[0031] In the embodiment of the present application, the expanded search statement refers to inputting the first search statement into the statement incremental optimization model, extracting the search semantic vector, performing L rounds of expansion iterations according to the matching incremental pre-training level, and calculating the semantic similarity to output the relevant search statement.

[0032] Specifically, the pre-trained sentence incremental optimization model takes as input an initial search sentence and a matching incremental pre-training level L. The output is a set of expanded search sentences generated after L rounds of semantic expansion iterations, which are used to construct a knowledge retrieval candidate library. This construction process, based on pre-trained language models (such as BERT and T5), first extracts the semantic vector of the initial search sentence through an encoder. It then integrates a semantic expansion module and a multi-task learning mechanism. The semantic expansion module consists of L rounds of iterative computations. Each round uses cosine similarity or contrastive learning to generate semantic neighborhoods and synonymous semantic neighborhoods, selects related entities to generate search sentences, and merges redundancies. The multi-task learning module integrates quality and resource scores, dynamically adjusting weights through a gating network such as Multi-Method Optimization (MMoE) to balance search quality and computational resource consumption. The training process utilizes an incremental pre-training strategy, combining new and old data to adjust the learning rate (e.g., cosine annealing) and adding original data to prevent forgetting. Contrastive learning (e.g., SimCSE and DiffCSE) is also introduced to optimize the semantic vector representation. The InfoNCE loss function is used to bring similar semantics closer together and push dissimilar semantics apart. In addition, the model integrates the mean information entropy through dynamic weight optimization (such as based on gradient or training progress), accurately controlling the quality of the candidate library while improving the search coverage, and ultimately outputting an expanded set of search statements containing rich semantic dimensions.

[0033] Next, the first search statement is input into the sentence incremental optimization model to extract the search semantic vector, and then L rounds of expansion iterations are performed according to the matching incremental pre-training level to obtain the relevant search statements. Then, the semantic vectors are extracted to calculate the semantic similarity, and the statements greater than the preset value are merged and output as the expanded search statements. The specific steps are described in detail in A410-A430.

[0034] Finally, based on these extended search statements, by integrating semantically related multi-search expressions, a knowledge retrieval candidate library with rich semantic dimensions is constructed. This process achieves a systematic expansion of the search coverage and a precise improvement of the candidate library quality through dynamic expansion and semantic optimization.

[0035] By dynamically matching incremental pre-training levels based on user retrieval behavior, combining a multi-level expansion mechanism of semantic neighborhoods and synonymous neighborhoods, and semantic similarity merging optimization, accurate expansion of retrieval statements is achieved, solving the problems of insufficient semantic coverage and resource waste caused by fixed levels in existing technologies, and providing a dynamically optimized model solution for the efficient generation of knowledge retrieval candidate libraries.

[0036] Furthermore, step A200 in the method provided in the embodiment of the present application includes:

[0037] A210: Perform weight calculation on the retrieval behavior feature vector to obtain a feature fusion index, map the feature fusion index to a target interval, and obtain a normalized co-occurrence frequency.

[0038] A220: The retrieval behavior feature vector includes a retrieval frequency feature vector representing the historical retrieval frequency, a generation time feature vector representing the time required to generate the retrieval candidate library, and a retrieval feedback index vector representing the number of feedbacks on the generated retrieval candidate library.

[0039] Specifically, first, a three-dimensional feature vector is extracted, including a retrieval frequency feature vector, a generation time feature vector, and a retrieval feedback index vector. For example, in the past 30 days, 12 retrievals were performed, each generation took an average of 3.5 seconds, and there were 9 valid feedbacks. Second, a linear weighted calculation is performed using a preset weighting system (e.g., 0.4 for retrieval frequency, 0.3 for generation time, and 0.3 for retrieval feedback; the specific values ​​are determined by those skilled in the art based on actual conditions). The feature fusion index is calculated as follows: fusion index = 0.4 × normalized frequency value + 0.3 × normalized duration value + 0.3 × normalized feedback value.

[0040] The specific conditions for determining the standardized values ​​of retrieval frequency, generation time, and retrieval feedback indicators are as follows: First, for the retrieval frequency feature vector, based on the historical number of retrievals within a preset period (such as the past 30 days), the minimum-maximum normalization method is used to map the actual number of retrievals to the [0,1] interval. For example, when the minimum value of the number of retrievals in the past 30 days is 0 and the maximum value is 15, the standardized value of 12 retrievals is (12-0) / (15-0)=0.8. Secondly, for the generation time feature vector, considering that shorter generation times are more efficient, we first convert the time data into an efficiency metric (e.g., 1 / generation time). Then, we normalize the vector using the minimum and maximum times within a preset period as the boundaries. For example, if the minimum time is 2 seconds and the maximum is 5 seconds, the average time of 3.5 seconds corresponds to an efficiency value of 1 / 3.5≈0.286. After normalization, this value becomes (0.286-1 / 5) / (1 / 2-1 / 5)=(0.286-0.2) / 0.3≈0.086 / 0.3≈0.29. The example data here can be adjusted based on actual conditions to ensure that the normalized value is within the [0,1] interval. Finally, for the retrieval feedback metric vector, we directly perform minimum-maximum normalization based on the number of valid feedback within the preset period. If the minimum number of feedback is 0 and the maximum is 10, the normalized value for 9 valid feedbacks is 9 / 10=0.9. The above standardization processes are based on the extreme values ​​within the preset period to ensure that the numerical range of each eigenvector is unified and provide standardized input for subsequent weight calculations.

[0041] Taking the above data as an example, the normalized frequency is 0.8, the normalized duration is 0.29, and the normalized feedback is 0.9, resulting in a calculated feature fusion index of 0.677. Finally, a linear mapping function is used to map the fusion index to the target interval [0, 1], resulting in a normalized co-occurrence frequency of 0.677. This value represents the comprehensive frequency characteristics of user search behavior. Since the weight sum is 1 and the normalized values ​​are all ≤ 1, the fusion index is naturally limited to the range [0, 1]. Therefore, the linear mapping function is used directly as the normalized co-occurrence frequency, which directly represents the comprehensive frequency characteristics of user search behavior.

[0042] By performing multi-dimensional weight calculation and target interval mapping on the retrieval behavior feature vector, an accurate quantitative representation of the user's retrieval behavior is achieved, providing a scientific numerical basis for the subsequent construction of a dynamic matching relationship between behavioral indicators and incremental pre-training levels.

[0043] Furthermore, step A300 in the method provided in the embodiment of the present application includes:

[0044] A310: Calculate the score value of each co-occurrence frequency sample in the normalized co-occurrence frequency sample and each incremental pre-training level in the incremental pre-training level sample, and output the score value sample.

[0045] A320: The score value is the information entropy average of the candidate library generation quality score value and the candidate library required computing resource score value.

[0046] A330: According to the score value sample, obtain the incremental pre-training level optimal solution of each co-occurrence frequency sample in the normalized co-occurrence frequency sample, fit the functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution, and output the mapping function relationship.

[0047] In this embodiment of the present application, the co-occurrence frequency sample is a normalized set of co-occurrence frequency values ​​used to construct a mapping function relationship with the incremental pre-training level sample. The candidate library is a knowledge retrieval candidate library generated based on the expanded search statement. The expanded search statement is generated by the statement incremental optimization model using matching incremental pre-training levels to perform multi-level expansion on the first search statement.

[0048] Optionally, in knowledge retrieval, existing technologies usually use a fixed incremental pre-training level when constructing a retrieval candidate library, such as uniformly setting it to 3 or 5 levels, without dynamic adjustment based on user retrieval behavior characteristics. This results in an increase in the semantic coverage loss rate due to insufficient levels in low-frequency scenarios with normalized co-occurrence frequency (such as <0.3), and a waste of computing resources in high-frequency scenarios (such as >0.7) due to excessively high levels.

[0049] To address the above issues, first, we establish a normalized co-occurrence frequency sample set, such as {0.1, 0.3, 0.5, 0.7, 0.9} and an incremental pre-training level sample set, such as {1, 2, 3, 4, 5}. Taking the normalized co-occurrence frequency of 0.1 and level 5 as an example, when calculating the score value, we first obtain the candidate library to generate the quality score value (the specific steps are detailed in A321-A322), such as the semantic query hit rate of 85%, the consistency score annotation information of 4.2 / 5, the semantic relevance weight of 0.7, the comprehensive quality score of 0.85×0.4+0.84×0.3+0.7×0.3=0.812 and the computing resource score value, such as GPU occupancy of 60% and time consumption of 1. 2 seconds, the resource score after standardization is 0.4, the information entropy of the two are H(quality) = -0.812×ln0.812-0.188×ln0.188≈0.36, H(resource) = -0.4×ln0.4-0.6×ln0.6≈0.67, the mean information entropy is (0.36+0.67) / 2≈0.515, that is, the score of the combination is 0.515, the calculation formula of information entropy can be simply described as: for a set of probability distributions ,in , its information entropy is In a binary distribution scenario, such as the probability distribution of the quality score and resource score above, the formula is simplified to ,in , and These correspond to the proportion of the candidate library generation quality score and the candidate library required computing resource score in the probability distribution.

[0050] Next, we traverse all combinations of normalized co-occurrence frequencies and incremental pre-training levels to obtain a sample matrix of rating values. For example, when the co-occurrence frequency is 0.7, the rating of level 2 is calculated to be 0.78, with a quality score of 0.92, a resource score of 0.85, and an information entropy mean of 0.81, the highest rating at this frequency. Therefore, level 2 is determined to be the optimal solution.

[0051] Based on all the optimal solution data, an exponentially decreasing function model was used for fitting, and a mapping relationship was obtained, such as 0.1 corresponding to level 5 and 0.9 corresponding to level 1, as shown in Table 1. The construction process of the exponentially decreasing function model is described in detail in A331.

[0052] By quantitatively analyzing the combined score of normalized co-occurrence frequency and incremental pre-training levels, combined with the information entropy mean evaluation mechanism and exponential fitting model, dynamic matching of retrieval behavior characteristics and pre-training resources is achieved, solving the problems of insufficient semantic coverage and resource waste caused by fixed levels in existing technologies, and providing a scientific mapping basis for the accurate generation of knowledge retrieval candidate libraries.

[0053] Table 1: Mapping table of normalized co-occurrence frequency samples and incremental pre-training levels

[0054]

[0055] Furthermore, step A330 in the method provided in the embodiment of the present application includes:

[0056] A331: Fitting a functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution through a fitting model, wherein the fitting model is an exponentially decreasing function model.

[0057] Specifically, an exponential decreasing function model is used to fit the mapping function relationship between the co-occurrence frequency samples and the corresponding incremental pre-training level optimal solution, which is in the form of , where α, β, and γ are fitting parameters, x is the normalized co-occurrence frequency, and y is the incremental pre-training level. The specific steps are as follows: First, based on the previously constructed score sample, obtain optimal level solutions corresponding to different normalized co-occurrence frequencies. For example, sample points (0.1, 5), (0.3, 4), (0.5, 3), (0.7, 2), and (0.9, 1) are collected. Second, the least squares method is used to optimize the parameters of the exponentially decreasing function, with the goal of minimizing the mean squared error of the sample points, and the values ​​of α, β, and γ are iteratively adjusted. Assuming that the initial parameters are α=4, β=2, and γ=1, calculate the error between the predicted level and the actual optimal solution. For example, when x=0.1, the predicted y=4·e^(-2×0.1)+1≈4×0.8187+1≈4.275, and the error with the actual optimal solution 5 is 0.725. Update the parameters to α=5, β=2.3, and γ=0.5 through the gradient descent method. At this time, when x=0.1, y=5·e^(-2.3×0.1)+0.5≈5×0.7945+0.5≈4.472, and the error is reduced to 0.528, eventually making the overall mean square error converge to below 0.1.

[0058] The model takes as input the normalized co-occurrence frequency in the range [0, 1], and outputs the matching incremental pre-training level, a positive integer. This model uses exponentially decreasing properties to achieve a dynamic mapping where higher frequencies correspond to lower levels. For example, when the input x = 0.2, the model outputs y ≈ 5·e^(-2.3×0.2)+0.5 ≈ 5×0.6387+0.5 ≈ 3.693, rounded to 4 levels. This improves the semantic query hit rate compared to the existing fixed level of 3. When x = 0.8, the model outputs y ≈ 5·e^(-2.3×0.8)+0.5 ≈ 5×0.1609+0.5 ≈ 1.304, rounded to 1 level, which reduces computational resource consumption compared to the fixed level of 5.

[0059] By constructing an exponentially decreasing function model and performing parameter fitting based on the scored optimal solution samples, nonlinear dynamic matching of normalized co-occurrence frequency and incremental pre-training levels is achieved, which solves the problems of insufficient semantic expansion or resource waste caused by fixed levels in existing technologies, and provides a quantitative mapping basis for the efficient generation of knowledge retrieval candidate libraries.

[0060] Furthermore, step A400 in the method provided in the embodiment of the present application includes:

[0061] A410: Input the first search statement into the statement incremental optimization model to extract the search semantic vector.

[0062] A420: Perform L rounds of expansion iterations on the search semantic vector according to the matching incremental pre-training levels to obtain L rounds of related search statements, where L is the number of matching incremental pre-training levels.

[0063] A430: Extracting the semantic vectors of the L rounds of related search statements to perform semantic similarity calculation, merging related search statements with a semantic similarity greater than a preset value, and outputting the processed related search statements as extended search statements.

[0064] Specifically, the model first inputs a search query, such as "Application of artificial intelligence in scientific and technological imaging," and uses the semantic analysis module to extract a search semantic vector containing the keyword. This process can be implemented using an encoder based on a Transformer architecture such as BERT, for example, by using a multi-layer self-attention mechanism to capture semantic associations within the query and generate a semantic vector with a dimension of 768.

[0065] Secondly, the retrieval semantic vector is expanded and iterated for L rounds according to the matching incremental pre-training level, including calculating its semantic neighborhood, selecting a round of semantically related entities from it to generate a round of related retrieval statements, and so on to output L rounds of related retrieval statements. The specific steps are described in detail in A421-A422.

[0066] Furthermore, the extended iterative method also includes obtaining the synonymous semantic neighborhood of the retrieval semantic vector, selecting synonymous semantically related entities therefrom to generate synonymous related retrieval statements, and updating them to the L-round related retrieval statements. The specific steps are described in detail in A423-A425.

[0067] Each iteration generates relevant search terms. For example, after four rounds of expansion, multiple relevant terms can be generated, such as "Application of deep learning in AI judgment of scientific images" and "Combination of scientific image processing and AI technology," covering semantic dimensions such as synonym replacement and field segmentation of the original terms.

[0068] Subsequently, the semantic vectors of all L-round expanded sentences are extracted, and the cosine similarity algorithm is used to calculate the semantic similarity between each sentence. The preset similarity threshold is 0.7, and redundant sentences with a similarity greater than the threshold are merged. For example, 5 of the 20 expanded sentences may be merged due to high semantic similarity (such as "AI technology" and "artificial intelligence technology"), and finally 15 streamlined expanded search sentences are output. This merging mechanism can improve the semantic query hit rate in low-frequency retrieval scenarios and reduce computing resource consumption in high-frequency retrieval scenarios (such as a normalized co-occurrence frequency of 0.8). By dynamically adjusting the expansion level and semantic similarity merging strategy, the retrieval efficiency in specific fields is significantly improved while maintaining the generalization ability of the model.

[0069] By building a dynamic expansion mechanism based on user search behavior characteristics, combined with semantic vector extraction, multi-level iterative expansion, and similarity merging optimization, we achieve precise expansion of search statements. Compared to the static expansion model with a fixed number of levels in existing technologies, this solution improves semantic coverage and resource utilization efficiency, providing a dynamically optimized model for efficiently generating knowledge retrieval candidate libraries.

[0070] Furthermore, step A420 in the method provided in the embodiment of the present application includes:

[0071] A421: Calculate a round of semantic neighborhood of the retrieval semantic vector.

[0072] A422: Select a round of semantically related entities from a round of semantic neighborhood of the search semantic vector to generate a round of related search statements, and so on to output L rounds of related search statements for the search semantic vector to perform L rounds of expansion iterations.

[0073] In the embodiments of the present application, a semantic neighborhood is a region adjacent to a search semantic vector in the semantic space, comprising a set of entities semantically related to the vector. Semantically related entities are entities selected from the semantic neighborhood of the search semantic vector that have semantic relevance to the original search statement and are used to generate related search statements.

[0074] Specifically, the first user-entered search phrase, such as "Application of artificial intelligence in scientific and technological imaging," is fed into the sentence incremental optimization model. A BERT-based encoder extracts a retrieval semantic vector of dimension 768. This process utilizes a multi-layer self-attention mechanism to capture semantic associations within the phrase, for example, identifying the domain relevance between "artificial intelligence" and "scientific and technological imaging." Subsequently, the retrieval semantic vector undergoes L rounds of iterative expansion, based on the previously determined number of matching incremental pre-training levels, such as 5 levels, determined by the mapping function.

[0075] In each iteration, the semantic neighborhood of the current semantic vector is first calculated. This process, based on the cosine similarity algorithm, searches for neighboring vectors in the pre-trained semantic space with a similarity greater than 0.6 to the current vector. For example, the first round of expansion might retrieve neighboring entities such as "deep learning" and "scientific image processing" from the knowledge graph, forming a semantic neighborhood containing 20 nodes.

[0076] Next, a round of semantically related entities is selected from this neighborhood, and a round of related search statements is generated using template filling or sequence generation models. For example, combining "deep learning" with "AI diagnosis of scientific imaging" generates multiple candidate statements, such as "Application of deep learning in AI diagnosis of scientific imaging." Repeat this step L times to ultimately generate a set of L expanded statements.

[0077] By constructing a dynamic neighborhood expansion mechanism based on semantic vector space, combined with cosine similarity calculation, knowledge graph entity retrieval and iterative statement generation strategy, a multi-dimensional expansion of retrieval semantics is achieved, providing an intelligent dynamic optimization model for the efficient generation of knowledge retrieval candidate libraries.

[0078] Furthermore, step A420 in the method provided in the embodiment of the present application includes:

[0079] A423: Obtain the synonymous semantic neighborhood of the retrieval semantic vector.

[0080] A424: Selecting synonymous semantically related entities from the synonymous semantic neighborhood of the search semantic vector to generate a synonymous related search statement of the search semantic vector.

[0081] A425: Update the synonymous related search statement to the L-round related search statement.

[0082] In one embodiment, the model first uses pre-trained sentence increment optimization to obtain synonymous semantic neighborhoods for the search semantic vectors. Taking the input sentence "The application of artificial intelligence in scientific and technological imaging" as an example, the model calculates the semantic vector for "scientific and technological imaging" using the BERT word vector model. It then searches for synonymous entities with a cosine similarity greater than 0.7 in the semantic space, such as "scientific imaging" and "image judgment," forming a synonymous semantic neighborhood set containing multiple synonymous entities.

[0083] Secondly, select 2-3 entities with the strongest semantic relevance to the original sentence from the neighborhood (such as "scientific imaging" and "image analysis"), and construct synonymous related search sentences through template generation method, such as "application of artificial intelligence in scientific imaging" and "technological application of AI in image analysis".

[0084] Finally, the newly generated synonymous sentences are updated to the set of related search sentences obtained by L rounds of expansion iterations to achieve dynamic expansion of semantic coverage.

[0085] By obtaining synonymous semantic neighborhoods, generating context-related sentences and updating the extension set, the problems of low synonymous extension coverage and poor adaptability in existing technologies are solved, providing an intelligent extension solution to improve the comprehensiveness and accuracy of the knowledge retrieval candidate library.

[0086] Furthermore, step A320 in the method provided in the embodiment of the present application includes:

[0087] A321: Obtain the semantic query hit rate, consistency score annotation information, and semantic relevance weight of the generated candidate library.

[0088] A322: Outputting a candidate library to generate a quality score value according to the calculation results of the semantic query hit rate, consistency score annotation information, and semantic relevance weight.

[0089] Optionally, first, obtain the semantic query hit rate, that is, search the candidate library through the query statement of the preset test set, and count the ratio of the number of correctly hit queries to the total number of queries, for example, 85 hits out of 100 test queries, and the hit rate is 85%. Secondly, collect consistency score annotation information, which is used by technical personnel in this field to score the semantic consistency of the search results and query statements in the candidate library (such as using a 5-point system). Assuming that the scores of 10 samples are 4, 5, 4, 3, 5, 4, 4, 5, 3, and 4 respectively, the average consistency score is 4.2 points. Finally, determine the semantic relevance weight, which is based on the experience of domain experts or historical data training, and is used to characterize the importance of semantic relevance in quality assessment, for example, set to 0.3.

[0090] During the calculation phase, the above metrics are quantified and integrated. Assuming the semantic query hit rate is standardized to 0.85 (out of a maximum score of 1), the consistency score is standardized to 0.84 (4.2 / 5), and combined with a semantic relevance weight of 0.3, a weighted summation formula is used: Quality Score = Hit Rate × 0.5 + Consistency Score × 0.2 + Relevance Weight × 0.3, i.e., 0.85 × 0.5 + 0.84 × 0.2 + 0.3 × 0.3 = 0.425 + 0.168 + 0.09 = 0.683. This calculation process comprehensively considers retrieval accuracy, result consistency, and semantic relevance, improving the credibility of the quality score compared to existing single-metric evaluation methods.

[0091] By obtaining semantic query hit rate, consistency score annotation information and semantic relevance weight from multiple dimensions, and integrating evaluation indicators based on scientific computational models, a comprehensive quantitative analysis of the quality of candidate library generation is achieved, solving the problems of single evaluation dimension and inaccurate results in existing technologies, and providing a reliable quality reference for optimizing incremental pre-training level matching and improving retrieval effects.

[0092] Furthermore, step A100 in the method provided in the embodiment of the present application includes:

[0093] A110: Get the preset cycle.

[0094] A120: Extracting valid historical search statement behaviors within the preset period from the historical search statement behavior library, analyzing the first search statement based on the valid historical search statement behaviors, and obtaining a search behavior feature vector.

[0095] In this embodiment of the present application, the preset period is a pre-set time range that limits the time span for extracting historical search behavior data. This period can be adjusted based on the changing characteristics of user search habits. Valid historical search statement behavior refers to historical search records that completed the search process within the preset period and generated valid feedback (such as clicks on search results, saved results, etc.). It does not include invalid search behavior caused by network failures, abnormal operation, etc.

[0096] In one embodiment, a preset period is first obtained. The specific data is adjusted by those skilled in the art based on the changing characteristics of user search habits. For example, a period of approximately 30 days is assumed, based on the short-term continuity characteristics of user behavior. Taking the user input "Application of machine learning in recommendation systems" as an example, valid historical search behaviors from May 25, 2025, to June 24, 2025, are extracted from the historical search statement behavior library based on the preset period. Valid behaviors are determined by records that complete the search process and generate feedback (e.g., clicks, favorites), excluding invalid requests due to network failures, etc.

[0097] After the extraction process described above, for example, the user has 12 searches for similar topics in the past 30 days, with an average candidate library generation time of 3.8 seconds each time and 10 relevant result feedbacks. Based on this valid data, the system analyzes the first search statement and converts the search frequency, generation time, and number of feedbacks into a feature vector with dimensions of [12, 3.8, 10]. After normalization, this vector accurately represents the user's current search behavior pattern.

[0098] By setting a preset period and extracting effective historical search behaviors within the period, we can accurately capture the characteristics of the user's current search behavior, providing timely and reliable data support for subsequent normalized co-occurrence frequency calculation and incremental pre-training level matching.

[0099] In summary, the method for generating a knowledge retrieval candidate library based on incremental pre-training optimization provided in the embodiments of the present application has the following technical effects:

[0100] This application collects the first search statement of the current user, calls the historical search statement behavior library to analyze it to obtain the search behavior feature vector, obtains the normalized co-occurrence frequency through weight calculation and mapping processing, constructs a mapping function relationship between the normalized co-occurrence frequency sample and the incremental pre-training level sample and obtains the matching incremental pre-training level, and uses the pre-trained statement incremental optimization model to perform multi-level expansion on the first search statement to generate an extended search statement, thereby accurately generating a knowledge retrieval candidate library, making the generation results of the knowledge retrieval candidate library more accurate and efficient, and achieving the technical effects of expanding the semantic coverage range, optimizing the configuration of computing resources and accurately generating the candidate library.

[0101] Example 2, as Figure 2 As shown, based on the same inventive concept as the aforementioned embodiment 1, the embodiment of the present application provides a knowledge retrieval candidate library generation system based on incremental pre-training optimization, the system comprising:

[0102] The retrieval behavior feature vector acquisition module 1 is used to collect the first retrieval statement of the current user, call the historical retrieval statement behavior library to analyze the first retrieval statement, and obtain the retrieval behavior feature vector.

[0103] The normalized co-occurrence frequency acquisition module 2 is used to calculate the search behavior feature vector to obtain the normalized co-occurrence frequency.

[0104] The mapping function relationship construction module 3 is used to construct a mapping function relationship between the normalized co-occurrence frequency samples and the incremental pre-training level samples, and use the mapping function relationship to obtain the matching incremental pre-training level corresponding to the normalized co-occurrence frequency.

[0105] The knowledge retrieval candidate library construction module 4 is used to pre-train the sentence increment optimization model. The sentence increment optimization model uses the matching incremental pre-training level to perform multi-level expansion on the first retrieval sentence, outputs the expanded retrieval sentence, and generates a knowledge retrieval candidate library based on the expanded retrieval sentence.

[0106] Furthermore, the normalized co-occurrence frequency acquisition module 2 is configured to perform the following steps:

[0107] The retrieval behavior feature vector is weighted to obtain a feature fusion index, and the feature fusion index is mapped to a target interval to obtain a normalized co-occurrence frequency; wherein the retrieval behavior feature vector includes a retrieval frequency feature vector representing the historical retrieval frequency, a generation time feature vector representing the time required to generate a retrieval candidate library, and a retrieval feedback index vector representing the number of feedbacks on the generated retrieval candidate library.

[0108] Furthermore, the mapping function relationship building module 3 is used to perform the following steps:

[0109] Calculate the score value of each co-occurrence frequency sample in the normalized co-occurrence frequency sample and each incremental pre-training level in the incremental pre-training level sample, and output the score value sample; wherein, the score value is the information entropy mean of the quality score value generated by the candidate library and the computing resource score value required for the candidate library; according to the score value sample, obtain the incremental pre-training level optimal solution of each co-occurrence frequency sample in the normalized co-occurrence frequency sample, fit the functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution, and output the mapping function relationship.

[0110] Furthermore, the mapping function relationship building module 3 is used to perform the following steps:

[0111] The functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution is fitted by a fitting model, and the fitting model is an exponentially decreasing function model.

[0112] Furthermore, the knowledge retrieval candidate library construction module 4 is used to perform the following steps:

[0113] Input the first search statement into the statement incremental optimization model to extract the search semantic vector; perform L rounds of expansion iterations on the search semantic vector according to the matching incremental pre-training levels to obtain L rounds of related search statements, where L is the number of matching incremental pre-training levels; extract the semantic vectors of the L rounds of related search statements to perform semantic similarity calculation, merge related search statements with a semantic similarity greater than a preset value, and output the processed related search statements as extended search statements.

[0114] Furthermore, the knowledge retrieval candidate library construction module 4 is used to perform the following steps:

[0115] Calculate a round of semantic neighborhood of the retrieval semantic vector; select a round of semantically related entities from the one round of semantic neighborhood of the retrieval semantic vector to generate a round of related retrieval statements, and so on to output L rounds of related retrieval statements for L rounds of expansion iterations of the retrieval semantic vector.

[0116] Furthermore, the knowledge retrieval candidate library construction module 4 is used to perform the following steps:

[0117] Obtaining a synonymous semantic neighborhood of the retrieval semantic vector; selecting synonymous semantically related entities from the synonymous semantic neighborhood of the retrieval semantic vector to generate a synonymous related retrieval statement of the retrieval semantic vector; and updating the synonymous related retrieval statement to the L-round related retrieval statement.

[0118] Furthermore, the mapping function relationship building module 3 is used to perform the following steps:

[0119] Obtain the semantic query hit rate, consistency score annotation information and semantic relevance weight of the generated candidate library; output the candidate library generation quality score value according to the calculation results of the semantic query hit rate, consistency score annotation information and semantic relevance weight.

[0120] Furthermore, the retrieval behavior feature vector acquisition module 1 is used to perform the following steps:

[0121] Obtaining a preset period; extracting valid historical search statement behaviors within the preset period from the historical search statement behavior library, analyzing the first search statement based on the valid historical search statement behaviors, and obtaining a search behavior feature vector.

[0122] The knowledge retrieval candidate library generation system based on incremental pre-training optimization provided by an embodiment of the present invention can execute the knowledge retrieval candidate library generation method based on incremental pre-training optimization provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0123] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.

[0124] The above specific embodiments do not constitute a limitation to the scope of protection of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of this application should be included in the scope of protection of this application. In some cases, the actions or steps recorded in this application can be performed in an order different from that in the embodiments and can still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for generating a knowledge retrieval candidate library based on incremental pre-training optimization, characterized in that: The method comprises: Collecting the first search statement of the current user, calling a historical search statement behavior library to analyze the first search statement, and obtaining a search behavior feature vector; Calculating the search behavior feature vector to obtain a normalized co-occurrence frequency; Constructing a mapping function relationship between normalized co-occurrence frequency samples and incremental pre-training level samples, and using the mapping function relationship to obtain a matching incremental pre-training level corresponding to the normalized co-occurrence frequency; A pre-trained statement incremental optimization model is provided, wherein the statement incremental optimization model uses the matching incremental pre-training level to perform multi-level expansion on the first search statement, outputs the expanded search statement, and generates a knowledge retrieval candidate library based on the expanded search statement.

2. The method according to claim 1, wherein Performing weight calculation on the retrieval behavior feature vector to obtain a feature fusion index, mapping the feature fusion index to a target interval to obtain a normalized co-occurrence frequency; The search behavior feature vector includes a search frequency feature vector representing the historical search frequency, a generation time feature vector representing the time required to generate a search candidate library, and a search feedback index vector representing the number of feedbacks on the generated search candidate library.

3. The method according to claim 1, wherein Constructing a mapping relationship between normalized co-occurrence frequency samples and incremental pre-training level samples, the method includes: Calculating the score value of each co-occurrence frequency sample in the normalized co-occurrence frequency sample and each incremental pre-training level in the incremental pre-training level sample, and outputting a score value sample; The score value is the information entropy mean of the candidate library generation quality score value and the candidate library required computing resource score value; According to the score value sample, the incremental pre-training level optimal solution of each co-occurrence frequency sample in the normalized co-occurrence frequency sample is obtained, the functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution is fitted, and the mapping function relationship is output.

4. The method according to claim 3, wherein The functional relationship between each co-occurrence frequency sample in the normalized co-occurrence frequency sample and the corresponding incremental pre-training level optimal solution is fitted by a fitting model, and the fitting model is an exponentially decreasing function model.

5. The method according to claim 1, wherein The sentence increment optimization model uses the matching increment pre-training level to perform multi-level expansion on the first search sentence and outputs the expanded search sentence, and the method includes: Inputting the first search sentence into the sentence increment optimization model to extract the search semantic vector; Performing L rounds of expansion iterations on the search semantic vector according to the matching incremental pre-training levels to obtain L rounds of related search statements, where L is the number of matching incremental pre-training levels; The semantic vectors of the L rounds of related search sentences are extracted to perform semantic similarity calculation, related search sentences with a semantic similarity greater than a preset value are merged, and the processed related search sentences are output as extended search sentences.

6. The method according to claim 5, wherein Performing L rounds of expansion iterations on the retrieval semantic vector according to the matching increment pre-training level, the method comprising: Calculating a semantic neighborhood of the retrieval semantic vector; A round of semantically related entities is selected from a round of semantic neighborhood of the search semantic vector to generate a round of related search sentences, and so on, L rounds of related search sentences are output after the search semantic vector is subjected to L rounds of expansion iterations.

7. The method according to claim 5, wherein Performing L rounds of expansion iterations on the retrieval semantic vector according to the matching increment pre-training level, the method further includes: Obtaining a synonymous semantic neighborhood of the search semantic vector; Selecting synonymous semantically related entities from the synonymous semantic neighborhood of the search semantic vector to generate a synonymous related search statement of the search semantic vector; The synonymous related search statement is updated to the L-round related search statement.

8. The method according to claim 3, wherein Methods for obtaining candidate library generation quality scores include: Obtain the semantic query hit rate, consistency score annotation information, and semantic relevance weight of the generated candidate library; According to the calculation results of the semantic query hit rate, consistency score annotation information and semantic relevance weight, the candidate library is output to generate a quality score value.

9. The method according to claim 1, wherein The method of calling a historical search statement behavior library to analyze the first search statement includes: Get the preset period; Valid historical search statement behaviors within the preset period are extracted from the historical search statement behavior library, and the first search statement is analyzed based on the valid historical search statement behaviors to obtain a search behavior feature vector.

10. A knowledge retrieval candidate library generation system based on incremental pre-training optimization, characterized by: A system for implementing the method for generating a knowledge retrieval candidate library based on incremental pre-training optimization according to any one of claims 1 to 9, comprising: A retrieval behavior feature vector acquisition module is used to collect the first retrieval statement of the current user, call the historical retrieval statement behavior library to analyze the first retrieval statement, and obtain a retrieval behavior feature vector; A normalized co-occurrence frequency acquisition module, configured to calculate the retrieval behavior feature vector to obtain a normalized co-occurrence frequency; A mapping function relationship construction module is used to construct a mapping function relationship between normalized co-occurrence frequency samples and incremental pre-training level samples, and use the mapping function relationship to obtain the matching incremental pre-training level corresponding to the normalized co-occurrence frequency; A knowledge retrieval candidate library construction module is used to pre-train a statement incremental optimization model. The statement incremental optimization model uses the matching incremental pre-training level to perform multi-level expansion on the first retrieval statement, outputs the expanded retrieval statement, and generates a knowledge retrieval candidate library based on the expanded retrieval statement.

Citation Information

Patent Citations

  • Query method and device

    CN107256267A

  • Query content library construction method and device, electronic equipment and readable storage medium

    CN114706841A

  • Intelligent retrieval method and system for unstructured asset content based on large model

    CN119646243A

  • Dynamic vector knowledge base construction and retrieval method based on multi-modal large model

    CN120277223A

  • Archive knowledge base construction and retrieval method and system based on multi-modal data fusion

    CN120407703A

Cited By

  • Keyword retrieval optimization system and method based on AI intelligent analysis

    CN120849594A

  • Information retrieval method based on large model deep search and real-time semantic analysis

    CN121636674A

  • Self-adaptive retrieval strategy optimization method and system based on artificial intelligence

    CN122286003A