Bid invitation information retrieval method and device based on multi-path mixed recall mechanism

By constructing a bidding information knowledge base with a multi-channel hybrid recall mechanism and combining the ERNIE model with LoRA incremental learning, the accuracy and policy adaptability issues of the retrieval system in power material bidding are solved, and efficient power material bidding information retrieval is achieved.

CN120763307AActive Publication Date: 2025-10-10JIANGSU ELECTRIC POWER INFORMATION TECH

Patent Information

Application Number
CN202511292443.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-10
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Traditional power material bidding retrieval systems find it difficult to balance accurate parameter queries with implicit semantic associations, resulting in low review efficiency and poor policy adaptability. Existing multi-channel recall methods fail to effectively solve the problem of contextual binding between technical terms and acceptance standards, resulting in a high rate of missed detection of key contextual information.

Method used

A bidding information knowledge base is constructed, adopting a multi-channel hybrid recall mechanism, including sparse channels, dense channels and hybrid channels, combined with the dynamic reordering of the ERNIE model and LoRA incremental learning, through the combination of TF-IDF index vectors, dense semantic vectors and Sentence-BERT and ColBERT models to achieve accurate parameter matching, term semantic expansion and clause context association.

Benefits of technology

It improves the accuracy and reliability of information retrieval for power material bidding, realizes real-time adaptation to policy changes, improves review efficiency and retrieval accuracy, and ensures that the system can dynamically adapt to business changes and maintain high retrieval accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763307A_ABST
    Figure CN120763307A_ABST
Patent Text Reader

Abstract

The invention provides a bid invitation information retrieval method and device based on a multi-path mixed recall mechanism. The method comprises the following steps: constructing a bid invitation information knowledge base; receiving a query statement input by a user; the query statement is analyzed, and matched knowledge slices are retrieved in the bid invitation information knowledge base to serve as a first recall result set; inputting the query statement into a Sension-BERT model to generate a query vector, and retrieving a matched knowledge slice as a second recall result set according to the query vector; forming query conditions by each knowledge slice and each query statement, inputting the query conditions into a ColBERT model for analysis, and outputting context-associated knowledge slices as a third recall result set; and integrating the first recall result set, the second recall result set and the third recall result set to generate a mixed recall result set, inputting the query statement and the mixed recall result set into an ERNIE model, and outputting a retrieval result. The method can improve the accuracy of bid invitation information retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a bidding information retrieval method and device based on a multi-path hybrid recall mechanism. Background Art

[0002] As the scale of power material bidding expands, bidding documents are becoming multimodal and fragmented, encompassing technical parameters (such as "transformer insulation rating"), qualification clauses (such as "supplier test report requirements"), and associated specifications (such as "material acceptance standards"). Traditional search systems rely on single algorithms (such as keyword matching or semantic models), making it difficult to balance accurate parameter queries with implicit semantic associations. This results in inefficient bidding review and poor policy adaptability.

[0003] Patent document CN118733732A discloses a multi-channel recall retrieval method and device for an electric power knowledge service. The method includes searching a search vector through a retriever to obtain retrieval information; matching the retrieval information in a vector database using a preset recall strategy to obtain matching results; reordering the matching results through a hybrid search; inputting the reordered matching results into a large electric power knowledge service model to generate answer results, and inputting the query information and answer results into a security protection model for judgment. This method uses a "sparse + dense" dual-channel recall, but does not introduce fine-grained interactive modeling. It cannot solve the problem of contextual binding between technical terms and acceptance criteria in electric power material bidding, resulting in a high rate of missed detection of key contextual information and difficulty in balancing retrieval efficiency and business coverage. Summary of the Invention

[0004] The present invention provides a bidding information retrieval method and device based on a multi-path hybrid recall mechanism, which can improve the accuracy and reliability of information retrieval.

[0005] A bidding information retrieval method based on a multi-channel hybrid recall mechanism includes: Constructing a tender information knowledge base, wherein the attributes of knowledge slices in the tender information knowledge base include standard text fragments, TF-IDF index vectors, and dense semantic vectors; Receive query statements input by users; Parsing the query statement, and searching the bidding information knowledge base for matching knowledge slices with attributes of TF-IDF index vectors as a first recall result set according to the parsing result; Inputting the query statement into the Sentence-BERT model to generate a query vector, and searching the bidding information knowledge base for matching knowledge slices with attributes of dense semantic vectors as a second recall result set according to the query vector; Traversing the tender information knowledge base, respectively forming query conditions from each knowledge slice and the query statement into the ColBERT model for analysis, and outputting the context-related knowledge slices as the third recall result set; The first recall result set, the second recall result set and the third recall result set are integrated to generate a mixed recall result set, the query statement and the mixed recall result set are input into the pre-trained ERNIE model for fine rearrangement, and the retrieval results are output.

[0006] Furthermore, a bidding information knowledge base is constructed, including: Perform structured parsing of the bidding documents, extracting technical parameter fields, qualification clause fields, and associated specification fields, and standardizing professional terminology to obtain standard text fragments on technical parameters, qualification clauses, and associated specifications; Generating a TF-IDF index vector based on the standard text segment; Input the standard text fragment into the Sentence-BERT model to generate a dense semantic vector; The standard text fragment, TF-IDF index vector and dense semantic vector are stored as knowledge slices.

[0007] Furthermore, the query statement is parsed, and according to the parsing result, matching knowledge slices with attributes of TF-IDF index vectors are retrieved from the tender information knowledge base as a first recall result set, including: Identify whether there is an operator in the query statement, and if so, generate a triple containing a field name, an operator, and a value according to the query statement; if not, extract keywords from the query statement; Calculating a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm, or calculating a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm; The knowledge slice corresponding to the triple whose first matching degree is greater than the preset matching degree threshold, or the knowledge slice whose second matching degree is greater than the preset matching threshold is used as the first recall result set.

[0008] Furthermore, according to the query vector, matching knowledge slices with attributes of dense semantic vectors are retrieved from the bidding information knowledge base as a second recall result set, including: Performing dimensionality reduction processing on the query vector; Clustering the knowledge slices whose attributes are dense semantic vectors in the bidding information knowledge base to obtain multiple clustering units; Calculate the distance between the query vector after dimensionality reduction and the cluster center of each cluster unit, and select the cluster unit with a distance less than the preset distance value as the target cluster unit; Performing PQ encoding on the knowledge slices within the target clustering unit to obtain a PQ encoding vector; The first cosine similarity between the query vector after dimensionality reduction processing and each of the PQ encoding vectors is calculated, and the knowledge slices corresponding to the PQ encoding vectors whose first cosine similarity is greater than a preset value are used as the second recall result set.

[0009] Furthermore, each knowledge slice and the query statement are respectively combined into query conditions and input into the ColBERT model for analysis, and the context-related knowledge slices are output as the third recall result set, including: The ColBERT model performs word segmentation processing on the query statement in the query condition to generate a first word-gram sequence, and performs word segmentation processing on the knowledge slice to generate a second word-gram sequence; Converting the first word-gram sequence and the second word-gram sequence in the same query condition to generate a first word embedding vector sequence and a second word embedding vector sequence; Calculate the second cosine similarity between each first word embedding vector in the first word embedding vector sequence and each second word embedding vector in the second word embedding vector sequence, and take the maximum value of the second cosine similarity as the association contribution value between the first word embedding vector and the corresponding knowledge slice; Sum up all the correlation contribution values ​​calculated in the same query condition to obtain the total correlation degree of the knowledge slice corresponding to the query statement; The knowledge slices whose total relevance is greater than the preset relevance threshold are context-bound and used as the third recall result set.

[0010] Furthermore, the first recall result set, the second recall result set, and the third recall result set are integrated to generate a mixed recall result set, including: The first recall result set, the second recall result set, and the third recall result set are combined and deduplicated to generate the mixed recall result set.

[0011] Furthermore, the ERNIE model includes an input layer, a Transformer encoder, an interaction aggregation module, and an output layer; The query statement and the mixed recall result set are input into the pre-trained ERNIE model for fine rearrangement, and the retrieval results are output, including: The input layer receives the query statement and the mixed recall result set, and performs splicing, encoding, and vector conversion on the query statement and each knowledge slice in the mixed recall result set to generate a splicing vector; The Transformer encoder performs perceptual interaction on the concatenated vector based on a multi-head attention mechanism to generate a context-aware vector; The interactive aggregation module aggregates global semantic information on the context-aware vector to generate a global semantic vector; The output layer performs task classification based on the global semantic vector and outputs the matching probability between the query statement and the corresponding knowledge slice; The knowledge slices in the mixed recall result set are arranged in descending order of matching probability as the retrieval result.

[0012] Furthermore, during the pre-training of the ERNIE model, the loss function is the sum of the cross entropy loss function and the business weight constraint term; The business weight constraint item is obtained by summing the absolute values ​​of the differences between all predicted business category probabilities and the corresponding preset business category weight values, and multiplying the sum by a hyperparameter.

[0013] Furthermore, the fully connected layer in the output layer of the ERNIE model introduces LoRA parameters; The method further comprises: Detecting whether the user has input the same query statement more than a preset number of times, and if the user has input the same query statement more than the preset number of times, generating a first incremental learning sample based on the search results corresponding to the same query statement; Freeze the main parameters of the ERNIE model, input the first incremental learning sample into the ERNIE model for training, and fine-tune the LoRA parameters; or, A second incremental learning sample is constructed based on the newly added bidding document, the main parameters of the ERNIE model are frozen, the second incremental learning sample is input into the ERNIE model for training, and the LoRA parameters are fine-tuned.

[0014] A bidding information retrieval device based on a multi-channel hybrid recall mechanism, comprising: A knowledge base construction module is used to construct a tender information knowledge base, wherein the attributes of the knowledge slices in the tender information knowledge base include standard text fragments, TF-IDF index vectors and dense semantic vectors; A receiving module, used for receiving a query statement input by a user; a first retrieval module, configured to parse the query statement and, based on the parsing result, retrieve matching knowledge slices with attributes of TF-IDF index vectors from the tender information knowledge base as a first recall result set; a second retrieval module, configured to input the query statement into a Sentence-BERT model to generate a query vector, and retrieve matching knowledge slices with attributes of dense semantic vectors from a tender information knowledge base according to the query vector as a second recall result set; A third retrieval module is configured to traverse the tender information knowledge base, combine each knowledge slice and the query statement into query conditions, input them into the ColBERT model for analysis, and output context-related knowledge slices as a third recall result set; The comprehensive retrieval module is used to integrate the first recall result set, the second recall result set and the third recall result set to generate a mixed recall result set, input the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement, and output the retrieval results.

[0015] Furthermore, the knowledge base construction module constructs a bidding information knowledge base, including: Perform structured parsing of the bidding documents, extracting technical parameter fields, qualification clause fields, and associated specification fields, and standardizing professional terminology to obtain standard text fragments on technical parameters, qualification clauses, and associated specifications; Generating a TF-IDF index vector based on the standard text segment; Input the standard text fragment into the Sentence-BERT model to generate a dense semantic vector; The standard text fragment, TF-IDF index vector and dense semantic vector are stored as knowledge slices.

[0016] Furthermore, the first retrieval module parses the query statement and retrieves matching knowledge slices with attributes of TF-IDF index vectors in the tender information knowledge base according to the parsing result as a first recall result set, including: Identify whether there is an operator in the query statement, and if so, generate a triple containing a field name, an operator, and a value according to the query statement; if not, extract keywords from the query statement; Calculating a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm, or calculating a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm; The knowledge slice corresponding to the triple whose first matching degree is greater than the preset matching degree threshold, or the knowledge slice whose second matching degree is greater than the preset matching threshold is used as the first recall result set.

[0017] Further, the second retrieval module retrieves, as a second recall result set, knowledge slices that match the query vector and have dense semantic attribute vectors in the tender information knowledge base according to the query vector, including: performing dimension reduction processing on the query vector; clustering and dividing the knowledge slices with dense semantic attribute vectors in the tender information knowledge base to obtain a plurality of clustering units; calculating distances between the query vector after dimension reduction processing and clustering centers of the clustering units, and selecting a clustering unit with a distance less than a preset distance value as a target clustering unit; performing PQ encoding on the knowledge slices in the target clustering unit to obtain a PQ encoding vector; calculating first cosine similarities between the query vector after dimension reduction processing and each PQ encoding vector, and selecting a knowledge slice corresponding to a PQ encoding vector with a first cosine similarity greater than a preset value as the second recall result set.

[0018] Further, the third retrieval module inputs each knowledge slice and the query statement into the ColBERT model as query conditions for analysis, respectively, and outputs context-related knowledge slices as a third recall result set, including: The ColBERT model performs word segmentation processing on the query statement in the query condition to generate a first token sequence, and performs word segmentation processing on the knowledge slice to generate a second token sequence; transforming the first token sequence and the second token sequence in the same query condition to generate a first word embedding vector sequence and a second word embedding vector sequence; calculating second cosine similarities between each first word embedding vector in the first word embedding vector sequence and each second word embedding vector in the second word embedding vector sequence, respectively, and taking a maximum value of the second cosine similarities as an association contribution value of the first word embedding vector and the corresponding knowledge slice; summing all association contribution values obtained in the same query condition to obtain a total association degree of the query statement to the knowledge slice; contextually binding the knowledge slice with a total association degree greater than a preset association degree threshold as the third recall result set.

[0019] Further, the comprehensive retrieval module integrates the first recall result set, the second recall result set, and the third recall result set to generate a hybrid recall result set, including: performing a set operation and a deduplication process on the first recall result set, the second recall result set, and the third recall result set to generate the hybrid recall result set.

[0020] Furthermore, the ERNIE model includes an input layer, a Transformer encoder, an interaction aggregation module, and an output layer; The comprehensive retrieval module inputs the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement and outputs the retrieval results, including: The input layer receives the query statement and the mixed recall result set, and performs splicing, encoding, and vector conversion on the query statement and each knowledge slice in the mixed recall result set to generate a splicing vector; The Transformer encoder performs perceptual interaction on the concatenated vector based on a multi-head attention mechanism to generate a context-aware vector; The interactive aggregation module aggregates global semantic information on the context-aware vector to generate a global semantic vector; The output layer performs task classification based on the global semantic vector and outputs the matching probability between the query statement and the corresponding knowledge slice; The knowledge slices in the mixed recall result set are arranged in descending order of matching probability as the retrieval result.

[0021] Furthermore, during the pre-training of the ERNIE model, the loss function is the sum of the cross entropy loss function and the business weight constraint term; The business weight constraint item is obtained by summing the absolute values ​​of the differences between all predicted business category probabilities and the corresponding preset business category weight values, and multiplying the sum by a hyperparameter.

[0022] Furthermore, the fully connected layer in the output layer of the ERNIE model introduces LoRA parameters; The device further includes an incremental learning module, configured to: Detecting whether the user has input the same query statement more than a preset number of times, and if the user has input the same query statement more than the preset number of times, generating a first incremental learning sample based on the search results corresponding to the same query statement; Freeze the main parameters of the ERNIE model, input the first incremental learning sample into the ERNIE model for training, and fine-tune the LoRA parameters; or, A second incremental learning sample is constructed based on the newly added bidding document, the main parameters of the ERNIE model are frozen, the second incremental learning sample is input into the ERNIE model for training, and the LoRA parameters are fine-tuned.

[0023] The bidding information retrieval method and device based on the multi-channel hybrid recall mechanism provided by the present invention have at least the following beneficial effects: (1) By constructing a bidding information knowledge base containing standardized knowledge slices, adopting a multi-channel hybrid recall mechanism of sparse channels, dense channels, and hybrid channels, and combining the dynamic rearrangement of the ERNIE model and LoRA incremental learning, the triple requirements of accurate parameter matching, term semantic extension, and clause context association in power material bidding are met at the same time, and real-time adaptation of policy changes is achieved, thereby breaking through the bottlenecks of retrieval accuracy and dynamic business coverage, improving the efficiency of bidding review, and providing efficient and accurate knowledge retrieval support for power material bidding; (2) The practicality of the search results is improved through dynamic re-ranking optimization of the ERNIE model. Combined with the LoRA incremental learning technology, the model parameters are updated when multiple clicks of users or policy updates are detected, ensuring that the system can dynamically adapt to business changes and maintain high search accuracy, effectively improving the efficiency of power bidding review. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 The present invention provides a flowchart of an embodiment of a bidding information retrieval method based on a multi-channel mixed recall mechanism.

[0025] Figure 2 The present invention provides a flowchart of an embodiment of constructing a tender information knowledge base in a tender information retrieval method with a multi-channel hybrid recall mechanism.

[0026] Figure 3 The present invention provides a flowchart of an embodiment of sparse channel recall in a bidding information retrieval method with a multi-channel mixed recall mechanism.

[0027] Figure 4 The present invention provides a flowchart of an embodiment of dense channel recall in a bidding information retrieval method with a multi-channel mixed recall mechanism.

[0028] Figure 5 The present invention provides a flowchart of an embodiment of mixed channel recall in a bidding information retrieval method with a multi-channel mixed recall mechanism.

[0029] Figure 6 The present invention provides a flowchart of an embodiment of fine rearrangement in a bidding information retrieval method with a multi-channel mixed recall mechanism.

[0030] Figure 7 This is a structural diagram of an embodiment of a bidding information retrieval device based on a multi-channel mixed recall mechanism provided by the present invention. DETAILED DESCRIPTION

[0031] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] refer to Figure 1 In some embodiments, a bidding information retrieval method based on a multi-channel hybrid recall mechanism is provided, comprising: S1. Build a tender information knowledge base, where the attributes of knowledge slices in the tender information knowledge base include standard text segments, TF-IDF index vectors, and dense semantic vectors; S2. Receive a query statement input by the user; S3. Parsing the query statement, and searching the bidding information knowledge base for matching knowledge slices whose attributes are TF-IDF index vectors as a first recall result set according to the parsing result; S4. Input the query statement into the Sentence-BERT model to generate a query vector, and retrieve matching knowledge slices with attributes of dense semantic vectors in the bidding information knowledge base according to the query vector as a second recall result set; S5. Traverse the tender information knowledge base, combine each knowledge slice and the query statement into query conditions, input them into the ColBERT model for analysis, and output the context-related knowledge slices as the third recall result set; S6. Integrate the first recall result set, the second recall result set and the third recall result set to generate a mixed recall result set, input the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement, and output the retrieval results.

[0033] Further, refer to Figure 2 In step S1, a bidding information knowledge base is constructed, including: S11. Perform structured parsing on the bidding documents, extract technical parameter fields, qualification clause fields, and associated specification fields, and perform professional terminology standardization to obtain standard text fragments on technical parameters, qualification clauses, and associated specifications; S12, generating a TF-IDF index vector based on the standard text segment; S13, inputting the standard text segment into the Sentence-BERT model to generate a dense semantic vector; S14. Store the standard text segment, TF-IDF index vector, and dense semantic vector as knowledge slices.

[0034] Specifically, building a bidding information knowledge base is the foundation of the entire retrieval method, and its quality directly affects the accuracy and efficiency of subsequent retrieval. This step mainly performs in-depth processing of power material bidding documents to generate standardized knowledge slices.

[0035] Perform structured parsing of power material bidding documents. Power material bidding documents typically contain a wealth of information, covering technical parameters, qualification clauses, and associated specifications. The technical parameter field may include specific indicators such as the transformer's insulation grade and rated capacity; the qualification clause field includes test report requirements and qualification certificate types required of suppliers; and the associated specification field covers material acceptance standards and industry implementation specifications. Through specialized parsing programs, these key fields are extracted from the document, converting the unstructured document into structured data, laying the foundation for subsequent processing.

[0036] Furthermore, the professional terminology of electricity is standardized. There are a large number of professional terms in the field of electricity, and some terms have aliases, abbreviations, etc. For example, "transformer" may be called "transformer" in different documents, which will affect the accuracy of knowledge retrieval. Therefore, it is necessary to establish a mapping table of power equipment terms, which contains standard terms, alias terms and associated parameter groups. For example, the standard term "high-voltage circuit breaker" may have an alias such as "high-voltage switch", and the associated parameter group includes rated voltage, rated current, etc. Through the term library, the extracted terms are matched with the standard terms to achieve term standardization, ensuring that the same term with different expressions can be correctly identified in the subsequent retrieval process.

[0037] Specifically, in step S12, the standard text fragment is segmented, a vocabulary is constructed, TF (term frequency) and IFD (inverse document frequency) are calculated, TD and IDF are multiplied to obtain the TF-IDF value, and the TF-IDF value is converted into a vector form according to the order of each word in the vocabulary to obtain a TF-IDF index vector.

[0038] Furthermore, in step S13, the Sentence-BERT model embeds the standard text fragment to generate a dense semantic vector.

[0039] Furthermore, in step S14, the resulting tender information knowledge base contains the original standard text snippets, the corresponding TF-IDF index vectors, and the dense semantic vectors. The original standard text snippets are key information fragments extracted from the document, preserving the original form of the information; the TF-IDF index vectors are used in subsequent sparse channel recall, which calculates the importance of terms in the document to achieve precise matching; and the dense semantic vectors generated by Sentence-BERT are used in dense channel recall, which can better capture the semantic information of the text and achieve semantic-level retrieval.

[0040] Further, refer to Figure 3Step S3 is sparse channel recall, parsing the query statement and searching the bidding information knowledge base for matching knowledge slices with attributes of TF-IDF index vectors as the first recall result set according to the parsing result, including: S31. Identify whether an operator exists in the query statement. If an operator exists, generate a triple containing a field name, an operator, and a value according to the query statement; if an operator does not exist, extract keywords from the query statement. S32. Calculate, based on the BM25 algorithm, a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base, or calculate, based on the BM25 algorithm, a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base; S33. Take the knowledge slice corresponding to the triple whose first matching degree is greater than the preset matching degree threshold, or the knowledge slice whose second matching degree is greater than the preset matching threshold as the first recall result set.

[0041] Specifically, in step S31, the query statement entered by the user may contain an operator, for example, "Transformer insulation grade ≥ Class F," or may not contain an operator, for example, "Transformer insulation grade." If an operator is present, a triple containing a field name, an operator, and a value is generated based on the query statement. The triple contains the field name, operator, and value; for example, "Transformer insulation grade ≥ Class F." If an operator is not present, the keyword is directly extracted.

[0042] Furthermore, in step S32, a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base is calculated based on the BM25 algorithm, or a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base is calculated based on the BM25 algorithm, and the calculation formula is: ; (1) Among them, score(D,Q) is the matching degree between knowledge slice D and query statement Q, n represents the number of word items in the query statement, which can be the number of elements in the triple or the number of keywords, IDF(q i ) represents the i-th participle item q i The inverse document frequency of knowledge slice D is represented by f(qi,D), which indicates the word frequency factor. The higher the word frequency is, the more important the word is in the document. k1 indicates the adjustment parameter, b indicates the length penalty parameter, |D| indicates the total number of word segments in knowledge slice D, and avqdl indicates the average number of word segments in all knowledge slices.

[0043] The inverse document frequency is calculated using the following formula: ; (2) wherein N represents the total number of knowledge slices, n(q i ) represents the number of knowledge slices containing the segmented term q i .

[0044] Inverse document frequency is used to measure the general importance of a term, the less frequently a term appears in documents, the higher its inverse document frequency, and the greater its importance.

[0045] Length normalization term is used to balance the influence of documents of different lengths on the search results, to avoid long documents from obtaining higher scores due to containing more terms.

[0046] Set the adjustment parameters control the saturation of term frequency, when the term frequency reaches a certain degree, its contribution to the score slows down; control the length penalty strength, the greater the value, the more obvious the penalty to long documents. By setting these parameters, the BM25 algorithm can more accurately calculate the matching degree of documents and queries, so as to retrieve knowledge slices that meet the precise parameter conditions.

[0047] Further, in step S33, for the query statement containing the operator, the knowledge slice corresponding to the triple with the first matching degree greater than the preset matching degree threshold is taken as the first recall result set. For the query statement not containing the operator, the knowledge slice with the second matching degree greater than the preset matching threshold is taken as the first recall result set.

[0048] Further, with reference to Figure 4 , step S4 is dense channel recall, and the knowledge slice matching the query vector and having a dense semantic attribute in the tender information knowledge base is retrieved as the second recall result set, including: S41, performing dimension reduction processing on the query vector; S42, clustering and dividing the knowledge slices with a dense semantic attribute in the tender information knowledge base to obtain a plurality of clustering units; S43, calculating the distance between the query vector after dimension reduction processing and the clustering centers of each clustering unit, and selecting a clustering unit with a distance less than a preset distance value as a target clustering unit; S44, performing PQ encoding on the knowledge slices in the target clustering unit to obtain a PQ encoding vector; S45, calculating the first cosine similarity between the query vector after dimension reduction processing and each PQ encoding vector, and taking the knowledge slice corresponding to the PQ encoding vector with a first cosine similarity greater than a preset value as the second recall result set.

[0049] Specifically, the query sentence is input into the Sentence-BERT model to generate a query vector. Sentence-BERT is a sentence embedding model based on BERT, which can convert sentences into dense semantic vectors and better capture the semantic information of the sentence.

[0050] Furthermore, in step S41, the query vector undergoes dimensionality reduction. Specifically, PCA is used to reduce the 768-dimensional dense vector generated by Sentence-BERT to 64 dimensions. High-dimensional vectors increase computational and storage costs, and PCA dimensionality reduction can reduce dimensionality while retaining key information, improving retrieval efficiency.

[0051] Furthermore, in step S42, the IVF algorithm is used to cluster the knowledge slices with attributes of dense semantic vectors in the bidding information knowledge base, and group similar vectors into the same unit, thereby obtaining multiple cluster units. Subsequent retrieval only requires searching in some related units, reducing the amount of calculation.

[0052] Furthermore, in step S43, the distance between the query vector after the dimension reduction process and the cluster center of each cluster unit is calculated, and the cluster unit with a distance less than a preset distance value is selected as the target cluster unit.

[0053] Furthermore, in step S44 and step S45, PQ (Product Quantization) encoding is used to quantize each vector into an 8-byte index code, and the first cosine similarity between the query vector after dimensionality reduction processing and each of the PQ encoding vectors is calculated, and the knowledge slices corresponding to the PQ encoding vectors whose first cosine similarity is greater than a preset value are taken as the second recall result set.

[0054] Further, refer to Figure 5 Step S5 is a hybrid channel recall, where each knowledge slice and the query statement are combined into query conditions and input into the ColBERT model for analysis, and the context-related knowledge slices are output as the third recall result set, including: S51, the ColBERT model performs word segmentation processing on the query statement in the query condition to generate a first word-gram sequence, and performs word segmentation processing on the knowledge slice to generate a second word-gram sequence; S52: Convert the first word-gram sequence and the second word-gram sequence based on the same query condition to generate a first word embedding vector sequence and a second word embedding vector sequence; S53, respectively calculating the second cosine similarity between each first word embedding vector in the first word embedding vector sequence and each second word embedding vector in the second word embedding vector sequence, and taking the maximum value of the second cosine similarity as the association contribution value between the first word embedding vector and the corresponding knowledge slice; S54. Sum all the calculated correlation contribution values ​​for the same query condition to obtain the total correlation degree of the knowledge slice corresponding to the query statement; S55. Context-bind the knowledge slices whose total relevance is greater than a preset relevance threshold value as the third recall result set.

[0055] Specifically, all knowledge slices in the tender information knowledge base are traversed, and each knowledge slice and the query statement are respectively combined into query conditions. Assuming that there are M knowledge slices in the tender information knowledge base, M query conditions are generated.

[0056] In step S51, the ColBERT model performs word segmentation processing on the query sentence Q in the query condition to generate a first word sequence: {q1,q2,q3...q n}.

[0057] Perform word segmentation on the knowledge slice D to generate the second word sequence: {d1, d2, d3...d m}.

[0058] Furthermore, in step S52, the ColBERT model converts the first word sequence and the second word sequence in the same query condition to generate a first word embedding vector sequence and the second word embedding vector sequence .

[0059] In step S53, the cosine similarity is calculated using the following formula: ; (3) Among them, E qi Represents the first word embedding vector sequence The i-th word embedding vector, E dj Represents the second word embedding vector sequence The j-th word embedding vector in .

[0060] The calculated cosine similarity is used as the correlation contribution value of the word embedding vector in the corresponding query condition.

[0061] Furthermore, in steps S54 and S55, all the calculated correlation contribution values ​​for the same query condition are summed to obtain the total correlation of the knowledge slices corresponding to the query statement in the query condition. The knowledge slices corresponding to the query conditions whose total correlation is greater than the preset correlation threshold are context-bound and serve as the third recall result set.

[0062] Furthermore, in step S6, the first recall result set, the second recall result set, and the third recall result set are integrated to generate a mixed recall result set, including: The first recall result set, the second recall result set, and the third recall result set are combined and deduplicated to generate the mixed recall result set.

[0063] Through the above processing, the retrieval results of the three channels can be integrated to avoid missing important information, while removing duplicates and reducing the redundancy of subsequent processing.

[0064] Furthermore, the ERNIE model includes an input layer, a Transformer encoder, an interaction aggregation module, and an output layer; refer to Figure 6 , input the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement, and output the retrieval results, including: S61, the input layer receives the query statement and the mixed recall result set, and performs splicing, encoding, and vector conversion on the query statement and each knowledge slice in the mixed recall result set to generate a splicing vector; S62, the Transformer encoder performs perceptual interaction on the splicing vector based on a multi-head attention mechanism to generate a context-aware vector; S63, the interaction aggregation module aggregates global semantic information on the context perception vector to generate a global semantic vector; S64, the output layer performs task classification based on the global semantic vector and outputs the matching probability between the query statement and the corresponding knowledge slice; S65. Arrange the knowledge slices in the mixed recall result set in descending order of matching probability as the retrieval result.

[0065] Specifically, the ERNIE model is a cross-encoder based on ERNIE, which can encode query statements and knowledge slices and directly calculate the matching degree between the two, with good ranking effect.

[0066] After the above processing, knowledge slices are output to the power bidding system in descending order of matching probability. Matching probability is the result of comprehensive consideration of multi-way mixed recall and dynamic reordering optimization. Knowledge slices with higher matching probabilities are more closely matched to user queries and are therefore more valuable to users.

[0067] When deploying the power bidding platform, the user query input interface is coupled to the State Grid ECP2.0 bidding system, making it convenient for users to perform query operations in the familiar bidding system; the knowledge slice output format follows the CIM model of the IEC61970 standard, ensuring that the output knowledge can be compatible and interactive with other systems in the power industry, thereby improving the integration and practicality of the entire system.

[0068] Furthermore, during the pre-training of the ERNIE model, the loss function is the sum of the cross entropy loss function and the business weight constraint term; The business weight constraint item is obtained by summing the absolute values ​​of the differences between all predicted business category probabilities and the corresponding preset business category weight values, and multiplying the sum by a hyperparameter.

[0069] Business category weights are pre-set according to power business rules. Business weight constraints follow the rule of qualification certificate weight > quotation weight > delivery date weight. This is because in power material bidding, supplier qualification certificates are the basis for ensuring material quality and supply capacity and therefore have the highest priority. Quotations directly impact bidding costs and are second highest priority. Delivery date, which affects the timeliness of material supply, has a relatively lower priority.

[0070] In ERNIE model training, the loss function introduces a business weight constraint term. The loss function is as follows: ; (4) in, represents the loss function of the ERNIE model, represents the cross entropy loss function, λ represents the hyperparameter, is the weight constraint term, w i is the probability of the i-th predicted business category, is the weight value of the i-th preset service category, and s is the total number of service categories.

[0071] For example, business categories include qualification certificates, quotations, and delivery dates, and business weight constraints are set. ,in, is the weight value of the business category of the qualification certificate, is the business category weight value, It is the weight value of the delivery period business category.

[0072] Hyperparameters Used to adjust the strength of business constraints, The larger the value, the greater the impact of business constraints on the loss function, and the more the model tends to follow the preset weight rules; Through the ERNIE model and the loss function, the mixed recall result set is sorted by weight priority, which improves the practicality of the retrieval results.

[0073] Furthermore, the fully connected layer in the output layer of the ERNIE model introduces LoRA parameters; The method further comprises: Detecting whether the user has input the same query statement more than a preset number of times, and if the user has input the same query statement more than the preset number of times, generating a first incremental learning sample based on the search results corresponding to the same query statement; Freeze the main parameters of the ERNIE model, input the first incremental learning sample into the ERNIE model for training, and fine-tune the LoRA parameters; or, A second incremental learning sample is constructed based on the newly added bidding document, the main parameters of the ERNIE model are frozen, the second incremental learning sample is input into the ERNIE model for training, and the LoRA parameters are fine-tuned.

[0074] Specifically, incremental learning adaptation is used to enable the system to adapt to changes in user behavior and policy updates, maintaining the stability and timeliness of retrieval performance.

[0075] Incremental learning is activated when users are detected searching for the same behavioral data or policy update instructions. Triggering conditions for the incremental learning module include users entering the same query statement more than a preset number of times and the addition of new tender documents. Multiple entries of the same query statement indicate that it may be more in line with user needs, and the system needs to learn this preference. Adding new tender documents indicates that business rules may have changed, and the system needs to adapt to new policy requirements in a timely manner.

[0076] The incremental learning module uses LoRA technology with RoPE position encoding; The rank decomposition matrix is ​​injected into the fully connected layer of the ERNI model's output layer, and its parameters are fine-tuned based on incremental learning samples. By freezing most of the pre-trained model's parameters and training only the low-rank decomposition matrix, LoRA technology reduces computational effort while enabling rapid model adaptation.

[0077] The LoRA bypass matrix is ​​constructed based on the formula , where W is the original 1-weight matrix, W' is the constructed rank decomposition matrix, and A and B are low-rank adaptation matrices. The operation process is: Decomposition of the original weight matrix , where d = 1024 is the input dimension and k = 4096 is the output dimension. The original weight matrix is ​​the parameter of the pre-trained model and contains a lot of knowledge.

[0078] Constructing a low-rank adaptation matrix and , r = 128 and r < 0.05k. The rank r of the low-rank matrix is ​​much smaller than the dimension of the original weight matrix. By training these two low-rank matrices to capture new knowledge, we can not only update the model but also reduce the number of parameters and computational cost.

[0079] refer to Figure 7 In some embodiments, a bidding information retrieval device based on a multi-channel hybrid recall mechanism is provided, comprising: The knowledge base construction module 201 is used to construct a tender information knowledge base, wherein the attributes of the knowledge slices in the tender information knowledge base include standard text segments, TF-IDF index vectors and dense semantic vectors; Receiving module 202, for receiving a query statement input by a user; A first retrieval module 203 is configured to parse the query statement and retrieve matching knowledge slices with attributes of TF-IDF index vectors from the tender information knowledge base according to the parsing result as a first recall result set; A second retrieval module 204 is configured to input the query statement into the Sentence-BERT model to generate a query vector, and retrieve matching knowledge slices with attributes of dense semantic vectors from the tender information knowledge base according to the query vector as a second recall result set; The third retrieval module 205 is used to traverse the tender information knowledge base, respectively form query conditions for each knowledge slice and the query statement, input them into the ColBERT model for analysis, and output context-related knowledge slices as a third recall result set; The comprehensive retrieval module 206 is used to integrate the first recall result set, the second recall result set and the third recall result set to generate a mixed recall result set, input the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement, and output the retrieval results.

[0080] Furthermore, the knowledge base construction module 201 constructs a bidding information knowledge base, including: Perform structured parsing of the bidding documents, extracting technical parameter fields, qualification clause fields, and associated specification fields, and standardizing professional terminology to obtain standard text fragments on technical parameters, qualification clauses, and associated specifications; Generating a TF-IDF index vector based on the standard text segment; Input the standard text fragment into the Sentence-BERT model to generate a dense semantic vector; The standard text fragment, TF-IDF index vector and dense semantic vector are stored as knowledge slices.

[0081] Furthermore, the first retrieval module 203 parses the query statement and searches the bidding information knowledge base for matching knowledge slices whose attributes are TF-IDF index vectors as a first recall result set according to the parsing result, including: Identify whether there is an operator in the query statement, and if so, generate a triple containing a field name, an operator, and a value according to the query statement; if not, extract keywords from the query statement; Calculating a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm, or calculating a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm; The knowledge slice corresponding to the triple whose first matching degree is greater than the preset matching degree threshold, or the knowledge slice whose second matching degree is greater than the preset matching threshold is used as the first recall result set.

[0082] Furthermore, the second retrieval module 204 retrieves matching knowledge slices with attributes of dense semantic vectors in the bidding information knowledge base according to the query vector as a second recall result set, including: Performing dimensionality reduction processing on the query vector; Clustering the knowledge slices whose attributes are dense semantic vectors in the bidding information knowledge base to obtain multiple clustering units; Calculate the distance between the query vector after dimensionality reduction and the cluster center of each cluster unit, and select the cluster unit with a distance less than the preset distance value as the target cluster unit; Performing PQ encoding on the knowledge slices within the target cluster unit to obtain a PQ encoding vector; The first cosine similarity between the query vector after dimensionality reduction processing and each of the PQ encoding vectors is calculated, and the knowledge slices corresponding to the PQ encoding vectors whose first cosine similarity is greater than a preset value are used as the second recall result set.

[0083] Furthermore, the third retrieval module 205 combines each knowledge slice and the query statement into query conditions and inputs them into the ColBERT model for analysis, and outputs the context-related knowledge slices as the third recall result set, including: The ColBERT model performs word segmentation processing on the query statement in the query condition to generate a first word-gram sequence, and performs word segmentation processing on the knowledge slice to generate a second word-gram sequence; Converting the first word-gram sequence and the second word-gram sequence in the same query condition to generate a first word embedding vector sequence and a second word embedding vector sequence; Calculate the second cosine similarity between each first word embedding vector in the first word embedding vector sequence and each second word embedding vector in the second word embedding vector sequence, and take the maximum value of the second cosine similarity as the association contribution value between the first word embedding vector and the corresponding knowledge slice; Sum up all the correlation contribution values ​​calculated in the same query condition to obtain the total correlation degree of the knowledge slice corresponding to the query statement; The knowledge slices whose total relevance is greater than the preset relevance threshold are context-bound and used as the third recall result set.

[0084] Furthermore, the comprehensive search module 206 integrates the first recall result set, the second recall result set, and the third recall result set to generate a mixed recall result set, including: The first recall result set, the second recall result set, and the third recall result set are combined and deduplicated to generate the mixed recall result set.

[0085] Furthermore, the ERNIE model includes an input layer, a Transformer encoder, an interaction aggregation module, and an output layer; The comprehensive retrieval module inputs the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement and outputs the retrieval results, including: The input layer receives the query statement and the mixed recall result set, and performs splicing, encoding, and vector conversion on the query statement and each knowledge slice in the mixed recall result set to generate a splicing vector; The Transformer encoder performs perceptual interaction on the concatenated vector based on a multi-head attention mechanism to generate a context-aware vector; The interactive aggregation module aggregates global semantic information on the context-aware vector to generate a global semantic vector; The output layer performs task classification based on the global semantic vector and outputs the matching probability between the query statement and the corresponding knowledge slice; The knowledge slices in the mixed recall result set are arranged in descending order of matching probability as the retrieval result.

[0086] Furthermore, during the pre-training of the ERNIE model, the loss function is the sum of the cross entropy loss function and the business weight constraint term; The business weight constraint item is obtained by summing the absolute values ​​of the differences between all predicted business category probabilities and the corresponding preset business category weight values, and multiplying the sum by a hyperparameter.

[0087] Furthermore, the fully connected layer in the output layer of the ERNIE model introduces LoRA parameters; The device further includes an incremental learning module, configured to: Detecting whether the user has input the same query statement more than a preset number of times, and if the user has input the same query statement more than the preset number of times, generating a first incremental learning sample based on the search results corresponding to the same query statement; Freeze the main parameters of the ERNIE model, input the first incremental learning sample into the ERNIE model for training, and fine-tune the LoRA parameters; or, A second incremental learning sample is constructed based on the newly added bidding document, the main parameters of the ERNIE model are frozen, the second incremental learning sample is input into the ERNIE model for training, and the LoRA parameters are fine-tuned.

[0088] The bidding information retrieval method and device based on the multi-channel hybrid recall mechanism provided in the above embodiment have at least the following beneficial effects: (1) By constructing a bidding information knowledge base containing standardized knowledge slices, adopting a multi-channel hybrid recall mechanism of sparse channels, dense channels, and hybrid channels, and combining the dynamic rearrangement of the ERNIE model and LoRA incremental learning, the triple requirements of accurate parameter matching, term semantic extension, and clause context association in power material bidding are met at the same time, and real-time adaptation of policy changes is achieved, thereby breaking through the bottlenecks of retrieval accuracy and dynamic business coverage, improving the efficiency of bidding review, and providing efficient and accurate knowledge retrieval support for power material bidding; (2) The practicality of the search results is improved through dynamic re-ranking optimization of the ERNIE model. Combined with the LoRA incremental learning technology, the model parameters are updated when multiple clicks of users or policy updates are detected, ensuring that the system can dynamically adapt to business changes and maintain high search accuracy, effectively improving the efficiency of power bidding review.

[0089] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A bidding information retrieval method based on a multi-channel hybrid recall mechanism, characterized in that: include: Constructing a tender information knowledge base, wherein the attributes of knowledge slices in the tender information knowledge base include standard text fragments, TF-IDF index vectors, and dense semantic vectors; Receive query statements input by users; Parsing the query statement, and searching the bidding information knowledge base for matching knowledge slices with attributes of TF-IDF index vectors as a first recall result set according to the parsing result; Inputting the query statement into the Sentence-BERT model to generate a query vector, and searching the bidding information knowledge base for matching knowledge slices with attributes of dense semantic vectors as a second recall result set according to the query vector; Traversing the tender information knowledge base, respectively forming query conditions from each knowledge slice and the query statement into the ColBERT model for analysis, and outputting the context-related knowledge slices as the third recall result set; The first recall result set, the second recall result set and the third recall result set are integrated to generate a mixed recall result set, the query statement and the mixed recall result set are input into the pre-trained ERNIE model for fine rearrangement, and the retrieval results are output.

2. The method according to claim 1, characterized in that Build a bidding information knowledge base, including: Perform structured parsing of the bidding documents, extracting technical parameter fields, qualification clause fields, and associated specification fields, and standardizing professional terminology to obtain standard text fragments on technical parameters, qualification clauses, and associated specifications; Generating a TF-IDF index vector based on the standard text segment; Input the standard text fragment into the Sentence-BERT model to generate a dense semantic vector; The standard text fragment, TF-IDF index vector and dense semantic vector are stored as knowledge slices.

3. The method according to claim 1, characterized in that Parse the query statement, and retrieve matching knowledge slices with attributes of TF-IDF index vectors from the tender information knowledge base according to the parsing result as a first recall result set, including: Identify whether there is an operator in the query statement, and if so, generate a triple containing a field name, an operator, and a value according to the query statement; if not, extract keywords from the query statement; Calculating a first matching degree between the triple and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm, or calculating a second matching degree between the keyword and each knowledge slice whose attribute is a TF-IDF index vector in the tender information knowledge base based on the BM25 algorithm; The knowledge slice corresponding to the triple whose first matching degree is greater than the preset matching degree threshold, or the knowledge slice whose second matching degree is greater than the preset matching threshold is used as the first recall result set.

4. The method according to claim 1, wherein According to the query vector, matching knowledge slices with attributes of dense semantic vectors are retrieved from the bidding information knowledge base as a second recall result set, including: Performing dimensionality reduction processing on the query vector; Clustering the knowledge slices whose attributes are dense semantic vectors in the bidding information knowledge base to obtain multiple clustering units; Calculate the distance between the query vector after dimensionality reduction and the cluster center of each cluster unit, and select the cluster unit with a distance less than the preset distance value as the target cluster unit; Performing PQ encoding on the knowledge slices within the target cluster unit to obtain a PQ encoding vector; The first cosine similarity between the query vector after dimensionality reduction processing and each of the PQ encoding vectors is calculated, and the knowledge slices corresponding to the PQ encoding vectors whose first cosine similarity is greater than a preset value are used as the second recall result set.

5. The method according to claim 1, wherein Each knowledge slice and the query statement are respectively composed of query conditions and input into the ColBERT model for analysis, and the context-related knowledge slices are output as the third recall result set, including: The ColBERT model performs word segmentation processing on the query statement in the query condition to generate a first word-gram sequence, and performs word segmentation processing on the knowledge slice to generate a second word-gram sequence; Converting the first word-gram sequence and the second word-gram sequence in the same query condition to generate a first word embedding vector sequence and a second word embedding vector sequence; Calculate the second cosine similarity between each first word embedding vector in the first word embedding vector sequence and each second word embedding vector in the second word embedding vector sequence, and take the maximum value of the second cosine similarity as the association contribution value between the first word embedding vector and the corresponding knowledge slice; Sum up all the correlation contribution values ​​calculated in the same query condition to obtain the total correlation degree of the knowledge slice corresponding to the query statement; The knowledge slices whose total relevance is greater than the preset relevance threshold are context-bound and used as the third recall result set.

6. The method according to claim 1, characterized in that Integrating the first recall result set, the second recall result set, and the third recall result set to generate a mixed recall result set, including: The first recall result set, the second recall result set, and the third recall result set are combined and deduplicated to generate the mixed recall result set.

7. The method according to claim 1, characterized in that The ERNIE model includes an input layer, a Transformer encoder, an interaction aggregation module, and an output layer; The query statement and the mixed recall result set are input into the pre-trained ERNIE model for fine rearrangement, and the retrieval results are output, including: The input layer receives the query statement and the mixed recall result set, and performs splicing, encoding, and vector conversion on the query statement and each knowledge slice in the mixed recall result set to generate a splicing vector; The Transformer encoder performs perceptual interaction on the concatenated vector based on a multi-head attention mechanism to generate a context-aware vector; The interactive aggregation module aggregates global semantic information on the context-aware vector to generate a global semantic vector; The output layer performs task classification based on the global semantic vector and outputs the matching probability between the query statement and the corresponding knowledge slice; The knowledge slices in the mixed recall result set are arranged in descending order of matching probability as the retrieval result.

8. The method according to claim 1 or 7, characterized in that During the pre-training of the ERNIE model, the loss function is the sum of the cross entropy loss function and the business weight constraint term; The business weight constraint item is obtained by summing the absolute values ​​of the differences between all predicted business category probabilities and the corresponding preset business category weight values, and multiplying the sum by a hyperparameter.

9. The method according to claim 7, characterized in that The fully connected layer in the output layer of the ERNIE model introduces LoRA parameters; The method further comprises: Detecting whether the user has input the same query statement more than a preset number of times, and if the user has input the same query statement more than the preset number of times, generating a first incremental learning sample based on the search results corresponding to the same query statement; Freeze the main parameters of the ERNIE model, input the first incremental learning sample into the ERNIE model for training, and fine-tune the LoRA parameters; or, A second incremental learning sample is constructed based on the newly added bidding document, the main parameters of the ERNIE model are frozen, the second incremental learning sample is input into the ERNIE model for training, and the LoRA parameters are fine-tuned.

10. A bidding information retrieval device based on a multi-channel hybrid recall mechanism, characterized in that: include: A knowledge base construction module is used to construct a tender information knowledge base, wherein the attributes of the knowledge slices in the tender information knowledge base include standard text fragments, TF-IDF index vectors and dense semantic vectors; A receiving module, used for receiving a query statement input by a user; a first retrieval module, configured to parse the query statement and, based on the parsing result, retrieve matching knowledge slices with attributes of TF-IDF index vectors from the tender information knowledge base as a first recall result set; a second retrieval module, configured to input the query statement into a Sentence-BERT model to generate a query vector, and retrieve matching knowledge slices with attributes of dense semantic vectors from a tender information knowledge base according to the query vector as a second recall result set; A third retrieval module is configured to traverse the tender information knowledge base, combine each knowledge slice and the query statement into query conditions, input the query conditions into the ColBERT model for analysis, and output context-related knowledge slices as a third recall result set; The comprehensive retrieval module is used to integrate the first recall result set, the second recall result set and the third recall result set to generate a mixed recall result set, input the query statement and the mixed recall result set into the pre-trained ERNIE model for fine rearrangement, and output the retrieval results.

Citation Information

Patent Citations

  • Method and system for retrieving context of objective questions of law test

    CN116805148A

  • Patent retrieval system combining patent images and text semantics

    CN118113810A

  • Table query method and device

    CN118277402A

  • Power knowledge service method and device for multi-path recall retrieval, equipment and medium

    CN118733732A

  • Text hotspot clustering method based on large model

    CN119474387A

Cited By

  • Generation method for improving quality of content generated by RAG technology based on semantic matching

    CN121599129A