Bidding document generation method based on retrieval enhancement generation and large language model
Through the method based on search-enhanced generation and large language model, the term recognition and compliance problems in bid generation are solved, and the precise extraction and compliance detection of bid content are realized, which improves the probability of winning bids and the adaptability of user preferences.
Patent Information
- Application Number
- CN202510780207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional bid creation methods rely on manual writing, making it difficult to accurately identify domain-specific terms and legal semantics. The existing search enhancement technology does not integrate a user feedback-driven weight optimization mechanism, resulting in an imbalance in the generated content and user preference bias and the integration of historical bid fragments with the latest technical specifications.
By using a method based on search enhancement generation and large language models, bidding requirements documents are obtained, technical parameters and legal terms are analyzed, and the industry knowledge base is used to search and match historical bid fragments and technical specifications, and combining dynamic weight allocation algorithms and legal semantic knowledge graphs to generate compliance verification and optimize bid content.
It realizes accurate extraction of bid content and legal compliance testing, reduces manual writing time, increases the probability of winning bids, and adapts to the update of industry norms and legal terms through a dynamic optimization mechanism driven by user feedback.
Smart Images

Figure CN120297242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a bid generation method based on retrieval-enhanced generation and large language models. Background Art
[0002] Traditional bid generation methods mainly rely on manual writing and template filling, and there are the following technical bottlenecks: Tender documents mostly contain unstructured text (such as nested tables, multi-page paragraphs). Although existing natural language processing technologies (such as CN119782470A "A method for generating a large language model based on retrieval enhancement") propose a retrieval-enhanced generation framework, they do not design a legal semantic knowledge graph and a multi-level conflict detection module for the tender scenario, and it is difficult to accurately identify domain-specific terms.
[0003] Existing retrieval enhancement technologies (such as CN118779425A "A retrieval-enhanced generation method for large language models based on hierarchical information expansion") have achieved chapter-by-chapter retrieval of documents, but have not integrated a user feedback-driven weight optimization mechanism, resulting in a significant deviation between the generated content and user preferences, and an imbalance in the fusion ratio of historical bid fragments and the latest technical specifications. Summary of the Invention
[0004] The present invention provides a bid generation method based on retrieval-enhanced generation and large language models, and its main purpose is to solve the problems of inaccurate extraction of technical parameters in bids, low efficiency of legal conflict detection, and unoptimized scoring rules.
[0005] To achieve the above object, a bid generation method based on retrieval-enhanced generation and large language models provided by the present invention includes: 1. A bid generation method based on retrieval-enhanced generation and large language models, characterized in that the method includes: S1. Obtain the tender requirement document of the target bid, and parse the technical parameters, legal terms, and scoring weights in the tender requirement document; S2. Based on the technical parameters, retrieve and match historical bid fragments and associated technical specifications from a preset industry knowledge base to generate a retrieval-enhanced data set; S3. Input the retrieval-enhanced data set into a pre-trained large language model, and fuse the historical bid fragments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial bid draft; S4. Perform compliance verification on the initial bid draft according to the legal terms, and identify and mark the conflicting terms; S5. Based on the scoring weights, perform structural optimization on the initial bid draft to generate a target bid that meets the tender scoring rules.
[0006] Optionally, parsing the technical parameters in the tender requirement document includes: Extracting the technical requirement paragraphs in the tender requirement document using natural language processing technology; Annotating the equipment models, performance indicators, and acceptance criteria in the technical requirement paragraphs through named entity recognition.
[0007] Optionally, retrieving matching historical tender fragments and associated technical specifications from a preset industry knowledge base based on the technical parameters includes: Generating multi-dimensional retrieval keywords according to the technical parameters; Based on the multi-dimensional retrieval keywords, using an inverted index to screen out candidate fragments from a preset industry knowledge base; Sorting the candidate fragments based on cosine similarity, and determining the historical tender fragments and associated technical specifications most relevant to the technical parameters according to the sorted candidate fragments.
[0008] Optionally, fusing the historical tender fragments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial tender draft includes: Calculating the semantic similarity between the historical tender fragments and the current technical parameters to generate a first weight coefficient; Extracting the timeliness indicators of the associated technical specifications, and generating a second weight coefficient based on the release time and revision version in the timeliness indicators, where the timeliness indicators include the national standard version number of the technical specifications and the validity period of industry certification; Obtaining the feedback score of the user on the historical tender fragments to generate a third weight coefficient; Normalizing the first weight coefficient, the second weight coefficient, and the third weight coefficient; Dynamically adjusting the fusion ratio of the historical tender fragments and the associated technical specifications according to the normalized weight coefficients, and generating an initial tender draft that meets the current technical requirements based on the fusion ratio.
[0009] Optionally, the calculation of the semantic similarity uses a sentence vector matching model based on BERT.
[0010] Optionally, performing compliance verification on the initial tender draft according to the legal provisions, and identifying and annotating conflicting clauses includes: Extracting the key obligations and restrictive conditions in the legal provisions, and constructing a legal semantic knowledge graph; Mapping the initial tender draft to the legal semantic knowledge graph paragraph by paragraph according to each legal provision, and detecting the coverage integrity of the initial tender draft for the legal provisions; Identifying the expressions in the initial tender draft that conflict with the legal provisions; Perform multi-level severity annotation on the identified conflicting clauses. Among them, the conflict levels of the conflicting clauses are divided into mandatory conflicts and recommended conflicts according to severity.
[0011] Optionally, after identifying and annotating the conflicting clauses, it further includes: Generate differentiated correction suggestions according to the conflict levels and embed them in the revision log of the initial tender draft; Retain the revision traces and legal bases of the conflicting clauses in the target tender. Among them, the revision traces include: the original text of the conflicting clause, the revised version, and the corresponding legal article numbers.
[0012] Optionally, the semantic matching model of the cross-document alignment module including the Transformer architecture is used to identify the expressions in the initial tender draft that conflict with the legal clauses.
[0013] Optionally, the structural optimization of the initial tender draft based on the scoring weights to generate a target tender that complies with the tender evaluation rules includes: Sort the priorities of each chapter of the initial tender draft according to the scoring weights; Based on the sorting results of the priority sorting, insert a strengthened description of the technical solution into the chapters where the weight ratio exceeds the preset ratio threshold; Generate a table of contents index for the target tender according to the key points of the tender evaluation rules and highlight the key contents in the target tender.
[0014] Optionally, after generating the initial tender draft, it further includes: Establish a user feedback-driven dynamic optimization mechanism to collect the editing behavior data of the tender fragments in the initial tender draft in real time. The editing behavior data includes the text modification track and the paragraph adoption rate; Convert the editing behavior data into a fine-tuning parameter set of the large language model through an online learning algorithm and dynamically inject the fine-tuning parameter set in the subsequent tender generation process to form a weight allocation strategy adaptive to user preferences; Construct a two-way association index between the user feedback data and the historical tender fragments. When it is detected that the user repeatedly corrects the technical parameters, automatically trigger the version update verification of the associated technical specifications in the industry knowledge base.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the collaboration of retrieval-augmented generation and the large language model, the intelligent integration of historical tender fragments and associated technical specifications is realized, significantly reducing the time cost of manually preparing tenders, and at the same time avoiding omissions or errors in technical parameters caused by lack of experience; 2. Combine the semantic knowledge graph of legal clauses with the structured analysis of scoring weights to ensure that the tender content complies with regulatory requirements, and focus on optimizing according to the tender evaluation rules to increase the probability of winning the bid; 3. Through a dynamic optimization mechanism driven by user feedback, enable the model to adapt to user preferences and industry standard updates in real time, and avoid the problem of outdated tender content caused by the iteration of technical standards or the revision of legal clauses. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flow chart of a tender generation method based on retrieval-enhanced generation and large language model provided by an embodiment of the present invention.
[0017] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] An embodiment of the present application provides a tender generation method based on retrieval-enhanced generation and large language model. The execution subject of the tender generation method based on retrieval-enhanced generation and large language model includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the tender generation method based on retrieval-enhanced generation and large language model can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0020] Refer to Figure 1 As shown, it is a schematic flow chart of a tender generation method based on retrieval-enhanced generation and large language model provided by an embodiment of the present invention. In this embodiment, the tender generation method based on retrieval-enhanced generation and large language model includes: S1. Obtain the tender demand document of the target tender, and analyze the technical parameters, legal clauses and scoring weights in the tender demand document.
[0021] In the embodiment of the present invention, the analysis of the technical parameters in the tender demand document includes: Extract the technical requirement paragraphs in the tender requirement document using natural language processing techniques; Annotate the equipment models, performance indicators, and acceptance criteria in the technical requirement paragraphs through named entity recognition.
[0022] Specifically, the tender requirement document refers to the official document issued by the tenderer that contains project technical requirements, legal constraints, and scoring rules, usually in PDF or Word format; technical parameters refer to the technical requirements such as equipment performance indicators (e.g., "Server CPU main frequency ≥ 3.0GHz") and acceptance criteria (e.g., "System response time < 2 seconds") clearly specified in the tender document; legal terms refer to the legally binding content related to contract performance, intellectual property rights, liability for breach of contract, etc. (e.g., "The bidder needs to provide three years of free quality assurance"); scoring weights refer to the quantitative indicators of the importance of each part of the tender document by the bid evaluation committee (e.g., "Technical solution accounts for 40%, business qualifications account for 30%").
[0023] Specifically, extracting the technical requirement paragraphs using natural language processing techniques includes the following steps: Use a PDF parsing library (such as PyPDF2 or PDFMiner) to extract the document layout information, identify chapter titles, numbering systems (e.g., "3.1.2 Performance Requirements"), and paragraph boundaries. For example, lock the content under the title "Chapter 4 Technical Specifications" and exclude non-technical chapters such as "Business Terms" and "Qualification Requirements"; Adopt a paragraph classifier based on a pre-trained model (such as BERT or RoBERTa). After inputting the paragraph text, output a binary classification result of "technical requirement / non-technical requirement". Among them, the training data of the classifier is a dataset of tender documents annotated with technical requirement paragraphs (positive samples) and non-technical paragraphs (negative samples such as legal statements and company introductions); If a paragraph continuously contains ≥ 3 technical keywords (such as "performance", "configuration", "test"), it is determined as a technical requirement paragraph. For example, when phrases such as "server configuration requirements" and "acceptance test criteria" appear in the paragraph, trigger the technical paragraph marker.
[0024] Furthermore, tender documents usually have unstructured features, and it is difficult for traditional regular expression matching to handle diverse layout formats. Combining the dual verification of layout analysis and semantic classification can accurately identify technical requirement paragraphs and avoid missing complex scenarios such as nested tables and cross-page content.
[0025] Specifically, annotating technical parameters through named entity recognition includes the following steps: entity label definition, model training, and industry term enhancement.
[0026] Specifically, entity label definition includes the following steps: Define three types of technical entity tags, for example: device models (such as "Huawei Atlas 800"), performance indicators (such as "Throughput ≥ 10 Gbps", "MTBF ≥ 100,000 hours"), acceptance criteria (such as "Needs to pass ISO9001 certification", "Third-party inspection report"); Use annotation tools (such as Label Studio) to perform entity annotation on historical tender documents to construct a training dataset; Adopt a BiLSTM-CRF model (or a domain-adapted BERT model) for training. The input of the model is the technical paragraph text, and the output is the entity tag sequence.
[0027] Specifically, industry term enhancement refers to loading industry dictionaries (such as "ICT Device Model Library", "National Standard Performance Indicator Terminology Table") as external feature inputs to improve the accuracy of entity recognition. For example, the dictionary presets "Intel Xeon Gold 6338" as a device model. When the model encounters an out-of-vocabulary word "AMD EPYC 7763", it can infer that it is a device model through character-level features.
[0028] Furthermore, format verification of the entities output by the model includes: the device model needs to conform to the pattern of "brand name + letter / number combination" (such as "Cisco Catalyst 9500"), and the performance indicator needs to contain a comparison operator (such as "≥", "<") or a dimension unit (such as "Gbps", "dB"). For example, extract "10 Gbps" as the performance indicator from the text "The throughput of the firewall is not less than 10 Gbps" and associate its constraint condition "not less than".
[0029] Generally speaking, traditional rule templates (such as regular expressions) are difficult to cover the diversity of technical parameter expressions (such as "≥ 10 Gbps", "greater than or equal to 10 Gbps", "10 Gbps and above"). The named entity recognition model based on deep learning can adapt to different expression methods and, combined with domain dictionary enhancement, significantly improve the recall rate and accuracy.
[0030] S2. Based on the technical parameters, retrieve matching historical tender fragments and associated technical specifications from a preset industry knowledge base to generate a retrieval-enhanced dataset.
[0031] In the embodiment of the present invention, the retrieving matching historical tender fragments and associated technical specifications from a preset industry knowledge base based on the technical parameters includes: Generate multi-dimensional retrieval keywords according to the technical parameters; Based on the multi-dimensional retrieval keywords, use an inverted index to screen out candidate fragments from a preset industry knowledge base; Sort the candidate segments based on cosine similarity, and determine the historical tender segments and associated technical specifications that are most relevant to the technical parameters according to the sorted candidate segments.
[0032] Specifically, the industry knowledge base refers to a structured database that contains historical tender texts and technical specification documents (such as national standards and industry white papers), and supports multi-dimensional retrieval; historical tender segments refer to the paragraph texts such as technical solutions, equipment configurations, and acceptance processes in past successfully bid tenders; associated technical specifications refer to document segments such as industry standards (such as "GB / T 20278-2023 Information Security Technology") and certification requirements (such as "ISO 27001 certification") that match the current tender requirements; the retrieval-enhanced dataset refers to a structured data set formed through retrieval and sorting, which contains historical tender segments and their associated technical specifications and is used for subsequent tender generation.
[0033] Specifically, generating multi-dimensional retrieval keywords includes the following steps: Use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to extract high-frequency key terms from the technical parameters. For example, in the technical parameter "Firewall throughput ≥ 10 Gbps", the TF-IDF scores of "firewall", "throughput", and "10 Gbps" are relatively high; Based on a domain knowledge graph (such as an ICT industry knowledge graph), expand synonyms and related terms. For example, "throughput" is expanded to "bandwidth" and "data transfer rate", and a pre-trained word vector model (such as Word2Vec or GloVe) is used to calculate word similarity, and the expanded words with a similarity > 0.7 are selected; Split the technical parameters into three categories: device model, performance index, and acceptance standard, and generate keyword groups respectively to form multi-dimensional retrieval conditions.
[0034] Generally speaking, traditional single-keyword retrieval is easily affected by expression differences (such as "10 Gbps" and "10 gigabits per second"), and multi-dimensional expansion ensures retrieval coverage and avoids missing related content with similar semantics but different expressions.
[0035] Specifically, screening candidate segments based on an inverted index includes the following steps: Use Elasticsearch to establish an inverted index for the tenders and technical specifications in the industry knowledge base. The index fields include text content, release time, document type, etc. Among them, the index strategy refers to setting different weights for device models (exact match), performance indicators (phrase match), and acceptance criteria (fuzzy match); Execute multi - condition combination queries, including: MUST condition: at least match one device model keyword (such as "HuaweiUSG6600"); SHOULD condition: match performance indicator or acceptance standard keywords, with bonus points according to TF - IDF weights; FILTER condition: limit the validity period of technical specifications (such as the release time within the past 3 years). According to the query score (BM25 algorithm of Elasticsearch), initially screen the top - N (such as Top200) candidate segments to avoid excessive data volume during semantic similarity calculation.
[0036] Generally speaking, the inverted index supports fast retrieval of large - scale text, combines boolean logic to precisely control the screening conditions, and ensures that the candidate segments simultaneously meet device matching and technical relevance.
[0037] Specifically, sorting the candidate segments based on cosine similarity includes the following steps: Use the Sentence - BERT model (pre - trained model: all - mpnet - base - v2) to convert technical parameters and candidate segments into 768 - dimensional sentence vectors. For example, the vector representation of the technical parameter "Firewall throughput ≥ 10Gbps" is 0.23, - 0.56,..., 0.78; Calculate the cosine similarity between the technical parameter vector and each candidate segment vector. The similarity range is [-1, 1], and the closer the value is to 1, the more similar the semantics; Sort the candidate segments in descending order of similarity, retain the segments with a score > 0.6 (the threshold is determined by validating historical data), and for technical specification documents, additionally add a timeliness weight (such as the weight of specifications released in 2023 × 1.2, and those before 2020 × 0.8) to adjust the final sorting.
[0038] Generally speaking, cosine similarity can effectively measure the semantic relevance of high - dimensional vectors, solving the limitation of the inverted index relying only on keyword matching. For example, although the candidate segment "Firewall bandwidth can support 10 gigabits per second" does not contain the keyword "10Gbps", it can still be recognized as highly relevant through vector similarity.
[0039] S3. Input the retrieved and enhanced dataset into a pre - trained large - language model, and fuse the historical tender segments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial tender draft.
[0040] In the embodiments of the present invention, fusing the historical tender segments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial tender draft includes: Calculate the semantic similarity between the historical tender segments and the current technical parameters to generate a first weight coefficient; Extract the timeliness index of the associated technical specification, and generate a second weight coefficient based on the release time and revision version in the timeliness index; Obtain the feedback score of the user on the historical tender fragment, and generate a third weight coefficient; Normalize the first weight coefficient, the second weight coefficient, and the third weight coefficient; Dynamically adjust the fusion ratio of the historical tender fragment and the associated technical specification according to the normalized weight coefficient, and generate a preliminary tender draft that meets the current technical requirements based on the fusion ratio.
[0041] Specifically, the pre-trained large language model refers to a generative model based on the Transformer architecture (such as GPT-4, BART), which obtains semantic understanding and generation capabilities through massive text pre-training; the dynamic weight allocation algorithm refers to the calculation rule for dynamically adjusting the fusion weight of historical fragments and technical specifications according to semantic relevance, timeliness, and user feedback; the timeliness index refers to the timeliness quantification parameter of the technical specification, including the national standard version number (such as "GB / T2023-001"), the validity period of industry certification (such as "ISO27001:2022 valid until 2025").
[0042] Specifically, the calculation of the semantic similarity adopts a sentence vector matching model based on BERT.
[0043] Specifically, the calculation method of the semantic similarity weight is the same as the similarity calculation method in sorting the candidate fragments based on the cosine similarity.
[0044] Specifically, Sentence-BERT performs better than the traditional bag-of-words model in semantic matching tasks, can recognize the semantic equivalence of "10Gbps" and "ten-gigabit bandwidth", and avoid the limitations of keyword literal matching.
[0045] Specifically, the calculation of the timeliness weight includes the following steps: extract the timeliness index, calculate the decay factor according to the time decay function, and determine the version number weight, and then generate the timeliness weight.
[0046] Specifically, the extraction of the timeliness index includes: the national standard version number, using regular expressions to match the "GB / TXXXX-YYYY" format, and extracting the release year YYYY (such as "GB / T36637-2018" → 2018); the validity period of industry certification, parsing the "valid until YYYY-MM-DD" field in the technical specification, and calculating the remaining number of valid days.
[0047] Furthermore, the decay factor represents the interval between the release time of the technical specification and the current time (unit: year), and the decay factor is calculated using the formula where, is the attenuation factor, is the current year, is the release time, is the attenuation rate (e.g., represents a 10% attenuation per year), is the current year. For example, for a specification released in 2022 (current year 2025), then .
[0048] Specifically, after linearly normalizing the version number "YYYY", the formula for calculating the version number weight is as follows:
[0049] where, is the version number weight, is the version year, is the base year, and the base year is the earliest standard year in the industry (e.g., 2010), is the current year.
[0050] Furthermore, the calculation of the comprehensive weight needs to combine time attenuation and version weight. For example:
[0051] Assume that for a certain related technical specification , then .
[0052] Specifically, the timeliness indicators include the national standard version number of the technical specification and the industry certification validity period.
[0053] Generally speaking, the dual timeliness indicators (release time and version number) avoid the limitations of single time attenuation. For example, an old version standard may still be valid due to not being updated (such as some national standards not being revised for ten years).
[0054] Specifically, the calculation of the user feedback score weight includes the following steps: Record the "adoption rate" (the number of times a fragment is retained / the total number of displays) and "edit distance" (the number of characters modified by the user / the length of the original text) of the user for historical tender fragments. For example: A certain fragment is displayed 10 times, and 8 of them are adopted, then the adoption rate is 0.8, and the average edit distance is 0.1 (minor modification); Assume that the adoption rate weight is equal to the adoption rate, and the edit distance weight is equal to 1 minus the edit distance. Then the formula for the comprehensive score is , and use the Sigmoid function to normalize to , then: , where is the slope parameter (such as ), which makes the score concentrated in the range of 0.3 - 0.7, avoiding the influence of extreme values.
[0055] Generally speaking, integrating the adoption rate and the degree of modification can more comprehensively reflect user preferences. For example, the weight of a frequently adopted but slightly modified fragment should be higher than that of a less frequently adopted but unmodified fragment.
[0056] Furthermore, the normalization process for the first weight coefficient, the second weight coefficient, and the third weight coefficient can be carried out using the following formula:
[0057] Assume the first weight coefficient , the second weight coefficient , and the third weight coefficient , then the normalized weight is .
[0058] Furthermore, according to the normalized weight coefficients, dynamically adjust the fusion ratio of the historical tender document fragments and the associated technical specifications, and generating an initial tender document draft that meets the current technical requirements includes: splicing the historical tender document fragments and the associated technical specifications according to the weights as input text, for example: [weight 0.41] Historical fragment A: Firewall throughput 12Gbps... [weight 0.36] Technical specification B: Requirements of GB / T36637-2018... [weight 0.23] Historical fragment C: Support for dual-machine hot standby... Furthermore, when inputting into the large language model, the high-weight text is placed at the front of the prompt words to enhance the attention allocation during generation.
[0059] Specifically, during the generation process of the large language model (such as GPT-4), weight information is injected through prefix tuning to constrain the consistency of the output with the high-weight content. For example, if the weight of the technical specification is high, the model preferentially generates a description of "meeting the requirements of GB / T36637-2018".
[0060] Generally speaking, Softmax normalization ensures that the sum of the weights is 1, avoiding the imbalance of the fusion ratio; the combination of weighted splicing and prefix tuning enables the model to explicitly perceive the input priorities and reduces the risk of the generated content deviating from the technical requirements.
[0061] Specifically, after generating the initial tender document draft, it further includes: Establish a user feedback-driven dynamic optimization mechanism to collect in real time the editing behavior data of the tender document fragments in the initial tender document draft. The editing behavior data includes text modification trajectories and paragraph adoption rates. Convert the editing behavior data into a fine-tuning parameter set for the large language model through an online learning algorithm, and dynamically inject the fine-tuning parameter set during the subsequent tender document generation process to form a weight allocation strategy adaptive to user preferences. Construct a two-way association index between user feedback data and the historical tender document fragments. When it is detected that the user repeatedly modifies the technical parameters, automatically trigger the version update verification of the associated technical specifications in the industry knowledge base.
[0062] Specifically, the editing behavior data refers to the direct operation records of the user on the tender document draft, including text modification trajectories (such as deletion, replacement, addition of content) and paragraph adoption rates (the number of times a fragment is retained / the total number of display times); the online learning algorithm refers to a machine learning method that can update model parameters in real time (such as incremental learning, continuous learning), without the need to retrain with all data; the two-way association index refers to a data structure that establishes a two-way mapping relationship between user feedback data and historical tender document fragments, supporting forward (tracing feedback from the fragment) and reverse (locating the fragment from the feedback) queries; the version update verification refers to an automated process of detecting the latest version of the technical specifications in the industry knowledge base and determining whether it is necessary to replace the old version content.
[0063] Specifically, the collection and structuring of editing behavior data include: using a difference algorithm (such as the Myers difference algorithm) to compare the version changes of the tender document draft and record the following information: Modification type: deletion (DEL), insertion (INS), replacement (REP); Modification location: paragraph ID, line number, character offset; Modified content: original text fragment and modified text.
[0064] Specifically, the collection and structuring of editing behavior data also include: calculating the paragraph adoption rate, counting the number of occurrences and the number of times retained of each historical fragment in the generated draft, and dividing the number of times retained by the number of occurrences to obtain the adoption rate. For example: a certain fragment is displayed 10 times, and 8 of them are not deleted by the user, then the adoption rate is 0.8.
[0065] Generally speaking, the difference algorithm accurately captures fine-grained modifications, providing a data basis for analyzing user preferences; the adoption rate quantifies the universality of the fragment, avoiding the accidental influence of single modifications.
[0066] Specifically, online learning and model fine-tuning include the following steps: converting editing behavior into training data, updating incremental learning parameters, and dynamically injecting fine-tuning parameters.
[0067] Further, the conversion of editing behavior into training data includes: using the tender document modified by the user as the positive sample and the original generated content as the negative sample to construct a contrastive learning dataset. For example, the positive sample is "Throughput ≥ 12 Gbps" (modified by the user); the negative sample is "Throughput ≥ 10 Gbps" (originally generated).
[0068] Specifically, the update of incremental learning parameters can adopt the HAT (Hierarchical Adaptive Transfer) algorithm. Based on a pre-trained large language model (such as GPT-4), it is fine-tuned through the following loss function:
[0069] Among them, is the cross-entropy loss, which forces the large language model to output close to the content modified by the user; is the KL divergence loss, which constrains the consistency of the output distributions of the new model and the original model to prevent catastrophic forgetting; is the balance coefficient such as , is the true value corresponding to the new content that represents the user's expected output of the large language model and is one of the goals for the model to learn, is the output value predicted by the model. During the model training process, the parameters are continuously adjusted to make close to , is the probability distribution output by the original model, reflecting the output characteristics of the pre-trained model before the current incremental learning, is the probability distribution output by the new model (the model after updating the parameters through incremental learning), prompting to be as similar as possible to .
[0070] Specifically, the dynamic injection of fine-tuning parameters means saving the fine-tuned model parameters (such as the weights of the attention layer ) as an incremental parameter set and dynamically loading them through parameter interpolation during subsequent generation:
[0071] Among them, is the injection intensity (such as ), and the injection intensity can be dynamically adjusted according to the confidence of user feedback.
[0072] Generally speaking, contrastive learning directly aligns with user preferences; the HAT algorithm avoids model forgetting caused by traditional fine-tuning and ensures generation stability.
[0073] Specifically, the construction of the bidirectional association index and the version update verification include steps such as index construction, duplicate correction detection, and version update verification.
[0074] Furthermore, index construction refers to using a graph database (such as Neo4j) to construct a bidirectional association index. The nodes include: user feedback nodes, such as: edit behavior ID, modified content, user ID; historical fragment nodes, such as: fragment ID, source tender document, technical parameters.
[0075] Specifically, duplicate correction means that the same technical parameter (such as "firewall model") is modified more than a threshold number of times (such as 3 times).
[0076] Specifically, when it is detected that "Model A" is changed to "Model B" 5 times, an associated technical specification update check is triggered.
[0077] Specifically, version update verification can be performed by accessing the specification update interface of the industry knowledge base (such as the API of the Standardization Administration of China) to check the following information of the associated specification: Latest version number: Compare the currently used version with the latest released version; Summary of updated content: Extract key change points (such as "the throughput requirement is increased from 10 Gbps to 12 Gbps").
[0078] Furthermore, if an update is detected, an automatic reminder is pushed and it is recommended to replace the old content.
[0079] Generally speaking, the graph database efficiently manages complex association relationships; the version verification API docking ensures the timeliness of technical specifications and avoids the disconnection between tender document content and the latest standards.
[0080] S4. Perform compliance verification on the initial tender document draft according to the legal provisions, and identify and mark the conflicting clauses.
[0081] In the embodiment of the present invention, the performing compliance verification on the initial tender document draft according to the legal provisions, and identifying and marking the conflicting clauses includes: Extract the key obligations and restrictive conditions in the legal provisions, and construct a legal semantic knowledge graph; Map the initial tender document draft to the legal semantic knowledge graph paragraph by paragraph according to each legal provision, and detect the coverage integrity of the initial tender document draft for the legal provisions; Identify the expressions in the initial tender document draft that conflict with the legal provisions; Perform multi-level severity marking on the identified conflicting clauses, where the conflict levels of the conflicting clauses are divided into mandatory conflicts and recommended conflicts according to severity.
[0082] Specifically, a legal semantic knowledge graph refers to a structured knowledge base centered around legal provisions, including "subject - obligation - condition" triples, logical relationships between provisions, and hierarchical systems; coverage integrity refers to the inspection result of whether the tender draft covers all mandatory obligations in legal provisions; conflicting provisions refer to statements in the tender content that have semantic contradictions or formal inconsistencies with legal provisions; the cross - document alignment module refers to a semantic matching model based on Transformer, which is used to identify semantic conflicts between different documents.
[0083] Specifically, the construction of a legal semantic knowledge graph includes the following steps: extraction of key obligations and restrictive conditions, storage of the graph, and definition of logical constraints.
[0084] Furthermore, use a pre - trained model in the legal domain (such as Legal - BERT) to identify the obligatory subjects (such as "supplier"), actions (such as "provide"), objects (such as "three - year quality guarantee"), and conditions (such as "within 30 days after acceptance") in the provisions, and use a sequence annotation model (such as SpanBERT) to construct "subject - obligation - condition" triples. For example: extract the triple: (supplier, submit inspection report, within 30 days after acceptance) from the provision "The supplier shall submit an inspection report within 30 days after acceptance"; Specifically, use a graph database (Neo4j) to store the triples and define logical rules: mandatory obligations: provisions containing modal words such as "must" and "shall" (such as "must provide a three - year quality guarantee"), recommended provisions: provisions containing modal words such as "recommend" and "should" (such as "recommend using energy - saving certified products"). For example: mark the node attribute {type: "mandatory", basis: "Article XX of the Contract Law"} in the knowledge graph.
[0085] Specifically, the section - by - section mapping and coverage integrity detection of the tender draft include the following steps: Segment the tender draft by chapter, and use the attention mechanism to calculate the correlation between each paragraph and legal provisions. For example: the "quality guarantee clause" chapter of the tender is automatically associated with the legal provision "quality assurance obligation in the Contract Law"; Traverse all mandatory obligation nodes in the knowledge graph, check whether there are corresponding statements in the tender. If a certain obligation node has no associated tender paragraph, mark it as "coverage missing". For example: there is a node "The supplier shall provide a three - year quality guarantee" in the knowledge graph, but the tender does not mention the quality guarantee period → trigger a missing alarm.
[0086] Specifically, traditional keyword matching is prone to missing semantic - equivalent expressions (such as "36 - month quality guarantee period" vs "three - year quality guarantee"), and the attention mechanism can capture deep semantic associations.
[0087] Specifically, the expression for identifying the conflict between the initial tender draft and the legal provisions adopts the semantic matching model of the cross-document alignment module of the Transformer architecture.
[0088] Specifically, a two-tower Transformer structure is adopted to encode the legal provisions (Tower A) and the tender paragraphs (Tower B) respectively, and semantic alignment is performed through the cross-attention layer.
[0089] For example: The model inputs are as follows: Tower A: [CLS] Legal provision text [SEP] Tower B: [CLS] Tender paragraph text [SEP] Furthermore, the model output is: the conflict probability value (0 - 1) and the conflict type (mandatory / recommendatory).
[0090] Specifically, an example of conflict detection is: the legal provision is "The warranty period shall not be less than three years", and the tender paragraph is "Provide a two-year warranty service", then the model output is a conflict probability of 0.92 and the type "mandatory conflict".
[0091] Specifically, the dataset of the semantic matching model is labeled with legal provision - tender paragraph pairs with conflict types (such as "conflict / no conflict").
[0092] Furthermore, the two-tower structure can independently process long texts, and the cross-attention can accurately locate the conflict points (such as the quantitative words "two years" vs "three years"), which is superior to traditional rule matching.
[0093] Specifically, the mandatory conflict determination rule refers to violating the legal mandatory obligations (such as insufficient warranty period), inconsistent subject qualifications (such as lack of legal certification). For example: the tender "Two-year warranty period" vs the law "≥ three years", then it belongs to a mandatory conflict; the recommendatory conflict determination rule refers to not conforming to industry practices or recommendatory terms (such as not using energy-saving products but the law only recommends). For example: the tender does not mention "It is recommended to use domestic equipment", then it belongs to a recommendatory conflict.
[0094] Furthermore, insert a mark in the tender draft, marked as: [Conflict: Mandatory] Warranty period: Two years (in accordance with Article XX of the "Contract Law").
[0095] Specifically, the hierarchical annotation guides differential processing (such as mandatory conflicts must be modified, and recommendatory conflicts can be selectively optimized), improving the correction efficiency.
[0096] Specifically, after identifying and marking the conflict clauses, it also includes: Generating differential correction suggestions according to the conflict level and embedding them into the revision log of the initial tender draft; Retain the revision traces and legal basis of the conflict clauses in the target tender document, where the revision traces include: the original text of the conflict clause, the revised version, and the corresponding legal article number.
[0097] Specifically, the differential amendment suggestions refer to the targeted modification plans generated according to the conflict level (mandatory / recommended), including the items that must be modified and the items recommended for optimization; the revision log refers to the structured document recording the modification process of the tender draft, including the conflict clauses, the modified content, the modifier, and the timestamp; the revision traces refer to the comparison information of the original text and the revised version of the conflict clauses retained in the final tender document, which is used for auditing and legal traceability; the legal article number refers to the specific articles of the laws and regulations corresponding to the conflict clauses (such as Article X of the "Contract Law").
[0098] Specifically, the mandatory conflict amendment rules are as follows: Pre-define a legal compliance template library, and automatically fill in the modified content according to the conflict type. Example: Conflict "Warranty period: two years" → Call the template "Warranty period shall be ≥ three years (based on Article XX of the 'Contract Law')"; Identify the conflict value and replace it with the legally required value. Example: Detect that "Report submission deadline: 20 days" conflicts with the law "≤ 15 days" → Replace it with "15 working days".
[0099] Specifically, the recommended conflict optimization suggestions are as follows: Provide multiple compliance options for users to choose from. Example: Conflict "Energy-saving products are not adopted" → Suggestion "Option A: Supplement energy-saving certification; Option B: Explain the legitimate reasons for not adopting"; Extract the common practices of similar projects from the industry knowledge base. Example: Suggestion "Refer to 90% of similar projects to supplement the explanation of the localization rate".
[0100] Generally speaking, template filling ensures strict compliance with legal clauses, multi-scheme recommendation improves the flexibility of user choice, and structured logs are convenient for traceability.
[0101] Specifically, use a version control tool (such as Git) in the final tender document to manage the revision history and retain the following information: Original text of the conflict clause: Displayed in strikethrough (such as "Warranty period: two years"); Revised version: Highlighted (such as "Warranty period: three years"); Legal basis: Marked in the form of footnotes or hyperlinks (such as "① Article XX of the 'Contract Law'").
[0102] Specifically, construct a legal article database (such as the "China Lawinfo" API), and reverse-match the article number through the clause content. Example: Input "The minimum warranty period is three years", and it will return "Article XX, Paragraph XX of the 'Contract Law'".
[0103] Specifically, natural language generation technology (such as the T5 model) converts the modification record into readable text, for example: "The original 'two-year warranty period' does not meet the requirements of Article XX of the Contract Law and has been amended to 'three years'. This modification was automatically completed by the system on August 20, 2023." Generally speaking, version control tools ensure the traceability of the revision process; legal provision APIs solve the problem of low efficiency in manual search; natural language generation improves the readability of the description.
[0104] S5. Optimize the structure of the initial tender draft based on the scoring weights to generate a target tender that complies with the tender evaluation rules.
[0105] In the embodiment of the present invention, the optimizing the structure of the initial tender draft based on the scoring weights to generate a target tender that complies with the tender evaluation rules includes: Sort the priorities of each chapter of the initial tender draft according to the scoring weights; Based on the sorting result of the priority sorting, insert a strengthened description of the technical solution into the chapter whose weight ratio exceeds the preset ratio threshold; Generate a table of contents index for the target tender according to the key points of the tender evaluation rules, and highlight the key content in the target tender.
[0106] Specifically, the preset ratio threshold refers to the minimum weight ratio that triggers content strengthening (for example, if the chapter weight ≥ 15%, it needs to be strengthened); the strengthened description of the technical solution refers to adding supplementary content such as technical details, data support, or differential advantages in the key chapter; the table of contents index refers to the navigation structure generated based on the evaluation rules, highlighting the hierarchical relationship of high-weight chapters; the key highlighting refers to visually enhancing the key content that affects the evaluation (such as bold, highlighting, sidebar notes).
[0107] Specifically, analyze the tender evaluation rules table of the tender document, establish the mapping relationship between the chapter name and the weight, use regular expressions to match the chapter titles of the tender draft (such as "A technical solution"), and associate the corresponding weights.
[0108] Specifically, sort the chapters in descending order of weight, and perform secondary sorting on the chapters with the same weight according to the following rules: the priority of the chapters containing keywords such as "technical parameters" and "innovation points" is increased; the priority of the chapters with a high user historical editing frequency is increased.
[0109] Generally speaking, weight mapping ensures that the sorting complies with the bid evaluation criteria, and secondary sorting introduces user behavior data to avoid mechanical weight allocation.
[0110] Specifically, based on the proportion = chapter weight / total score × 100%, calculate the proportion of the chapter weight. Assume that the set proportion threshold is T (e.g., T = 15%). For the chapters with a weight proportion greater than the proportion threshold T (such as the technical solution accounts for 40 / 100 = 40%), trigger content enhancement.
[0111] Specifically, the enhanced content generation includes: retrieving the technical indicators of similar projects from the industry knowledge base and inserting a comparison table; using a large language model (such as GPT-4) to generate a summary of technical advantages. For example: This solution adopts a distributed architecture, which improves the throughput by 50% compared with the traditional solution; insert the content at the beginning of the chapter (abstract enhancement) or after the technical parameter list (detail enhancement), and determine the best position through the semantic similarity of paragraphs.
[0112] Specifically, the threshold mechanism focuses on high-value chapters; data comparison and differential description directly enhance the potential of the bid evaluation score.
[0113] Specifically, generate a multi-level table of contents according to the priority ranking. Promote the high-weight chapters to the first-level headings (such as "1. Technical Solution"), and the subheadings reflect the bid evaluation rules. For example: 1. Technical Solution (40 points), 1.1 Core Technical Indicators (15 points), and 1.2 Innovation Points (10 points).
[0114] Specifically, the key content marking includes: rule key matching and visual enhancement.
[0115] Furthermore, rule key matching means using a keyword extraction algorithm to extract high-frequency terms (such as "innovation" and "energy saving") from the bid evaluation rules and marking the corresponding content in the bid text.
[0116] Specifically, the visual enhancement includes: adding a background color (such as light yellow) to the sentences that match the bid evaluation terms; adding a reference to the bid evaluation criteria at the edge of the page (such as "▲ Corresponding to bid evaluation rule 1.2: The highest score for innovation is 10 points").
[0117] Specifically, the table of contents weight label helps the bid evaluation committee quickly locate the key points; the visual marking ensures that the key scoring points are not missed.
[0118] In several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways.
[0119] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0120] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, and technology that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A tender document generation method based on retrieval-augmented generation and large language models, characterized in that, The method includes: S1. Obtain the tender requirement document of the target tender, and parse the technical parameters, legal terms, and scoring weights in the tender requirement document; S2. Based on the technical parameters, retrieve the matching historical tender fragments and associated technical specifications from a preset industry knowledge base, and generate a retrieval-enhanced dataset; S3. Input the retrieval-enhanced dataset into a pre-trained large language model, and fuse the historical tender fragments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial tender draft; S4. Conduct a compliance check on the initial tender draft according to the legal terms, and identify and mark the conflicting terms; S5. Optimize the structure of the initial tender draft based on the scoring weights to generate a target tender that meets the tender scoring rules.
2. The bid document generation method based on retrieval-augmented generation and large language model according to claim 1, characterized in that, The parsing of the technical parameters in the tender requirement document includes: Using natural language processing technology to extract the technical requirement paragraphs in the tender requirement document; Marking the equipment models, performance indicators, and acceptance criteria in the technical requirement paragraphs through named entity recognition.
3. The bid document generation method based on retrieval-augmented generation and large language model according to claim 1, characterized in that, The retrieving of the matching historical tender fragments and associated technical specifications from a preset industry knowledge base based on the technical parameters includes: Generating multi-dimensional retrieval keywords according to the technical parameters; Based on the multi-dimensional retrieval keywords, using an inverted index to screen out candidate fragments from a preset industry knowledge base; Sorting the candidate fragments based on cosine similarity, and determining the historical tender fragments and associated technical specifications most relevant to the technical parameters according to the sorted candidate fragments.
4. The bid document generation method based on retrieval augmented generation and large language model according to claim 1, wherein, The fusing of the historical tender fragments and the associated technical specifications through a dynamic weight allocation algorithm to generate an initial tender draft includes: Calculating the semantic similarity between the historical tender fragments and the current technical parameters to generate a first weight coefficient; Extracting the timeliness indicators of the associated technical specifications, and generating a second weight coefficient based on the release time and revision version in the timeliness indicators, where the timeliness indicators include the national standard version number of the technical specifications and the industry certification validity period; Obtaining the feedback score of the user on the historical tender fragments to generate a third weight coefficient; Normalizing the first weight coefficient, the second weight coefficient, and the third weight coefficient; Dynamically adjusting the fusion ratio of the historical tender fragments and the associated technical specifications according to the normalized weight coefficients, and generating an initial tender draft that meets the current technical requirements based on the fusion ratio.
5. The method for generating a tender document based on retrieval-augmented generation and a large language model according to claim 4, wherein, The calculation of the semantic similarity uses a sentence vector matching model based on BERT.
6. The method for generating a tender document based on retrieval-augmented generation and a large language model according to claim 1, wherein The conducting of a compliance check on the initial tender draft according to the legal terms, and identifying and marking the conflicting terms includes: Extracting the key obligations and restrictive conditions in the legal terms, and constructing a legal semantic knowledge graph; Mapping the initial tender draft to the legal semantic knowledge graph paragraph by paragraph according to each legal term, and detecting the coverage integrity of the initial tender draft for the legal terms; Identifying the expressions in the initial tender draft that conflict with the legal terms; Perform multi-level severity annotation on the identified conflicting clauses. Among them, the conflict levels of the conflicting clauses are classified into mandatory conflicts and recommended conflicts according to severity.
7. The method for generating a tender document based on retrieval-augmented generation and a large language model according to claim 6, wherein, After identifying and annotating the conflicting clauses, it also includes: Generate differentiated amendment suggestions according to the conflict levels and embed them into the revision log of the initial tender draft. Retain the revision traces and legal bases of the conflicting clauses in the target tender. Among them, the revision traces include: the original text of the conflicting clause, the revised version, and the corresponding legal article number.
8. The method for generating tender documents based on retrieval-augmented generation and large language models according to claim 6, wherein The semantic matching model of the cross-document alignment module including the Transformer architecture is used to identify the expressions in the initial tender draft that conflict with the legal clauses.
9. The bid generation method based on retrieval-augmented generation and large language model according to claim 1, wherein The structural optimization of the initial tender draft based on the scoring weights to generate a target tender that complies with the tender evaluation rules includes: Sort the priorities of each chapter of the initial tender draft according to the scoring weights. Based on the sorting results of the priority sorting, insert a strengthened description of the technical solution into the chapter where the weight ratio exceeds the preset ratio threshold. Generate a table of contents index for the target tender according to the key points of the tender evaluation rules and highlight the key contents in the target tender.
10. The method for generating a tender document based on retrieval-augmented generation and a large language model according to claim 1, wherein, After generating the initial tender draft, it also includes: Establish a user feedback-driven dynamic optimization mechanism to collect the editing behavior data of the tender fragments in the initial tender draft in real time. The editing behavior data includes the text modification track and the paragraph adoption rate. Convert the editing behavior data into a fine-tuning parameter set of the large language model through an online learning algorithm and dynamically inject the fine-tuning parameter set during the subsequent tender generation process to form a weight allocation strategy that adapts to user preferences. Construct a two-way association index between the user feedback data and the historical tender fragments. When it is detected that the user repeatedly modifies the technical parameters, automatically trigger the version update verification of the associated technical specifications in the industry knowledge base.
Citation Information
Patent Citations
Large language model retrieval enhancement generation method based on hierarchical information expansion
CN118779425A
Large language model generation method based on retrieval enhancement
CN119782470A
Cited By
Bid invitation purchasing bid evaluation method and system based on big language model technology
CN120805927A
Laboratory quality management document intelligent generation method and system based on retrieval enhancement
CN120951955A
Data segmentation method and system based on machine learning
CN121009890A
Bank intelligent operation knowledge base implementation method and system based on large language model and RAG
CN121071162A
Implementation method and system of bank intelligent operation knowledge base based on large language model and RAG
CN121071162B