Model training method and device, electronic equipment, storage medium and program product

By defining structured specifications and training methods for the evaluation knowledge base of large language models, the problem of unreliable evaluation results is solved, generating structured and auditable evaluation reports, thereby improving the accuracy and efficiency of evaluation decisions.

CN121808370APending Publication Date: 2026-04-07CHINA SHENHUA INT CONSTR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, when using large language models for bid evaluation, the generated bid evaluation results cannot guarantee that they contain the necessary audit elements, the internal logic is unreliable, and they cannot meet the rigid requirements of the bidding and tendering field for auditable processes.

Method used

By defining a structured and standardized training sample set and evaluation knowledge base, a retrieval-enhanced supervised fine-tuning method is used to train the initial evaluation model, ensuring that the evaluation reports generated by the model are machine-readable data with clear structure and well-defined fields, thus ensuring that the output results conform to the rigorous reasoning and compliance of evaluation experts.

Benefits of technology

It achieves the standardization and auditability of bid evaluation results, significantly improves the accuracy and reliability of bid evaluation decisions, reduces the risk of misjudgment due to limited knowledge, and supports a fully automated intelligent bid evaluation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808370A_ABST
    Figure CN121808370A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of bid evaluation, and discloses a model training method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining an initial bid evaluation model; a training sample set is obtained, each training sample in the training sample set is a thinking chain conforming to a structured specification, the structured specification defines a plurality of fields, and the plurality of fields comprise bid evaluation items, bid evaluation bases, technical analysis, compliance judgment and scores; and training the initial bid evaluation model by adopting a retrieval enhanced supervised fine tuning method, training the initial bid evaluation model to generate an output result conforming to the structured specification, and enabling the content of the bid evaluation basis in the output result to belong to the retrieved bid evaluation knowledge. According to the mode, the bid evaluation model is trained to perform reasoning according to verifiable external knowledge, so that'illusion 'caused by the fact that the model depends on internal parameterized knowledge is fundamentally avoided, the accuracy and the reliability of a bid evaluation result are improved, and the bid evaluation result contains necessary auditing elements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of bid evaluation, and in particular, relate to a model training method and device, electronic equipment, storage medium and program product. BACKGROUND

[0002] With the development of intelligent bid evaluation, the use of large language model reasoning to conduct intelligent bid evaluation is also attracting more and more attention. How to use large language models to better perform the bid evaluation task is a hot issue in intelligent bid evaluation.

[0003] In related technologies, a general Chain-of-Thought (CoT) dataset containing logical reasoning steps is used to perform SFT (Supervised Fine-Tuning) on a base model (a pre-trained large language model) to make the base model learn to "think". Then, in the application stage, the bid documents and related regulations documents are constructed into a vector database. When a bid evaluation question is raised, the relevant text fragments are first retrieved from the vector database based on semantic similarity, and these fragments are input into the base model fine-tuned with the CoT dataset together with the bid evaluation question. The base model synthesizes the information and outputs a bid evaluation result containing a reasoning process in the form of free text.

[0004] However, the way of generating bid evaluation results in the form of free text cannot guarantee the inclusion of necessary elements for auditing, and even less can it guarantee the reliability of the internal logic. SUMMARY

[0005] The purpose of the present application is to at least provide a model training method and device, electronic equipment, storage medium and program product, which can at least solve the problem that the bid evaluation result cannot guarantee the inclusion of necessary elements for auditing and the reliability of the internal logic, and can at least achieve the effect of generating a bid evaluation result containing necessary elements for auditing and reliable internal logic.

[0006] To solve the above technical problems, at least one embodiment of the present application provides a model training method, comprising: obtaining an initial bid evaluation model; obtaining a training sample set, each training sample in the training sample set being a structured thinking chain, the structured thinking chain defining a plurality of fields, the plurality of fields including a bid evaluation item, a bid evaluation basis, a technical analysis, a compliance judgment, and a score, the bid evaluation item being used to describe the evaluation content required in the bidding document, the bid evaluation basis being used to describe the bid evaluation knowledge cited for the bid evaluation item, the technical analysis being used to describe the technical analysis of the corresponding content in the bid document for the bid evaluation item, the compliance judgment being used to describe the compliance conclusion of the bid document with respect to the bid evaluation item, and the score being used to describe the score of the bid document on the bid evaluation item; and training the initial bid evaluation model by using a retrieval-enhanced supervised fine-tuning method; wherein, in the training process, for any training sample, bid evaluation knowledge matching the bid evaluation item in the any training sample is retrieved from a pre-constructed bid evaluation knowledge base, and prompt content is constructed based on the retrieved bid evaluation knowledge and the bid evaluation item in the any training sample, so that the initial bid evaluation model is trained to generate an output result conforming to the structured thinking chain based on the prompt content and the any training sample, and the content of the bid evaluation basis in the output result belongs to the retrieved bid evaluation knowledge.

[0007] By defining the mandatory structured specification including the bid evaluation item, the bid evaluation basis, the technical analysis, the compliance judgment, and the score, the bid evaluation report (output result) generated by the bid evaluation model is standardized data with clear structure and explicit fields, ensuring the normativity and auditability of the output result, greatly facilitating machine automatic auditing and traceability, and meeting the rigid demand for process auditability in the bidding field. By retrieving bid evaluation knowledge matching the bid evaluation item from the pre-constructed bid evaluation knowledge base and forcing the content of the bid evaluation basis in the output result to belong to the retrieved bid evaluation knowledge, the bid evaluation model is trained to necessarily reason based on verifiable external knowledge (such as regulations), fundamentally avoiding the "hallucination" (i.e. generating seemingly reasonable but actually incorrect conclusions) of the model due to reliance on internal parameterized knowledge, and significantly improving the accuracy and reliability of bid evaluation decisions.

[0008] Further, the large language model (initial bid evaluation model) is fine-tuned by the field-specific training sample and the bid evaluation knowledge base, effectively injecting the thinking mode and field knowledge of bid evaluation experts into the model, enabling the model to transform from a general-purpose dialogue tool into a professional assistant proficient in bidding business, and capable of handling high-specification and strong-compliance professional evaluation tasks.

[0009] In some examples, the obtaining the training sample set comprises: obtaining a pre-trained language model; generating an initial train of thought set conforming to the structured specification based on the content of the bidding document and the bid document by using the pre-trained language model; receiving a modification instruction, and modifying each initial train of thought in the initial train of thought set according to the modification instruction to review, revise and supplement the initial train of thought according to the opinions of the bid evaluation expert, to obtain the training sample set.

[0010] The pre-trained language model is used to automatically generate an initial train of thought, which greatly reduces the labor cost and time cost of constructing high-quality training data. By reviewing, revising and supplementing the initial train of thought according to the modification instruction, i.e., introducing the bid evaluation expert for fine bidding, the data quality in terms of professionalism, accuracy and compliance is ensured. Through the data production mode of human-computer cooperation, the efficiency of data preparation is significantly improved under the premise of ensuring data quality. Since the training sample ultimately reflects the review opinions of the bid evaluation expert, the model can learn not only the surface format but also the deep reasoning mode and compliance judgment standard possessed by the expert, thereby directly improving the ability of the model to make correct judgments in complex and fuzzy scenarios.

[0011] In some examples, the bid evaluation knowledge base stores at least one of bidding regulations, industry standards, historical cases and bidding document templates.

[0012] By explicitly specifying that the knowledge base includes bidding regulations, industry standards, historical cases and bidding document templates, the model retrieval and referenced knowledge are ensured to have high authority and coverage, so that the model can obtain multi-dimensional and multi-level domain knowledge during training and reasoning, whether it is macro legal regulations or micro specific project requirements, which can be effectively supported, thereby making its decision-making basis more sufficient and reliable, and reducing the risk of misjudgment due to one-sided knowledge.

[0013] In some examples, the output result is in a machine-readable structured data format.

[0014] Limiting the output result to machine-readable structured data (e.g., JSON, XML) breaks down the barriers between AI decision-making and existing IT systems. The output bid evaluation report can be directly read, parsed and stored by audit programs, archive management systems, etc., without the need for manual intervention for format conversion or content extraction. This greatly improves the overall efficiency of bid evaluation work and provides a technical foundation for fully automatic and high-concurrency intelligent bid evaluation processes.

[0015] In some examples, the reward model used during the initial evaluation model training process includes at least one of knowledge grounding reward, structure format reward, and final compliance reward; the knowledge grounding reward is positively correlated with the number of target knowledge items in the evaluation criteria of the output result and negatively correlated with the number of non-target knowledge items in the evaluation criteria of the output result. Knowledge that exists in the evaluation criteria of the output result, the retrieved evaluation knowledge, and the evaluation criteria of any training sample is target knowledge, and the content in the evaluation criteria of the output result other than target knowledge is non-target knowledge; the structure format reward is used to evaluate whether the format of the output result conforms to a machine-readable structured data format and whether the data structure conforms to the structured specification; the final compliance reward is used to evaluate whether the compliance judgment content in the output result is consistent with the compliance judgment content in any training sample, and / or, to evaluate the degree of consistency between the score in the output result and the score in any training sample.

[0016] The knowledge grounding reward directly drives the model to generate "evaluation criteria" that are highly consistent with the retrieved evaluation knowledge, strengthening the factual basis and verifiability of the output. The structure format reward ensures that the model's output strictly follows the preset structured specifications, guaranteeing the consistency and machine readability of the output results, which is key to achieving automated auditing. Finally, the compliance reward, by comparing the consistency between the model output and the training samples, forces the model to align with expert standards in key decisions such as "compliance judgment" and "score," systematically reducing the risk of the model generating illegal conclusions.

[0017] Furthermore, the synergistic optimization of the three reward objectives enables the trained model to not only perform well in a single dimension, but also achieve a balanced and high level of performance in accuracy, standardization, and compliance, thus becoming a truly reliable and trustworthy intelligent evaluation system.

[0018] In some examples, the expression for the knowledge-based reward is:

[0019] in, Knowledge-based rewards and All are preset coefficients greater than 0. The number of target knowledge items. The number of non-target knowledge items; The expression for the reward structure is:

[0020] in, Rewards are given for structural formatting; , , and All are preset coefficients greater than 0; The value is 1 when the output result is in a machine-readable structured data format, and -1 or 0 when the output result is not in a machine-readable structured data format. The number of fields defined by the structured specification included in the output result; The number of fields in the output that do not belong to the structured specification definition; The number of fields missing from the structured specification definition in the output result; When the final compliance reward is used to evaluate whether the compliance judgment in the output result is consistent with the compliance judgment in any training sample, and to evaluate the degree of consistency between the score in the output result and the score in any training sample, the expression for the final compliance reward is:

[0021] in, For ultimate compliance rewards; in, For the score in any of the training samples, the The score in the output result; , All are preset coefficients greater than 0; The value is 1 when the compliance judgment content in the output result is consistent with the compliance judgment content in any training sample, and 0 or -1 when the compliance judgment content in the output result is inconsistent with the compliance judgment content in any training sample.

[0022] At least one embodiment of this application also provides a model training apparatus, comprising: a first acquisition module for acquiring an initial evaluation model; and a second acquisition module for acquiring a training sample set, wherein each training sample in the training sample set is a thought chain conforming to a structured specification, the structured specification defining multiple fields, the multiple fields including evaluation items, evaluation basis, technical analysis, compliance judgment, and score, wherein the evaluation items are used to describe the evaluation content required in the tender documents, the evaluation basis is used to describe the evaluation knowledge referenced for the evaluation items, the technical analysis is used to describe the technical analysis of the corresponding content in the tender documents for the evaluation items, and the compliance judgment is used to describe the tender documents. The score, relative to the compliance conclusion of the evaluation item, is used to describe the score of the bid document on the evaluation item; the training module is used to train the initial evaluation model using a retrieval-enhanced supervised fine-tuning method; wherein, during the training process, for any training sample, evaluation knowledge matching the evaluation item in the pre-built evaluation knowledge base is retrieved from the pre-built evaluation knowledge base, and prompt content is constructed based on the retrieved evaluation knowledge and the evaluation item in the pre-built evaluation knowledge. Based on the prompt content and the pre-built evaluation knowledge, the initial evaluation model is trained to generate output results that conform to the structured specifications, and the content of the evaluation basis in the output results belongs to the retrieved evaluation knowledge.

[0023] In some optional embodiments, the second acquisition module is used to acquire a pre-trained language model; using the pre-trained language model, based on the content of the tender documents and bid documents, generate an initial thought chain set that conforms to the structured specifications; receive modification instructions, and modify each initial thought chain in the initial thought chain set according to the modification instructions, so as to review, correct and supplement the initial thought chains according to the opinions of the evaluation experts, and obtain the training sample set.

[0024] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the model training method described above.

[0025] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the model training method described above.

[0026] At least one embodiment of this application also provides a computer program product, including a computer program that, when executed by a processor, implements the model training method described above. Attached Figure Description

[0027] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0028] Figure 1 This is a flowchart of a model training method provided in one embodiment of this application; Figure 2 This is a flowchart of a bid evaluation method provided in another embodiment of this application; Figure 3 This is a schematic diagram of a model training apparatus provided in another embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0030] To facilitate understanding of the embodiments of this application, the relevant content on model training in intelligent evaluation will be introduced first.

[0031] Large Language Models and Chain of Reasoning (CoT): Large Language Models (LLMs) have achieved breakthroughs in natural language understanding and generation; CoT hints or fine-tuning techniques significantly improve model performance on arithmetic, common sense, and symbolic reasoning tasks by guiding the model to generate intermediate reasoning steps. Currently, mainstream methods for improving reasoning ability include: Supervised fine-tuning (SFT): Fine-tuning a pre-trained model using a high-quality "question-thought chain-answer" triplet dataset; Preference alignment: Such as reinforcement learning based on human feedback or direct preference optimization, which aligns the model behavior with human preferences by comparing the merits of different outputs.

[0032] Retrieval-Augmented Generation (RAG): Before generating an answer, relevant information is retrieved from an external knowledge base as context to reduce model illusions and improve the factual accuracy of the answer.

[0033] Business and regulatory background of bidding and evaluation: Bidding and evaluation is an economic activity strictly regulated by law; its core requirements are "fairness, impartiality, science, and selection of the best." Evaluation experts must conduct independent and recorded reviews of each bid in strict accordance with the bidding documents (containing technical requirements, commercial terms, and evaluation criteria) and relevant national regulations; any evaluation conclusion must be supported by clear clauses, and the entire process needs to be archived for future reference.

[0034] In related technologies, a general-purpose thought chain dataset containing logical reasoning steps is used to perform SFT on the base model, enabling the base model to "think". Then, in the application phase, tender documents and relevant regulatory documents are constructed into a vector database. When an evaluation question is posed, several relevant text fragments are first retrieved from the vector database through semantic similarity. These fragments, along with the evaluation question, are then input into the base model fine-tuned using the CoT dataset. After the model synthesizes the information, it outputs an evaluation result containing the reasoning process in free text form.

[0035] The problems existing in the related technologies are as follows: Decoupling of training and inference: CoT's "reasoning ability" training and RAG's "knowledge acquisition" are largely separate. The model learns general reasoning patterns in the SFT stage and is not specifically trained on how to accurately and critically integrate the retrieved, noisy context into a rigorous, structured evaluation domain reasoning framework.

[0036] Lack of structured and compliance constraints: The model output remains unstructured free text, failing to guarantee the inclusion of all necessary audit elements (such as supporting clauses, analysis process, and judgment conclusions), and further failing to guarantee that its internal logic strictly conforms to the compliance requirements of bid evaluation. Furthermore, there is no explicit optimization objective for compliance during training. The free text output leads to the unstructured and unauditable nature of the decision-making path, a serious flaw for bid evaluation activities requiring rigorous, efficient, and large-scale audits. Regulatory agencies or review panels cannot quickly verify the source and logical validity of each judgment step in a machine-readable manner, making the decision-making process a "black box" and failing to meet the requirements of transparency and traceability in bidding processes. The core of bid evaluation is compliance judgment, whose logic is deterministic and verifiable. However, the training paradigm of LLM in related technologies is probabilistic, with alignment objectives of "usefulness" and "harmlessness," lacking an explicit and powerful optimization objective for "legal / clause compliance." This allows the model to generate seemingly reasonable conclusions that violate specific mandatory clauses, leading to significant legal and commercial risks.

[0037] Poor data quality and domain adaptability: The reasoning paradigm trained using the general CoT dataset differs greatly from the rigorous, reference-driven thinking mode of bidding evaluation. This means that even if the model obtains the correct knowledge fragments, it may not be able to apply them correctly. Consequently, when the model reasons in professional fields such as bidding evaluation, the intermediate steps and conclusions it generates lack mandatory references and bindings to specific bidding documents, laws and regulations, and other "foundational facts." This can easily generate "reasonable illusions" that contradict the facts, which is unacceptable in bidding evaluation activities.

[0038] To address the aforementioned technical problems, this application proposes a model training method. The implementation details of the model training method in this embodiment are described below. The following content is only for ease of understanding and is not necessary for implementing this solution.

[0039] Example 1: The model training method in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. Its specific process can be as follows: Figure 1 As shown, it includes S101-S103.

[0040] S101, Obtain the initial evaluation model.

[0041] The initial evaluation model can be any model capable of performing evaluation reasoning tasks.

[0042] In some cases, the initial evaluation model is a large language model, such as DeepSeek and the GPT series.

[0043] In some cases, the initial evaluation model is a large language model with a small number of parameters that is easy to deploy and maintain.

[0044] S102, Obtain the training sample set. Each training sample in the training sample set is a thought chain that conforms to the structured specification. The structured specification defines multiple fields, including the evaluation items, evaluation basis, technical analysis, compliance judgment, and score.

[0045] The evaluation item describes the evaluation content required in the bidding documents; the evaluation basis describes the evaluation knowledge referenced for the evaluation item; the technical analysis describes the technical analysis of the corresponding content in the bid documents for the evaluation item; the compliance judgment describes the compliance conclusion of the bid documents relative to the evaluation item; and the score describes the score of the bid documents on the evaluation item.

[0046] In some examples, the structured specifications can be as shown in Table 1 below.

[0047] Table 1

[0048] This structured specification ensures that the evaluation model explicitly cites evidence during the reasoning process, rather than using vague statements.

[0049] In some examples, obtaining a training sample set may include: obtaining a pre-trained language model; using the pre-trained language model to generate an initial thought chain set that conforms to the structured specifications based on the content of the tender documents and bid documents; receiving modification instructions and modifying each initial thought chain in the initial thought chain set according to the modification instructions, so as to review, correct and supplement the initial thought chains according to the opinions of the evaluation experts, thereby obtaining the training sample set.

[0050] By automatically generating initial thought chains using pre-trained language models, the manual and time costs of constructing high-quality training data are significantly reduced. The process of "reviewing, correcting, and supplementing the initial thought chains according to modification instructions"—that is, introducing expert evaluation for rigorous review—ensures high standards of data professionalism, accuracy, and compliance. Through a human-machine collaborative data production model, the efficiency of data preparation is significantly improved while maintaining data quality. Since the training samples ultimately reflect the review opinions of the expert evaluation team, the model learns not just the superficial format, but the deep-seated, rigorous reasoning patterns and compliance judgment standards possessed by the experts, directly enhancing the model's ability to make correct judgments in complex and ambiguous scenarios.

[0051] The embodiments of this application do not limit the specific type of pre-trained language model. For example, the pre-trained language model can be GPT-4, Qwen, etc.

[0052] In some examples, the pre-trained language model is used to generate an initial set of thought chains that conforms to the structured specifications based on the content of the tender documents and bid documents. This may include: constructing prompt words based on the content that needs to be reviewed in the tender documents and the content in the bid documents that corresponds to the content that needs to be reviewed; inputting the prompt words into the pre-trained language model, which then generates an initial set of thought chains that conforms to the structured specifications.

[0053] For example, the contents of the tender documents include: Section 3.2.1 Storage System Requirements: To ensure data security and reliability, the server storage system provided by the bidder must support RAID 5 (Redundant Array of Independent Disks Level 5) disk redundancy technology, and the bidder shall provide an official technical white paper or relevant supporting materials in the bid documents.

[0054] The contents of the tender documents include: Server Configuration: We recommend using the XX model server, whose storage controller supports multiple RAID configurations, including RAID 0, RAID 1, and RAID 10, providing high performance and data protection.

[0055] Based on the content of the tender document and the tender documents, the following prompts can be constructed: "You are a bid evaluation expert. Please strictly follow the contents of the provided 'Tender Documents' and 'Bid Documents' and generate a bid evaluation thought process according to the following machine-readable structured data format."

[0056] Evaluation item: Review whether the server in the bid documents meets the technical requirement of 'supporting RAID 5' in the tender documents.

[0057] Evaluation criteria: The specific clauses of the bidding documents or relevant laws and regulations must be directly cited.

[0058] Technical Analysis: Compare and analyze the content of the tender documents with the bidding requirements.

[0059] Compliance assessment: Provide a clear conclusion of 'compliant' or 'non-compliant'.

[0060] Score: Based on the evaluation method in the bidding documents, give the score and explain the reasons.

[0061] [Tender Document Contents]: Section 3.2.1, Storage System Requirements: To ensure data security and reliability, the server storage system provided by the bidder must support RAID 5 disk redundancy technology, and the bidder shall provide an official technical white paper or relevant supporting materials in the tender document.

[0062] [Tender Document Contents]: Server Configuration: We recommend using the XX model server. This server's storage controller supports multiple RAID configurations, including RAID 0, RAID 1, and RAID 10, providing high performance and data protection.

[0063] After inputting the prompt word into the pre-trained language model, the initial draft of the model's thought process output is as follows: { "Evaluation Item": "Evaluate whether Bidder A's server meets the technical requirement of 'supporting RAID 5' in the tender document." Evaluation criteria: "The tender documents require that the server must support RAID5 technology." Technical Analysis: "The bidder's technical proposal mentions support for RAID 0, 1, and 10. RAID 5 is a striping technology with parity checking, which differs from RAID 10. Based on the description, the bidder may not support RAID 5." "Compliance assessment": "Does not comply" "Score": "This technical indicator is not met; according to the scoring criteria, this item receives 0 points." }

[0064] The evaluation experts analyzed that the initial thought process had the following problems: 1. Inaccurate evaluation criteria: The evaluation criteria only vaguely mention "requirements in the tender documents" without citing specific clause numbers and original text, which does not meet the auditability requirements; 2. Technical analysis is speculative: it uses uncertain terms such as "maybe" and lacks professionalism; 3. The scoring derivation is not rigorous: the specific clauses of the scoring criteria on which it is based are not explained.

[0065] Based on the problems existing in the initial thought chain, the evaluation experts modify the initial thought chain using the electronic device. The modification opinions of the evaluation experts are input through the input device provided by the electronic device, so that the electronic device generates corresponding modification instructions. The initial thought chain is modified according to the modification instructions, so as to review, correct and supplement the initial thought chain according to the modification opinions of the evaluation experts, and obtain the training sample corresponding to the initial thought chain.

[0066] For example, the training sample is as follows: { "Evaluation Item": "Evaluate whether the bidder's server meets the technical requirement of 'supporting RAID5' in Clause 3.2.1 of the tender document." Evaluation Criteria: "According to Article 3.2.1 of the tender document, 'The server storage system provided by the bidder must support RAID5 disk redundancy technology.'" Technical Analysis: "Upon review of the 'Server Configuration' section of the tender documents, it clearly lists supported RAID levels as RAID 0, RAID 1, and RAID 10, but does not list RAID 5, and no official technical documentation supporting RAID 5 is provided." "Compliance assessment": "Does not comply" "Score": "According to the provisions in the attached table of the 'Evaluation Method' for this project," If the key technical indicator is not met, this item will receive 0 points. 'Support RAID 5' is a key technical indicator specified in Clause 3.2.1 of the tender document (with...). (The score for this item is 0, therefore the score is 0.) }

[0067] In another example, obtaining the training sample set can include: using zero-shot or self-consistency hint engineering, a language model generates a large number of evaluation thought chains according to structured specifications, which are then screened and quality-filtered by evaluation experts or a more powerful language model to obtain multiple filtered evaluation thought chains, which constitute the training sample set. This method can reduce the cost of manual annotation.

[0068] S103, the initial evaluation model is trained using a retrieval-enhanced supervised fine-tuning method.

[0069] During the training process, for any training sample, evaluation knowledge matching the evaluation items in the pre-built evaluation knowledge base is retrieved from the pre-built evaluation knowledge base. Based on the retrieved evaluation knowledge and the evaluation items in the pre-built evaluation knowledge base, prompt content is constructed. Based on the prompt content and the pre-built evaluation knowledge base, the initial evaluation model is trained to generate output results that conform to the structured specifications, and the content of the evaluation basis in the output results belongs to the retrieved evaluation knowledge.

[0070] The embodiments of this application do not limit how to retrieve evaluation knowledge that matches the evaluation items in any training sample from the pre-built evaluation knowledge base.

[0071] In some cases, the similarity between the evaluation items in any training sample and each piece of evaluation knowledge in the evaluation knowledge base can be calculated, and the top n (n is an integer greater than 0) pieces of evaluation knowledge in the similarity ranking can be selected as the evaluation knowledge that matches the evaluation items in any training sample.

[0072] The embodiments of this application do not limit the method used for similarity calculation. For example, cosine similarity, L1 similarity, etc., can be used.

[0073] For example, the training samples are: { "Evaluation Item": "Evaluate whether the server in bid A meets the RAID support requirements". "Evaluation Criteria": "According to the requirements of Tender_2024_001_Sec3.2.1 in the tender document: 'Storage System: The tendered products must support RAID 5 disk redundancy technology'" "Technical Analysis": "Upon reviewing Bidder A's technical specifications, it was found that the supported RAID levels are clearly listed as 0, 1, and 10, but RAID 5 is not mentioned." "Compliance assessment": "Does not comply" "Score Derivation / Conclusion": "This technical requirement is not met; according to the evaluation method, this item scores 0 points." }

[0074] The evaluation knowledge retrieved from the evaluation knowledge base based on the evaluation item "whether the server of bid A meets the RAID support requirements" in the training sample is as follows: 1. Document ID: First code, content is: Storage system: The tendered product must support RAID 5 disk redundancy technology and provide official technical white paper proof.

[0075] 2. Document ID: Second code, content is: The bid evaluation committee shall conduct a systematic review and comparison of the bid documents in accordance with the bid evaluation standards and methods stipulated in the bidding documents.

[0076] 3. Document ID: Third code, content is: Our (Bidder A) server supports multiple RAID configurations, including RAID0, RAID 1, and RAID 10, with excellent performance.

[0077] Based on the retrieved evaluation knowledge and the evaluation items in the training samples, the suggested content can be: "You are an intelligent bid evaluation expert. Please generate a detailed bid evaluation reasoning chain based strictly on the provided [bid evaluation items] and [relevant bid evaluation knowledge]."

[0078] [Evaluation Item]: Evaluate whether the server in bid A meets the RAID support requirements.

[0079] [Relevant Bidding Evaluation Knowledge]: 1. Document ID: First code, content is: Storage System: The tendered product must support RAID 5 disk redundancy technology and provide official technical white paper proof; 2. Document ID: Second code, content is: The bid evaluation committee shall conduct a systematic review and comparison of the bid documents in accordance with the bid evaluation standards and methods stipulated in the bidding documents; 3. Document ID: Second code, content is: Our (Bidder A) server supports multiple RAID configurations, including RAID0, RAID 1, and RAID 10, with excellent performance.

[0080] Please output a machine-readable structured data chain, which must include the following fields: evaluation items, evaluation basis, technical analysis, compliance judgment, and score. Additionally, ensure that the content of the "evaluation basis" field is directly quoted from specific clauses in the [Relevant Evaluation Knowledge].

[0081] In some examples, training the initial evaluation model based on the prompt content and any training sample may include: inputting the prompt content into the initial evaluation model, generating an output result from the initial evaluation model; calculating a reward based on the output result and the any training sample; and optimizing and adjusting the parameters of the initial evaluation model based on the reward.

[0082] In some cases, a bid evaluation knowledge base needs to be built before training the initial bid evaluation model. This knowledge base may store at least one of the following: bidding regulations, industry standards, historical precedents, and bid document templates.

[0083] Taking the bid evaluation knowledge base as an example, which stores bidding regulations, industry standards, historical precedents and bidding document templates, the construction process of the bid evaluation knowledge base may include: slicing all documents (including bidding regulations, industry standards, historical precedents and bidding document templates) into multiple knowledge segments (each knowledge segment corresponds to one knowledge), then using an embedding model to generate a corresponding vector for each knowledge segment, and storing the vectors corresponding to each knowledge segment into the bid evaluation knowledge base.

[0084] In some cases, after storing the vectors corresponding to each knowledge fragment into the evaluation knowledge base, an index can be created for the vectors corresponding to each knowledge fragment.

[0085] For example, a knowledge fragment in the evaluation knowledge base can be: Document ID: First Code Content: "'Storage System: Bidding products must support RAID 5 disk redundancy technology and provide official technical white paper proof' (vector form)" Source: Article 3.2.1 of the "XX Project Tender Document - Technical Specifications".

[0086] By explicitly including "bidding regulations, industry standards, historical precedents, and bidding document templates" in the knowledge base, the model ensures that the knowledge retrieved and referenced by the model has high authority and coverage. This allows the model to acquire multi-dimensional and multi-level domain knowledge during training and inference, effectively supporting both macro-level laws and regulations and micro-level specific project requirements. This makes the model's decision-making basis more sufficient and reliable, reducing the risk of misjudgment due to incomplete knowledge.

[0087] By defining mandatory structured specifications that include "evaluation items, evaluation basis, technical analysis, compliance judgment, and scores," the evaluation reports (output results) generated by the evaluation model are standardized data with clear structure and well-defined fields. This ensures the standardization and auditability of the output results, greatly facilitating automated machine review and traceability, and meeting the rigid requirements of auditable processes in the bidding and tendering field. By "retrieving evaluation knowledge matching the evaluation items from a pre-built evaluation knowledge base" and forcing "the content of the evaluation basis in the output results to belong to the retrieved evaluation knowledge," the evaluation model is trained to reason based on verifiable external knowledge (such as legal provisions). This fundamentally avoids the "illusion" (i.e., generating seemingly reasonable but factually incorrect conclusions) caused by the model's reliance on internal parameterized knowledge, significantly improving the accuracy and reliability of evaluation decisions.

[0088] Furthermore, the large language model (initial evaluation model) is fine-tuned through domain-specific training samples and an evaluation knowledge base, effectively injecting the thinking patterns and domain knowledge of evaluation experts into the model. This transforms the model from a general dialogue tool into a professional assistant proficient in bidding and tendering business, capable of handling high-standard and highly compliant professional review tasks.

[0089] In some cases, the evaluation model outputs in a machine-readable structured data format. Both during training and application, the model's output is constrained to be in a machine-readable structured data format.

[0090] Limiting the output to machine-readable structured data (e.g., JSON, XML) breaks down the barriers between AI decision-making and existing IT systems. The output evaluation reports can be directly read, parsed, and stored by auditing programs, document management systems, etc., without manual intervention for format conversion or content extraction. This significantly improves the overall efficiency of the evaluation process and provides the technological foundation for fully automated, high-concurrency intelligent evaluation workflows.

[0091] In some cases, the reward model used during the initial evaluation model training process includes at least one of the following: knowledge grounding reward, structure format reward, and final compliance reward.

[0092] Among them, the knowledge grounding reward is positively correlated with the number of target knowledge items in the evaluation criteria of the output results and negatively correlated with the number of non-target knowledge items in the evaluation criteria of the output results. Knowledge that exists in the evaluation criteria of the output results, the retrieved evaluation knowledge, and the evaluation criteria of any training sample is considered target knowledge, while content other than target knowledge in the evaluation criteria of the output results is considered non-target knowledge.

[0093] The structure format award is used to evaluate whether the output format conforms to a machine-readable structured data format and whether the data structure conforms to structured specifications.

[0094] The final compliance reward is used to evaluate whether the compliance judgments in the output are consistent with the compliance judgments in any training sample, and / or to evaluate the degree of consistency between the score in the output and the score in any training sample.

[0095] The knowledge grounding reward directly drives the model to generate "evaluation criteria" that are highly consistent with the retrieved evaluation knowledge, strengthening the factual basis and verifiability of the output. The structure format reward ensures that the model's output strictly follows the preset structured specifications, guaranteeing the consistency and machine readability of the output results, which is key to achieving automated auditing. Finally, the compliance reward, by comparing the consistency between the model output and the training samples, forces the model to align with expert standards in key decisions such as "compliance judgment" and "score," systematically reducing the risk of the model generating illegal conclusions.

[0096] Furthermore, the synergistic optimization of the three reward objectives enables the trained model to not only perform well in a single dimension, but also achieve a balanced and high level of performance in accuracy, standardization, and compliance, thus becoming a truly reliable and trustworthy intelligent evaluation system.

[0097] In some examples, the expression for knowledge-based reward is:

[0098] in, Knowledge-based rewards and All are preset coefficients greater than 0. The number of target knowledge items. The number of non-target knowledge items; The expression for the structured reward is:

[0099] in, Rewards are given for structural formatting; , , and All are preset coefficients greater than 0; The value is 1 when the output result is in a machine-readable structured data format, and -1 or 0 when the output result is not in a machine-readable structured data format. The number of fields defined by the structured specification included in the output; This represents the number of fields in the output that are not defined by the structured specification. This represents the number of fields in the output that are missing from the structured specification definition. When the final compliance reward is used to evaluate whether the compliance judgments in the output are consistent with those in any training sample, and to evaluate the degree of consistency between the score in the output and the score in any training sample, the expression for the final compliance reward is:

[0100] in, For ultimate compliance rewards; in, For any training sample, the score is... The score in the output results; , All are preset coefficients greater than 0; The value is 1 when the compliance judgment content in the output result is consistent with the compliance judgment content in any training sample, and 0 or -1 when the compliance judgment content in the output result is inconsistent with the compliance judgment content in any training sample.

[0101] In some examples, the reward model is in, , and These are preset weighting coefficients. In some examples, Greater than and Greater than .

[0102] Compared to large language models trained using general CoT and traditional rule matching systems, the model training method in this embodiment offers advantages in terms of systematic improvement in compliance, interpretability, and domain depth. The leap from "potentially correct" to "provably correct": Large language models trained on general CoT, combined with RAG, may "potentially" utilize correct knowledge. However, this application, through mandatory retrieval and citation mechanisms during the training phase, and multi-objective rewards, ensures that the model's reasoning process is built upon verifiable evidence. Each step of its output logic possesses deterministic traceability, fundamentally guaranteeing the authenticity and reliability of the reasoning process.

[0103] Systematic compliance risk control: Existing technologies lack direct optimization for compliance. This invention transforms compliance from an implicit expectation into an explicit, high-weight optimization objective by leveraging expert judgments in expert thought chain (training sample) data and the strong penalty of the final compliance reward in multi-objective rewards, thus systematically reducing the risk of the model generating illegal or non-compliant conclusions.

[0104] This application achieves end-to-end automated auditing: existing technologies still require manual auditing of free text output. The model output in this application is in a machine-readable structured data format, making the "AI decision-making process" itself a data object that can be automatically parsed and verified by another program. This provides a technological foundation for achieving 100% automated and highly efficient AI decision auditing.

[0105] In some cases, techniques such as Task Arithmetic (vector fusion) can be used to fuse expert thought chains under structured norms in the evaluation domain with other general reasoning or safety alignment vectors. The fused thought chains can then be used to train the evaluation model, thereby improving the model's generalization ability and robustness while maintaining its professionalism in the evaluation domain.

[0106] Example 2: The evaluation method of this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. Its specific process can be as follows: Figure 2 As shown, it includes S201-S204.

[0107] S201, Receive input containing the evaluation task.

[0108] S202, retrieve the evaluation knowledge that matches the input from the evaluation knowledge base.

[0109] It should be noted that the bidding documents and tender documents corresponding to this evaluation task have been pre-stored in the evaluation knowledge base.

[0110] S203, construct prompt content based on the input and the evaluation knowledge that matches the input.

[0111] S204. Input the content of this question into the evaluation model, and use the evaluation model to generate a thought chain that conforms to the structured specifications.

[0112] The evaluation model is a model trained based on the model training method in Example 1 above.

[0113] For the specific implementation of S201-S204, please refer to the above embodiment 1, which will not be repeated here.

[0114] Example 3: Another embodiment of this application relates to a model training device. The implementation details of the model training device in this embodiment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution. A schematic diagram of the model training device in this embodiment can be seen as follows: Figure 3As shown, it includes: a first acquisition module 31, used to acquire an initial evaluation model; and a second acquisition module 32, used to acquire a training sample set, where each training sample is a thought chain conforming to a structured specification. The structured specification defines multiple fields, including evaluation items, evaluation basis, technical analysis, compliance judgment, and score. The evaluation items describe the review content required in the bidding documents; the evaluation basis describes the evaluation knowledge referenced for the evaluation items; the technical analysis describes the technical analysis of the corresponding content in the bidding documents for the evaluation items; and the compliance judgment describes the bidding documents relative to the evaluation items. The compliance conclusion score describes the score of the bid document on the evaluation items; the training module 33 is used to train the initial evaluation model using a retrieval-enhanced supervised fine-tuning method; during the training process, for any training sample, evaluation knowledge matching the evaluation items in any training sample is retrieved from the pre-built evaluation knowledge base, and prompt content is constructed based on the retrieved evaluation knowledge and the evaluation items in any training sample, so as to train the initial evaluation model to generate output results that conform to the structured specifications based on the prompt content and any training sample, and make the content of the evaluation basis in the output results belong to the retrieved evaluation knowledge.

[0115] In some optional embodiments, the second acquisition module 32 is used to acquire a pre-trained language model; using the pre-trained language model, based on the content of the tender documents and bid documents, generate an initial thought chain set that conforms to the structured specifications; receive modification instructions, and modify each initial thought chain in the initial thought chain set according to the modification instructions, so as to review, correct and supplement the initial thought chains according to the opinions of the evaluation experts, and obtain a training sample set.

[0116] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0117] Example 4: Another embodiment of this application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the model training methods in the above embodiments.

[0118] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0119] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory, on the other hand, is used to store data used by the processor during operation.

[0120] Example 5: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0121] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0122] Example 6: Another embodiment of this application relates to a computer program product, including a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0123] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. A model training method, characterized in that, include: Obtain the initial evaluation model; Obtain a training sample set, where each training sample is a thought chain conforming to a structured specification. The structured specification defines multiple fields, including evaluation items, evaluation basis, technical analysis, compliance judgment, and score. The evaluation items describe the review content required in the tender documents. The evaluation basis describes the evaluation knowledge referenced for the evaluation items. The technical analysis describes the technical analysis of the corresponding content in the tender documents for the evaluation items. The compliance judgment describes the compliance conclusion of the tender documents relative to the evaluation items. The score describes the score of the tender documents on the evaluation items. The initial evaluation model is trained using a retrieval-enhanced supervised fine-tuning method. During the training process, for any training sample, evaluation knowledge matching the evaluation items in the pre-built evaluation knowledge base is retrieved from the pre-built evaluation knowledge base. Based on the retrieved evaluation knowledge and the evaluation items in the pre-built evaluation knowledge base, prompt content is constructed. Based on the prompt content and the pre-built evaluation knowledge base, the initial evaluation model is trained to generate output results that conform to the structured specifications, and the content of the evaluation basis in the output results belongs to the retrieved evaluation knowledge.

2. The model training method according to claim 1, characterized in that, The acquisition of the training sample set includes: Obtain a pre-trained language model; Using the pre-trained language model, an initial set of thought chains conforming to the structured specifications is generated based on the content of the tender documents and bid documents; Receive modification instructions, modify each initial thought chain in the initial thought chain set according to the modification instructions, so as to review, correct and supplement the initial thought chains according to the opinions of the evaluation experts, and obtain the training sample set.

3. The model training method according to claim 1, characterized in that, The bid evaluation knowledge base stores at least one of the following: bidding regulations, industry standards, historical precedents, and bidding document templates.

4. The model training method according to claim 1, characterized in that, The output is in a machine-readable structured data format.

5. The model training method according to claim 1, characterized in that, The reward model used in the initial evaluation model training process includes at least one of the following: knowledge grounding reward, structural format reward, and final compliance reward; The knowledge grounding reward is positively correlated with the number of target knowledge items in the evaluation criteria of the output results and negatively correlated with the number of non-target knowledge items in the evaluation criteria of the output results. Knowledge that exists in the evaluation criteria of the output results, the retrieved evaluation knowledge, and the evaluation criteria of any training sample is target knowledge, and the content in the evaluation criteria of the output results other than target knowledge is non-target knowledge. The structure format reward is used to evaluate whether the output format conforms to a machine-readable structured data format and whether the data structure conforms to the structure specification; The final compliance reward is used to evaluate whether the compliance judgment in the output result is consistent with the compliance judgment in any training sample, and / or to evaluate the degree of consistency between the score in the output result and the score in any training sample.

6. The model training method according to claim 5, characterized in that, The expression for the knowledge-based reward is: in, Knowledge-based rewards and All are preset coefficients greater than 0. The number of target knowledge items. The number of non-target knowledge items; The expression for the reward structure is: in, Rewards are given for structural formatting; , , and All are preset coefficients greater than 0; The value is 1 when the output result is in a machine-readable structured data format, and -1 or 0 when the output result is not in a machine-readable structured data format. The number of fields defined by the structured specification included in the output result; The number of fields in the output that do not belong to the structured specification definition; The number of fields missing from the structured specification definition in the output result; When the final compliance reward is used to evaluate whether the compliance judgment in the output result is consistent with the compliance judgment in any training sample, and to evaluate the degree of consistency between the score in the output result and the score in any training sample, the expression for the final compliance reward is: in, For ultimate compliance rewards; in, For the score in any of the training samples, the The score in the output result; , All are preset coefficients greater than 0; The value is 1 when the compliance judgment content in the output result is consistent with the compliance judgment content in any training sample, and 0 or -1 when the compliance judgment content in the output result is inconsistent with the compliance judgment content in any training sample.

7. A model training device, characterized in that, include: The first acquisition module is used to acquire the initial evaluation model; The second acquisition module is used to acquire a training sample set. Each training sample in the training sample set is a thought chain that conforms to a structured specification. The structured specification defines multiple fields, including evaluation items, evaluation basis, technical analysis, compliance judgment, and score. The evaluation items describe the review content required in the bidding documents. The evaluation basis describes the evaluation knowledge referenced for the evaluation items. The technical analysis describes the technical analysis of the corresponding content in the bid documents for the evaluation items. The compliance judgment describes the compliance conclusion of the bid documents relative to the evaluation items. The score describes the score of the bid documents on the evaluation items. The training module is used to train the initial evaluation model using a retrieval-enhanced supervised fine-tuning method; During the training process, for any training sample, evaluation knowledge matching the evaluation items in the pre-built evaluation knowledge base is retrieved from the pre-built evaluation knowledge base. Based on the retrieved evaluation knowledge and the evaluation items in the pre-built evaluation knowledge base, prompt content is constructed. Based on the prompt content and the pre-built evaluation knowledge base, the initial evaluation model is trained to generate output results that conform to the structured specifications, and the content of the evaluation basis in the output results belongs to the retrieved evaluation knowledge.

8. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the model training method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the model training method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the model training method according to any one of claims 1 to 6.