Bid and tendering document analysis method, device and equipment based on large language model and medium

By using a bidding document analysis method based on a large language model, the automated analysis of bidding documents has been achieved, solving the problems of low efficiency, strong subjectivity, and high compliance risks in traditional review processes. This improves review efficiency and accuracy, and ensures the objectivity and compliance of review results.

CN121786192APending Publication Date: 2026-04-03INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional bidding and tendering reviews are inefficient, highly subjective, and carry high compliance risks. Existing technologies lack deep semantic parsing and automated analysis capabilities.

Method used

This paper adopts a bidding document analysis method based on a large language model. Through semantic segmentation, knowledge base storage, and a large model-driven analysis process, it realizes automated analysis of bidding documents, including semantic segmentation, vectorized storage, and relationship establishment. It uses a pre-set large language model to extract target bidding information that meets the core bidding information standards and performs multi-dimensional analysis.

Benefits of technology

It improved review efficiency, reduced manual retrieval time, lowered the rate of missed detections, ensured the objectivity and compliance of review results, provided digital audit evidence, and reduced the risk of legal disputes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786192A_ABST
    Figure CN121786192A_ABST
Patent Text Reader

Abstract

The invention discloses a bidding and tendering file analysis method and device based on a large language model, equipment and a medium, and relates to the field of artificial intelligence. The target file comprises a bidding file and a bidding file; performing semantic segmentation on the target file to convert the target file into structured chapter segment information; uploading the chapter segment information to a target knowledge base for vectorization storage, and establishing a chapter segment association relationship between chapter segments of the bidding document and chapter segments of the bidding document; extracting target bid invitation information corresponding to the bid invitation file from the target knowledge base by using a preset large language model, and constructing a bid invitation file information base according to the target bid invitation information; and analyzing the bidding document according to the bidding document information base, the bidding document chapter segments and the chapter segment association relationship among the bidding document chapter segments to obtain an analysis result, and determining a review report based on the analysis result. According to the invention, the efficiency, accuracy and compliance of bidding and tendering review are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, equipment and medium for analyzing bidding documents based on a large language model. Background Technology

[0002] With the widespread adoption of digital bidding platforms, the quantity and complexity of bidding documents have increased significantly. Traditional bidding review relies on manual operation, which has the following drawbacks: 1. Low efficiency: Evaluation experts need to manually search document content, which is time-consuming; 2. High subjectivity: Business / technical scores are easily influenced by expert experience, and different experts may have different evaluation conclusions for the same bid document; 3. High compliance risk: Manual verification is prone to missing key information, leading to subsequent audit disputes. Most existing bidding management systems only realize document uploading and basic information storage, lacking the ability to perform in-depth semantic analysis and automated analysis of document content. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, apparatus, device, and medium for analyzing bidding documents based on a large language model. Through semantic segmentation, knowledge base storage, and a large model-driven analysis process, it achieves automated analysis of bidding documents, solving the problems of low efficiency, strong subjectivity, and high compliance risks associated with traditional manual review. The specific solution is as follows:

[0004] Firstly, this application provides a method for analyzing bidding documents based on a large language model, including:

[0005] The target file is received based on a preset interface; the target file includes tender documents and bid documents.

[0006] The target file is semantically segmented to transform it into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, file type to which the chapter fragment belongs, and position information of the chapter fragment in the corresponding target file;

[0007] The chapter fragment information is uploaded to the target knowledge base for vectorized storage, and the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document is established based on the preset semantic matching method;

[0008] Using a pre-defined large language model, target bidding information that meets the pre-defined core bidding information standards corresponding to the bidding documents is extracted from the target knowledge base, and a bidding document information database is constructed based on the target bidding information;

[0009] The tender documents are analyzed based on the relationship between the tender document information database, the chapter segments of the tender documents, and the chapter segments of the tender documents, and the analysis results are obtained. The review report is determined based on the analysis results.

[0010] Optionally, receiving the target file based on a preset interface includes:

[0011] The system receives target files uploaded by users based on a preset interface and generates a unique task ID to identify the file analysis task.

[0012] The task ID and the corresponding initial task status are returned to the user terminal through the preset interface;

[0013] Accordingly, the bidding document analysis method based on a large language model also includes:

[0014] During the analysis of bidding documents, the current task status corresponding to the task ID is dynamically updated so that the user terminal can initiate a task status query request based on the task ID and receive the current task status corresponding to the task status query request.

[0015] Optionally, the step of semantically segmenting the target file to convert it into structured chapter fragment information includes:

[0016] The target file is subjected to text content extraction to obtain corresponding plain text data, and noise in the plain text data is removed to obtain the target text data;

[0017] The target text data is matched with chapter titles based on a preset regular expression pattern, and the target text data is split into chapter segments corresponding to each chapter title.

[0018] The target semantic role in the chapter segment is identified by a pre-trained target language model that meets the preset core semantic standard, and the segment segmentation error in the chapter segment is corrected according to the associated content of the target semantic role to obtain the target chapter segment;

[0019] Identify the chapter subheadings corresponding to the chapter segments whose length exceeds a preset length threshold in the target chapter segment, and split the chapter segments whose length exceeds the preset length threshold according to the chapter subheadings to obtain chapter sub-segments;

[0020] Based on the metadata corresponding to each target chapter segment and each chapter sub-segment, structured chapter segment information is obtained; the metadata includes the chapter segment title, chapter segment content, the file type to which the chapter segment belongs, and the position information of the chapter segment in the corresponding target file.

[0021] Optionally, uploading the chapter fragment information to the target knowledge base for vectorized storage, and establishing the chapter fragment association relationship between the bidding document chapter fragments and the tender document chapter fragments based on a preset semantic matching method, includes:

[0022] The chapter fragment content in the chapter fragment information is encoded to obtain the corresponding target semantic vector, and the target semantic vector and the corresponding metadata are saved to the target knowledge base;

[0023] The semantic similarity between chapter fragments of the tender document and chapter fragments of the bid document is calculated using the cosine similarity algorithm;

[0024] If the semantic similarity reaches a preset similarity threshold, the chapter segment association between the bidding document chapter segment and the bid document chapter segment is marked as an associated segment; if the semantic similarity does not reach the preset similarity threshold, the chapter segment association between the bidding document chapter segment and the bid document chapter segment is marked as an unassociated segment.

[0025] Optionally, after uploading the chapter fragment information to the target knowledge base for vectorized storage, the method further includes:

[0026] A semantic index is constructed using the chapter segment titles in the chapter segment information as the primary index and the file types to which the chapter segments belong as secondary indexes.

[0027] Accordingly, the step of extracting target bidding information that meets the preset core bidding information standards corresponding to the bidding documents from the target knowledge base includes:

[0028] Obtain the core bidding information extraction instruction, determine the keywords corresponding to the core bidding information extraction instruction, and extract the core bidding information of the document type as bidding document from the target knowledge base according to the semantic index and by keyword retrieval.

[0029] The core bidding information is converted according to a preset format to obtain the target bidding information; the target bidding information includes keywords, descriptive information, and scoring rules.

[0030] Optionally, the analysis of the tender documents may include price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis.

[0031] Optionally, determining the review report based on the analysis results includes:

[0032] Structured data is generated based on the results of price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis.

[0033] The structured data is rendered into a review report in a preset format using a preset template engine, and the review report is returned to the user through the preset interface.

[0034] Secondly, this application provides a tender document analysis device based on a large language model, comprising:

[0035] The file receiving module is used to receive target files based on a preset interface; the target files include tender documents and bid documents.

[0036] The semantic segmentation module is used to perform semantic segmentation on the target file to convert the target file into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, the file type to which the chapter fragment belongs, and the position information of the chapter fragment in the corresponding target file;

[0037] The data storage module is used to upload the chapter fragment information to the target knowledge base for vectorized storage, and to establish the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document based on a preset semantic matching method;

[0038] The information database construction module is used to extract target bidding information that meets the preset core bidding information standards from the target knowledge base using a preset large language model, and to construct a bidding document information database based on the target bidding information.

[0039] The document analysis module is used to analyze the tender documents based on the relationship between the tender document information database, the chapter segments of the tender documents, and the chapter segments of the tender documents, to obtain the analysis results, and to determine the review report based on the analysis results.

[0040] Thirdly, this application provides an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute the computer program to implement the aforementioned method for analyzing bidding documents based on a large language model.

[0043] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned method for analyzing bidding documents based on a large language model.

[0044] In this application, target files are received based on a preset interface; the target files include tender documents and bid documents; the target files are semantically segmented to convert them into structured chapter fragment information; the chapter fragment information includes chapter fragment titles, chapter fragment content, the file type to which the chapter fragment belongs, and the position information of the chapter fragment in the corresponding target file; the chapter fragment information is uploaded to a target knowledge base for vectorized storage, and a chapter fragment association relationship is established between the chapter fragments of the tender documents and the chapter fragments of the bid documents based on a preset semantic matching method; target tender information corresponding to the tender documents and meeting preset core tender information standards is extracted from the target knowledge base using a preset large language model, and a tender document information database is constructed based on the target tender information; the bid documents are analyzed based on the tender document information database, the chapter fragments of the tender documents, and the chapter fragment association relationship between the chapter fragments of the bid documents, and the analysis results are obtained, and a review report is determined based on the analysis results. As described above, this application transforms unstructured bidding / tender documents into structured chapter fragments through semantic segmentation, avoiding manual page-by-page information searching. These chapter fragments are uploaded to a target knowledge base for vectorized storage, and a "bidding-tender" chapter association is established. Subsequent analysis can directly retrieve associated fragments without repeated searches. A pre-defined large language model automatically extracts target bidding information that meets core bidding information standards from the knowledge base, replacing the manual extraction of key information. This process automates manual retrieval and supports concurrent analysis of multiple tender documents, significantly improving review efficiency. A bidding document information database is constructed based on the target bidding information extracted from the pre-defined large language model, clarifying a unified review benchmark and avoiding differences in expert understanding of review standards. Combining the bidding document information database with the "bidding-tender" chapter association, quantitative analysis of tender documents is performed, yielding analysis results that replace subjective judgment. This process combines quantitative analysis results with standardized benchmarks, ensuring consistent review logic across different tender documents and reducing interference from subjective human factors. Furthermore, the chapter fragment information includes the position information of the chapter fragment within the corresponding target document, accurately locating key content in the tender document and avoiding omissions due to scattered content during manual verification. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 This application discloses a flowchart of a bidding document analysis method based on a large language model.

[0047] Figure 2 This is a schematic diagram illustrating a specific method for analyzing bidding documents based on a large language model disclosed in this application;

[0048] Figure 3 This application discloses a system architecture diagram for a bidding document analysis method based on a large language model.

[0049] Figure 4 This is a schematic diagram of the structure of a bidding document analysis device based on a large language model disclosed in this application;

[0050] Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in this application. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] With the widespread adoption of digital bidding platforms, the quantity and complexity of bidding documents have increased significantly. Traditional bidding review relies on manual operation, resulting in inefficiency, strong subjectivity, and high compliance risks. Most existing bidding management systems only handle document uploading and basic information storage, lacking the ability for deep semantic analysis and automated analysis of document content. Therefore, this application provides a bidding document analysis method based on a large language model. Through semantic segmentation, knowledge base storage, and a large model-driven analysis process, it achieves automated analysis of bidding documents, solving the problems of low efficiency, strong subjectivity, and high compliance risks associated with traditional manual review.

[0053] See Figure 1 As shown in the figure, this application discloses a method for analyzing bidding documents based on a large language model, including:

[0054] Step S11: Receive target files based on a preset interface; the target files include tender documents and bid documents.

[0055] In this embodiment, the target file uploaded by the user client can first be received via a preset interface, and a unique task ID identifying the file analysis task can be generated. Then, the task ID and its corresponding initial task status can be returned to the user client via the preset interface. Accordingly, during subsequent analysis of the bidding documents, the current task status corresponding to the task ID can be dynamically updated, allowing the user client to initiate a task status query request based on the task ID and receive the current task status corresponding to the query request.

[0056] For example, the system receives tender documents and bid documents through the API (Application Programming Interface) provided to the outside world. The API can adopt a RESTful architecture, support HTTP (Hypertext Transfer Protocol) / HTTPS (Hypertext Transfer Protocol Secure), and be compatible with common document formats such as PDF (Portable Document Format) and Word.

[0057] When uploading the target file, the client can authenticate its identity using JWT (JSON Web Token, an open standard for authentication and authorization) to ensure that the API caller is a legitimate user, such as a bidding platform administrator or bidder. After the target file is uploaded, the system generates a unique task ID (Identity document) to identify this file analysis task and returns the task ID and its initial status to the client. During subsequent file analysis, the task status is updated in real time, and the client can query the current task status, such as "Parsing in progress" or "Parsing complete," through the API interface.

[0058] Step S12: Perform semantic segmentation on the target file to convert the target file into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, file type to which the chapter fragment belongs, and position information of the chapter fragment in the corresponding target file.

[0059] In this embodiment, the text content of the target file is first extracted to obtain the corresponding plain text data, and noise in the plain text data is removed to obtain the target text data. Then, the chapter titles in the target text data are matched based on a preset regular expression pattern, and the target text data is split into chapter segments corresponding to each chapter title. Next, a pre-trained target language model is used to identify target semantic roles in the chapter segments that meet preset core semantic standards, and segmentation errors in the chapter segments are corrected based on the associated content of the target semantic roles to obtain target chapter segments. Further, the chapter subtitles corresponding to the chapter segments whose length exceeds a preset length threshold are determined, and the chapter segments whose length exceeds the preset length threshold are split according to the chapter subtitles to obtain chapter sub-segments. Finally, structured chapter segment information can be obtained based on the metadata corresponding to each target chapter segment and each chapter sub-segment; wherein, the metadata includes, but is not limited to, chapter segment title, chapter segment content, the file type to which the chapter segment belongs, and the position information of the chapter segment in the corresponding target file.

[0060] For example, for unstructured files such as PDFs and scanned documents, OCR (Optical Character Recognition) technology, such as Tesseract, can be used to extract the text content from the target file, remove noise and redundant information, and preserve the original text order. For semi-structured / structured files, such as Word documents, document parsing tools, such as Apache Tika, can be used to extract the text content and preserve the original formatting tags, such as heading styles and table row and column relationships. Then, semantic segmentation of the extracted text content is performed based on NLP (Natural Language Processing) technology, specifically using the following strategies:

[0061] 1. Regular Expression Matching: Preset regular expression patterns for chapter titles, such as ^Part [I II III IV].*$“(I) Bidder Qualification Requirements”“1\. Quotation List”. Use regular expression matching to quickly locate key titles, such as “Qualification Requirements”, “Evaluation Method”, and “Technical Solution”.

[0062] 2. Semantic Role Labeling (SRL): Using pre-trained language models, such as BERT-base-chinese (BERT, Bidirectional Encoder Representation from Transformers), semantic role labeling is performed on the text content to identify core semantic roles, such as "bidder", "tenderer", and "evaluation committee", and chapters are divided according to the content associated with the roles.

[0063] 3. Nested segmentation: Long text chapters are further segmented by subheadings, such as "Technical Parameters", "Implementation Steps", and "Acceptance Criteria". Subheading boundaries are identified through syntactic analysis, such as dependency parsing.

[0064] The file is ultimately divided into independent, structured chapter segments.

[0065] Step S13: Upload the chapter fragment information to the target knowledge base for vectorized storage, and establish the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document based on the preset semantic matching method.

[0066] In this embodiment, the content of each chapter segment in the chapter segment information is first encoded to obtain the corresponding target semantic vector, and the target semantic vector and corresponding metadata are saved to the target knowledge base. Then, the semantic similarity between the chapter segments of the bidding document and the chapter segments of the tender document can be calculated using a cosine similarity algorithm. If the semantic similarity reaches a preset similarity threshold, the chapter segment association between the chapter segments of the bidding document and the chapter segments of the tender document is marked as associated segments; if the semantic similarity does not reach the preset similarity threshold, the chapter segment association between the chapter segments of the bidding document and the chapter segments of the tender document is marked as unassociated segments.

[0067] For example, word embedding models such as Sentence-BERT can be used to encode the content of each chapter segment, generating a semantic vector with a dimension of 768. The vectorization process needs to retain the metadata corresponding to the chapter segment content so that it can be associated with the context during subsequent retrieval. Then, the semantic vector and its corresponding metadata are saved to a vector database, such as Milvus or the built-in vector storage module of FastGPT (FastGenerative Pre-trained Transformer, a language model).

[0068] Furthermore, semantic matching algorithms, such as cosine similarity, are used to establish the relationship between chapter segments in the bidding documents and chapter segments in the tender documents. For example, if the semantic similarity between the "Evaluation Method - Price Score Calculation Rules" segment in the bidding documents and the "Commercial Terms - Quotation Explanation" segment in the tender documents is greater than or equal to 0.6, then the relationship between their chapter segments can be marked as "related segments".

[0069] Step S14: Extract target bidding information that meets the preset core bidding information standards from the target knowledge base using a preset large language model, and construct a bidding document information database based on the target bidding information.

[0070] In this embodiment, after uploading the chapter fragment information to the target knowledge base for vectorized storage in step S13, it may further include: constructing a semantic index using the chapter fragment title as the primary index and the file type to which the chapter fragment belongs as an auxiliary index. Correspondingly, extracting target bidding information corresponding to the bidding document and meeting the preset core bidding information standards from the target knowledge base may include: obtaining a core bidding information extraction instruction, determining the keywords corresponding to the core bidding information extraction instruction, and extracting core bidding information of the document type "bidding document" from the target knowledge base based on the semantic index and keyword retrieval. Then, the core bidding information is converted according to a preset format to obtain the target bidding information; wherein, the target bidding information includes, but is not limited to, keywords, descriptive information, and scoring rules.

[0071] Understandably, this step primarily involves using the workflow engine to call the Large Language Model (LLM) to extract target bidding information that meets preset core bidding information standards from the core sections of the bidding documents, and then generating a bidding document information database based on this target bidding information. For example:

[0072] Workflow engines, such as Apache Airflow, pre-configure task flows, including:

[0073] Task Node 1: Call the LLM to parse the "Qualification Requirements" section and extract the qualification conditions that bidders must meet, such as "Grade II or above for general contracting of building construction projects" and "ISO9001 certification".

[0074] Task Node 2: Call LLM to parse the "Evaluation Method" section and extract the scoring criteria, such as "Business score accounts for 40%, technical score accounts for 60%" and "Technical solution compliance score = (number of compliant items / total number of items) × 60".

[0075] Task Node 3: Call the LLM to parse the "Quotation Requirements" section and extract pricing rules, such as "Quotations must include VAT" and "Tax rate range 7%-13%".

[0076] LLM extracts core bidding information through commands such as "Please extract the qualification requirements that bidders must meet from the following text and return it in JSON format," outputting structured target bidding information, including:

[0077] Keywords: such as "certificate of qualification" and "ISO9001";

[0078] Description information: such as "Bidders must provide valid ISO9001 quality management system certification";

[0079] Scoring rules: For example, "If the qualification is missing, the bid will be rejected; if the certification expires, 5 points will be deducted."

[0080] Finally, a structured tender document information database, such as JSON format, can be built based on the extracted target tender information. This database can be queried via API interfaces, such as GET / bid-core-info / {taskId}, and can serve as a benchmark for subsequent tender document analysis.

[0081] Step S15: Analyze the tender documents according to the relationship between the chapter fragments of the tender documents and the chapter fragments of the tender documents, obtain the analysis results, and determine the review report based on the analysis results.

[0082] In this embodiment, the analysis of the tender documents may include price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis. Then, structured data is generated based on the results of these analyses. Finally, a preset template engine renders the structured data into a preset format review report, and the review report is returned to the user through a preset interface. For example:

[0083] 1. Price Correctness Analysis: A predefined rule engine, such as Drools, is invoked to validate the price data in the bid documents based on the "Pricing Requirements" rules in the tender document database.

[0084] (1) Comparison of uppercase and lowercase amounts: Extract the uppercase and lowercase amounts from the tender letter and verify their consistency through numerical conversion and string matching.

[0085] (2) Tax rate compliance detection: Identify the tax rate field in the quotation details and determine whether it falls within the range of valid tax rates stipulated by the state.

[0086] (3) Verification of consistency between details and total price: sum up the item amounts in the quotation list and compare them with the total bid price. If there is an error, mark it as "the total price is inconsistent with the details".

[0087] 2. Format Integrity Analysis: Using text matching and structure parsing techniques, the table of contents of the tender document is compared with the list of chapters required by the tender document.

[0088] (1) Chapter missing detection: Extract the table of contents information of the tender document, such as matching “Chapter [X] [Chapter Title]” with regular expressions, and compare it with the chapter list of the tender document one by one to identify the missing chapters.

[0089] 3. Bidder Qualification Analysis: Using entity recognition technology, such as SpaCy's Named Entity Recognition model, qualification certificate information from the bid documents, such as "Business License Number: XXX" and "ISO9001 Certification Valid Until 2026," is extracted and matched against the "Qualification Requirements" rules in the bidding document database.

[0090] (1) Qualification type matching: Determine whether the tender documents contain all the qualification types required by the tender documents, such as "Grade II General Contracting for Building Construction".

[0091] (2) Qualification validity verification: Verify the validity period of the certificate, such as whether "ISO9001 certification is valid until 2026" is later than the bid closing date.

[0092] (3) Matching degree calculation: Calculate the proportion of the number of qualified qualifications to the total number of qualifications, and output the "qualification matching degree".

[0093] 4. Business Response Analysis: Calculate the semantic similarity between the business terms in the tender documents and the "Business Requirements" rules in the tender document database.

[0094] (1) Response score: Output “Business terms responsiveness” based on similarity score.

[0095] (2) Deviation identification: If the similarity is lower than the preset threshold, extract the specific deviation clause.

[0096] 5. Technical Compliance Analysis: Semantic matching of the technical solutions in the tender documents with the "Technical Requirements" rules in the tender document database:

[0097] (1) Parameter compliance test: Extract the technical parameters in the bidding technical solution and compare them with the technical requirements in the bidding documents to determine whether they are met.

[0098] (2) Acceptance criteria matching degree: The consistency between the acceptance criteria of the tender technical solution and the requirements of the tender document is evaluated by text similarity calculation.

[0099] (3) Compliance score: Calculate the proportion of technical items that meet the requirements to the total number of technical items, and output the "Technical Compliance Score".

[0100] Furthermore, the results of the multi-dimensional analysis are summarized to generate structured data containing the following: risk warnings, such as "inconsistency between the amount in words and figures in the bid" and "missing technical solution chapters"; scoring suggestions, such as "business score: 80 points, technical score: 75 points"; links to the original text; and correlation analysis, such as "the similarity between the technical solution in the bid document and the technical requirements in the tender document is 0.92".

[0101] Then, using a template engine such as Thymeleaf or Jinja2, the structured data is rendered into an HTML / PDF review report. The review report supports the following features: dynamic tables displaying analysis results across various dimensions; color-coded risk warnings (e.g., red for high risk, yellow for medium risk); and embedded original text, allowing for quick access to screenshots or links to excerpts from the original tender documents. Finally, the review report is returned to the user via an API interface, supporting downloading in PDF / HTML format or online viewing. Report metadata, such as generation time and analyst information, is synchronously stored in a database for subsequent auditing and traceability.

[0102] The following is based on Figure 2 The process in this embodiment is summarized as follows:

[0103] First, the system receives the tender documents and bid documents uploaded by users via a pre-defined interface. Then, it performs semantic segmentation and structuring processing on the tender documents and bid documents to transform them into structured chapter fragments. These chapter fragments are then uploaded to a target knowledge base for vectorized storage. A pre-defined semantic matching method is used to establish chapter fragment relationships between the tender document and bid document chapter fragments. Next, a pre-defined large language model is used to extract target tender information from the target knowledge base that meets the pre-defined core tender information standards, achieving core information extraction from the tender documents. A tender document information database is then constructed based on this target tender information. Further, based on the tender document information database, the chapter fragment relationships between the tender documents and bid document chapter fragments, a multi-dimensional analysis is performed on the bid documents, including price correctness analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis. The analysis results are then used to generate a review report.

[0104] As shown above, this embodiment improves evaluation efficiency by avoiding the time-consuming nature of manual retrieval through semantic segmentation and knowledge base retrieval. Rule verification and semantic matching driven by a large model reduce the false negative rate and enhance accuracy. Based on structured rules and semantic similarity calculation, it reduces interference from subjective human factors, making the review results traceable and verifiable, thus ensuring objectivity. Automatically generated review reports provide digital evidence for auditing, reducing the risk of legal disputes and strengthening compliance.

[0105] See Figure 3 As shown in the figure, this application also discloses a system architecture diagram of a bidding document analysis method based on a large language model, including the following functional modules:

[0106] 1. API Interface Module: Developed based on the Spring Boot framework, it provides file upload, supports multipart / form-data format, task status query (e.g., GET / task / {taskId}), report download (e.g., GET / report / {taskId}), and supports JWT authentication to ensure interface security.

[0107] 2. File parsing module: Integrates Apache Tika, Tesseract, and spaCy tools to achieve content extraction and semantic segmentation of PDF / Word documents.

[0108] 3. Knowledge Base Module: Built on the FastGPT open-source framework, it supports storing semantic vectors in a vector database (Milvus) and provides a similarity retrieval interface.

[0109] 4. Workflow Engine Module: It uses Apache Airflow to schedule tasks and call large language models, such as calling Zhipu AI's GLM (Generalized Linear Model) through API to perform tasks such as keyword extraction and description generation, driving the analysis process.

[0110] 5. Review and Analysis Module: The built-in rule engine (Drools) stores price verification rules and format rules, and integrates semantic similarity algorithms, such as Sentence-BERT, to achieve qualification matching and business / technical scoring.

[0111] 6. Report Generation Module: Generates review reports in HTML / PDF format based on the template engine (Thymeleaf), supporting the embedding of original text screenshots, risk level annotations, and scoring suggestions.

[0112] As can be seen from the above, this embodiment solves the problems of low efficiency, strong subjectivity, and high compliance risk in traditional manual bidding by the synergistic effect of various modules, and significantly improves the level of digitalization and intelligence in the bidding process.

[0113] See Figure 4 As shown in the embodiments of this application, a tender document analysis device based on a large language model is also disclosed, including:

[0114] The file receiving module 11 is used to receive target files based on a preset interface; the target files include tender documents and bid documents.

[0115] Semantic segmentation module 12 is used to perform semantic segmentation on the target file to convert the target file into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, file type to which the chapter fragment belongs, and position information of the chapter fragment in the corresponding target file;

[0116] The data storage module 13 is used to upload the chapter fragment information to the target knowledge base for vectorized storage, and to establish the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document based on a preset semantic matching method;

[0117] The information database construction module 14 is used to extract target bidding information that meets the preset core bidding information standards from the target knowledge base using a preset large language model, and to construct a bidding document information database based on the target bidding information.

[0118] The document analysis module 15 is used to analyze the tender documents based on the relationship between the chapter fragments of the tender documents and the chapter fragments of the tender documents, obtain the analysis results, and determine the review report based on the analysis results.

[0119] In some specific embodiments, the file receiving module 11 includes:

[0120] The file receiving unit is used to receive target files uploaded by the user terminal based on a preset interface and generate a unique task ID to identify the file analysis task.

[0121] An information return unit is used to return the task ID and the corresponding initial task status to the user terminal through the preset interface;

[0122] Correspondingly, the bidding document analysis device based on a large language model also includes:

[0123] The status update unit is used to dynamically update the current task status corresponding to the task ID during the analysis of bidding documents, so that the user terminal can initiate a task status query request based on the task ID and receive the current task status corresponding to the task status query request.

[0124] In some specific embodiments, the semantic segmentation module 12 includes:

[0125] The data processing unit is used to extract the text content of the target file to obtain the corresponding plain text data, and remove noise from the plain text data to obtain the target text data.

[0126] The first chapter splitting unit is used to match the chapter titles in the target text data based on a preset regular expression pattern, and split the target text data into chapter segments corresponding to each chapter title;

[0127] The error correction unit is used to identify target semantic roles in the chapter segment that meet the preset core semantic standards using a pre-trained target language model, and to correct the segment segmentation errors in the chapter segment according to the associated content of the target semantic roles, so as to obtain the target chapter segment.

[0128] The second chapter splitting unit is used to determine the chapter subheadings corresponding to the chapter segments whose length exceeds a preset length threshold in the target chapter segment, and to split the chapter segments whose length exceeds the preset length threshold in the target chapter segment according to the chapter subheadings to obtain chapter sub-segments;

[0129] The information generation unit is used to obtain structured chapter fragment information based on the metadata corresponding to each target chapter fragment and each chapter sub-fragment; the metadata includes the chapter fragment title, chapter fragment content, the file type to which the chapter fragment belongs, and the position information of the chapter fragment in the corresponding target file.

[0130] In some specific embodiments, the data storage module 13 includes:

[0131] The vectorization processing unit is used to encode the chapter fragment content in the chapter fragment information to obtain the corresponding target semantic vector, and save the target semantic vector and the corresponding metadata to the target knowledge base;

[0132] The similarity calculation unit is used to calculate the semantic similarity between chapter fragments of the tender document and chapter fragments of the bid document using the cosine similarity algorithm;

[0133] The relationship determination unit is used to mark the chapter segment association relationship between the chapter segment of the bidding document and the chapter segment of the tender document as an associated segment if the semantic similarity reaches a preset similarity threshold, and to mark the chapter segment association relationship between the chapter segment of the bidding document and the chapter segment of the tender document as an unassociated segment if the semantic similarity does not reach the preset similarity threshold.

[0134] In some specific embodiments, the data storage module 13 further includes:

[0135] An index building unit is used to build a semantic index by using the chapter segment title in the chapter segment information as the main index and the file type to which the chapter segment belongs in the chapter segment information as the auxiliary index.

[0136] Accordingly, the information database construction module 14 includes:

[0137] The information extraction unit is used to obtain the core bidding information extraction instruction, determine the keywords corresponding to the core bidding information extraction instruction, and extract the core bidding information of the document type as bidding document from the target knowledge base according to the semantic index and through keyword retrieval.

[0138] The information conversion unit is used to convert the core bidding information into target bidding information according to a preset format; the target bidding information includes keywords, descriptive information and scoring rules.

[0139] In some specific implementations, the analysis of tender documents includes price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis.

[0140] In some specific embodiments, the file analysis module 15 includes:

[0141] The data generation unit is used to generate structured data based on the results of price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis.

[0142] The report generation unit is used to render the structured data into a review report in a preset format using a preset template engine, and return the review report to the user terminal through the preset interface.

[0143] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0144] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the bidding document analysis method based on a large language model disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0145] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0146] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0147] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the bidding document analysis method based on a large language model executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0148] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for analyzing bidding documents based on a large language model. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0150] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0151] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0152] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0153] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for analyzing bidding documents based on a large language model, characterized in that, include: Receive the target file based on the preset interface; The target documents include tender documents and bid documents; The target file is semantically segmented to transform it into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, file type to which the chapter fragment belongs, and position information of the chapter fragment in the corresponding target file; The chapter fragment information is uploaded to the target knowledge base for vectorized storage, and the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document is established based on the preset semantic matching method; Using a pre-defined large language model, target bidding information that meets the pre-defined core bidding information standards corresponding to the bidding documents is extracted from the target knowledge base, and a bidding document information database is constructed based on the target bidding information; The tender documents are analyzed based on the relationship between the tender document information database, the chapter segments of the tender documents, and the chapter segments of the tender documents, and the analysis results are obtained. The review report is determined based on the analysis results.

2. The bidding document analysis method based on a large language model according to claim 1, characterized in that, The step of receiving the target file based on the preset interface includes: The system receives target files uploaded by users based on a preset interface and generates a unique task ID to identify the file analysis task. The task ID and the corresponding initial task status are returned to the user terminal through the preset interface; Accordingly, the bidding document analysis method based on a large language model also includes: During the analysis of bidding documents, the current task status corresponding to the task ID is dynamically updated so that the user terminal can initiate a task status query request based on the task ID and receive the current task status corresponding to the task status query request.

3. The bidding document analysis method based on a large language model according to claim 1, characterized in that, The step of semantically segmenting the target file to transform it into structured chapter fragment information includes: The target file is subjected to text content extraction to obtain corresponding plain text data, and noise in the plain text data is removed to obtain the target text data; The target text data is matched with chapter titles based on a preset regular expression pattern, and the target text data is split into chapter segments corresponding to each chapter title. The target semantic role in the chapter segment is identified by a pre-trained target language model that meets the preset core semantic standard, and the segment segmentation error in the chapter segment is corrected according to the associated content of the target semantic role to obtain the target chapter segment; Identify the chapter subheadings corresponding to the chapter segments whose length exceeds a preset length threshold in the target chapter segment, and split the chapter segments whose length exceeds the preset length threshold according to the chapter subheadings to obtain chapter sub-segments; Based on the metadata corresponding to each target chapter segment and each chapter sub-segment, structured chapter segment information is obtained; the metadata includes the chapter segment title, chapter segment content, the file type to which the chapter segment belongs, and the position information of the chapter segment in the corresponding target file.

4. The bidding document analysis method based on a large language model according to claim 3, characterized in that, The step of uploading the chapter fragment information to the target knowledge base for vectorized storage, and establishing the chapter fragment association relationship between the bidding document chapter fragments and the tender document chapter fragments based on a preset semantic matching method, includes: The chapter fragment content in the chapter fragment information is encoded to obtain the corresponding target semantic vector, and the target semantic vector and the corresponding metadata are saved to the target knowledge base; The semantic similarity between chapter fragments of the tender document and chapter fragments of the bid document is calculated using the cosine similarity algorithm; If the semantic similarity reaches a preset similarity threshold, the chapter segment association between the bidding document chapter segment and the bid document chapter segment is marked as an associated segment; if the semantic similarity does not reach the preset similarity threshold, the chapter segment association between the bidding document chapter segment and the bid document chapter segment is marked as an unassociated segment.

5. The bidding document analysis method based on a large language model according to claim 1, characterized in that, After uploading the chapter fragment information to the target knowledge base for vectorized storage, the method further includes: A semantic index is constructed using the chapter segment titles in the chapter segment information as the primary index and the file types to which the chapter segments belong as secondary indexes. Accordingly, the step of extracting target bidding information that meets the preset core bidding information standards corresponding to the bidding documents from the target knowledge base includes: Obtain the core bidding information extraction instruction, determine the keywords corresponding to the core bidding information extraction instruction, and extract the core bidding information of the document type as bidding document from the target knowledge base according to the semantic index and by keyword retrieval. The core bidding information is converted according to a preset format to obtain the target bidding information; the target bidding information includes keywords, descriptive information, and scoring rules.

6. The bidding document analysis method based on a large language model according to claim 1, characterized in that, The analysis of the tender documents includes price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis.

7. The bidding document analysis method based on a large language model according to claim 6, characterized in that, The process of determining the review report based on the analysis results includes: Structured data is generated based on the results of price accuracy analysis, format integrity analysis, bidder qualification analysis, business responsiveness analysis, and technical compliance analysis. The structured data is rendered into a review report in a preset format using a preset template engine, and the review report is returned to the user through the preset interface.

8. A bidding document analysis device based on a large language model, characterized in that, include: The file receiving module is used to receive target files based on a preset interface. The target documents include tender documents and bid documents; The semantic segmentation module is used to perform semantic segmentation on the target file to convert the target file into structured chapter fragment information; the chapter fragment information includes chapter fragment title, chapter fragment content, the file type to which the chapter fragment belongs, and the position information of the chapter fragment in the corresponding target file; The data storage module is used to upload the chapter fragment information to the target knowledge base for vectorized storage, and to establish the chapter fragment association relationship between the chapter fragments of the bidding document and the chapter fragments of the tender document based on a preset semantic matching method; The information database construction module is used to extract target bidding information that meets the preset core bidding information standards from the target knowledge base using a preset large language model, and to construct a bidding document information database based on the target bidding information. The document analysis module is used to analyze the tender documents based on the relationship between the tender document information database, the chapter segments of the tender documents, and the chapter segments of the tender documents, to obtain the analysis results, and to determine the review report based on the analysis results.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the bidding document analysis method based on a large language model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs, which, when executed by a processor, implement the bidding document analysis method based on a large language model as described in any one of claims 1 to 7.