Bid document abstract generation method and system based on large language model
By constructing a domain terminology database and semantic chain for tender documents and adjusting the semantic parsing weights of the large language model, personalized tender document summaries are generated. This solves the problems of time-consuming and labor-intensive manual extraction and insufficient accuracy of general models in existing technologies, and achieves efficient and accurate tender document summary generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-20
AI Technical Summary
Existing methods for generating bid document summaries rely on manual extraction, which is time-consuming, labor-intensive, and lacks accuracy and comprehensiveness. General language models lack an understanding of the professional knowledge in the bid document field, resulting in inconsistent quality of generated summaries that cannot meet the need for quickly obtaining core information.
By extracting domain terms and constructing semantic chains from tender documents, a domain terminology library for tender documents is formed. The semantic parsing weights of the large language model are adjusted to generate a domain-adapted semantic parsing model. The core expression content of semantic nodes is extracted, and the summary content is optimized based on reader feedback to achieve personalized customization.
It improves the accuracy and coherence of bid document summaries, accurately captures key information, and generates summaries that comprehensively reflect the core content of bid documents. This greatly facilitates readers in quickly obtaining key information and improves the efficiency and quality of bidding and tendering work.
Smart Images

Figure CN121301563B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, in particular to a bid document abstract generation method and system based on a large language model. BACKGROUND
[0002] In the field of bidding, a bid document contains a large amount of complex and professional information, covering project plans, technical specifications, business terms, and other aspects. For the reader, it is a very challenging task to quickly obtain key information from the vast amount of bid document content. The traditional bid document abstract generation method mainly relies on manual extraction, which not only requires a large amount of manpower and time cost, but also due to the differences and subjectivity of human understanding, it is difficult to ensure the accuracy and comprehensiveness of the abstract. With the development of natural language processing technology, some automatic abstract methods based on general language models have begun to be applied to bid document processing, but these general models lack in-depth understanding of the professional knowledge of the bid document field, and cannot accurately grasp the key terms and semantic logical relationships in the bid document, resulting in uneven quality of the generated abstract, which cannot well meet the needs of bid document readers to quickly and accurately obtain core information. SUMMARY
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a bid document abstract generation method based on a large language model, which comprises:
[0004] Extracting the domain terms from the inputted set of bid documents to be processed to form a bid document domain term library, and based on the term association relationship in each bid document unit, splitting each bid document unit into semantic nodes, and connecting the semantic nodes through the term association relationship to form a bid document domain semantic chain;
[0005] Inputting the bid document domain term library into a pre-trained large language model, adjusting the semantic parsing weight of the pre-trained large language model to adapt the pre-trained large language model to the semantic parsing requirements of the bid document domain, and obtaining a domain-adapted semantic parsing model;
[0006] Inputting each semantic node in the bid document domain semantic chain into the domain-adapted semantic parsing model, extracting the core expression content of each semantic node through the domain-adapted semantic parsing model, forming a dynamic abstract segment corresponding to each semantic node, and integrating all dynamic abstract segments to obtain a dynamic abstract segment set;
[0007] According to the connection relationship of the semantic nodes in the bid document domain semantic chain, determining the association order of each dynamic abstract segment in the dynamic abstract segment set, performing semantic connection processing on each dynamic abstract segment according to the association order, and obtaining a preliminary fusion abstract text;
[0008] Receiving semantic feedback information of the tender document reader on the preliminary fusion summary text, adjusting the expression content and the association order of each dynamic summary fragment based on the semantic feedback information, updating the tender document field semantic chain, generating a final tender document summary and outputting to a target receiving end.
[0009] In still another aspect, the present application also provides a tender document summary generation system based on a large language model, comprising a processor, a machine-readable storage medium, the machine-readable storage medium and the processor are connected, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to realize the above-mentioned method.
[0010] Based on the above aspects, the present application extracts field terms from the set of tender documents to be processed and constructs a tender document field term library, simultaneously splits the tender document units based on the term association relationship to form a field semantic chain, inputs the field term library into a pre-trained large language model and adjusts the semantic parsing weight, obtains a semantic parsing model that adapts to the semantic parsing requirements of the tender document field, makes the semantic parsing model able to deeply understand the professional semantics in the tender document, and improves the accuracy of semantic parsing. The semantic parsing model is used to extract the core expression content of the semantic nodes to form a set of dynamic summary fragments, which can accurately capture the key information of each semantic node. The association order of the dynamic summary fragments is determined according to the connection relationship of the semantic nodes and the semantic connection processing is performed, a preliminary fusion summary text is obtained, and the coherence and logic of the summary are ensured. The expression content and the association order of the summary fragments are adjusted according to the semantic feedback information of the reader, and the field semantic chain is updated to generate a final summary, which realizes the individual customization and continuous optimization of the summary. The final generated tender document summary can accurately and comprehensively reflect the core content of the tender document, greatly facilitates the tender document reader to quickly obtain the key information, and improves the efficiency and quality of the tendering and bidding work. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is the execution flow diagram of the tender document summary generation method based on a large language model provided by the embodiment of the present application.
[0012] Figure 2 is a schematic diagram of exemplary hardware and software components of the tender document summary generation system based on a large language model provided by the embodiment of the present application.
[0013] In the figure, 100 is a tender document summary generation system based on a large language model, 110 is a network port, 120 is a processor, 130 is a communication bus, 140 is a storage medium, and 150 is an I / O interface. DETAILED DESCRIPTION
[0014] The application will be specifically described below with reference to the drawings of the specification, Figure 1 is a flowchart of a large language model-based bid file abstract generation method provided by an embodiment of the application. The large language model-based bid file abstract generation method will be described in detail below.
[0015] Step S110: Extracting domain terms from the input set of to-be-processed bid files to form a bid file domain term library, and based on the term association relationship in each bid file unit, splitting each bid file unit into semantic nodes, and connecting the semantic nodes through the term association relationship to form a bid file domain semantic chain.
[0016] In this embodiment, the set of to-be-processed bid files is the "industrial plant steel structure installation project" bid file, which includes four bid file units: project requirement response file, technical implementation scheme file, commercial offer file, and construction organization plan file. Through term extraction, semantic node splitting and association construction, a semantic chain that fits the industrial steel structure installation field is formed.
[0017] Step S111: Scanning each bid file unit text in the set of to-be-processed bid files sentence by sentence, identifying words with bid file domain attributes, including project requirement expression words, technical scheme expression words, commercial clause expression words, and implementation plan expression words, to complete the domain term extraction of the set of to-be-processed bid files.
[0018] Start the text scanning program to parse the four bid file units sentence by sentence. In the project requirement response file, identify project requirement expression words such as "steel structure bearing capacity grade", "column span requirement", and "roof load standard", which directly correspond to the core performance requirements of industrial plant steel structure.
[0019] In the technical implementation scheme file, identify technical scheme expression words such as "steel member welding process", "high-strength bolt connection standard", "steel structure hoisting scheme", and "anticorrosion coating process", which cover the key technical links of steel structure installation.
[0020] In the commercial offer file, identify commercial clause expression words such as "steel member unit price", "installation labor cost", "performance bond ratio", and "payment cycle agreement", which reflect the core business content of the bid.
[0021] In the construction organization plan file, identify implementation plan expression words such as "construction period arrangement", "construction team configuration", "large equipment access plan", and "acceptance node setting", which embody the time and resource planning of construction execution.
[0022] Store all identified words in a temporary term buffer area to ensure that the domain core words of each unit of the bid file are covered.
[0023] Step S112: Count the frequency of the identified words with the bidding document field attribute, retain the words with frequency meeting the bidding document field core term standard, delete the repeated words, and form the bidding document field term library.
[0024] Call the word frequency counting module to count the words in the temporary term buffer. For example, "steel member welding process" appears multiple times in the technical solution document, "high-strength bolt connection standard" is mentioned in multiple technical chapters, "construction period arrangement" is repeatedly emphasized in the construction plan, and "temporary material storage area" appears only once.
[0025] Set the bidding document field core term standard to "appear in at least two bidding document units or appear in a single unit with a frequency not lower than a certain benchmark", and filter the words according to this standard. Delete repeated words and low-frequency words that do not meet the standard, and classify and organize the filtered words into four categories: "project requirements", "technical solutions", "commercial terms", and "construction organization".
[0026] For example, the project requirements category includes "steel structure bearing capacity level" and "column span requirement", the technical solutions category includes "steel member welding process" and "anticorrosion coating process", the commercial terms category includes "steel member unit price" and "payment cycle agreement", and the construction organization category includes "construction period arrangement" and "large equipment access plan", and finally forms a structured bidding document field term library, stored in JSON format.
[0027] Step S113: Use the terms in the bidding document field term library as the splitting marker to divide each bidding document unit into multiple text segments, each containing at least one field term and the corresponding expression content, and each text segment as a semantic node.
[0028] Use the terms in the field term library as anchor points to split the text of each bidding document unit. In the technical implementation scheme document, use "steel member welding process" as the splitting marker to divide the continuous text containing the term and "welding material model", "welding current control", "weld detection standard" and other expressions into one segment; use "high-strength bolt connection standard" as the splitting marker to divide the text containing the term and "bolt pre-tightening force requirement", "connection surface treatment method", "torque detection method" and other expressions into another segment.
[0029] In the construction organization plan document, the text containing the terms such as "construction period arrangement", "foundation construction stage length", "steel member installation stage progress", "roof sealing node time", etc. is divided into a segment with "construction period arrangement" as the split symbol; the text containing the terms such as "large equipment access plan", "tower crane access time", "car hoist operation range", "equipment installation and debugging period", etc. is divided into a segment with "large equipment access plan" as the split symbol.
[0030] Each split text segment contains at least one domain term, and the expression content is complete and independent, and each segment is a semantic node. Assign a unique identifier to each semantic node, and the identification rule is "category abbreviation-sequence number", such as "JS-001" and "JS-002" for the technical scheme class semantic node, and "SG-001" and "SG-002" for the construction organization class.
[0031] Step S114: Compare the domain terms in different semantic nodes, identify the combination of semantic nodes containing the same domain terms, and determine the term association relationship between the combination of semantic nodes containing the same domain terms, including term reference relationship, term subordinate relationship and term supplement relationship.
[0032] The association relationship analysis is completed through the following sub-steps:
[0033] Step S1141: Traverse the text content of each semantic node, and filter out the words belonging to the bid document domain term library from the text content to form a domain term set corresponding to each semantic node.
[0034] Read the text content of each semantic node one by one, and filter out the words belonging to the domain term library from the text content through the term matching algorithm. For example, the text of semantic node "JS-001" is "steel member welding process adopts submerged arc automatic welding, welding material selects low alloy steel welding wire, welding current is controlled within a reasonable range, and weld needs to be detected by ultrasonic", the filtered term is "steel member welding process", and the term set {steel member welding process} is formed; the text of semantic node "JS-002" is "the weld quality grade of steel member welding process needs to reach level two, and the deformation after welding needs to be controlled within the allowable range", and the filtered term is also "steel member welding process", and the term set {steel member welding process} is formed; the text of semantic node "JS-003" is "high-strength bolt connection standard requires bolt pre-tightening force to reach design value, connection surface needs to be sandblasted and rusted, and torque value needs to be rechecked after installation", and the filtered term is "high-strength bolt connection standard", and the term set {high-strength bolt connection standard} is formed.
[0035] Step S1142: Compare the domain term sets of any two semantic nodes one by one, record the semantic node pairs with the same domain terms, and form the combination of semantic nodes containing the same domain terms.
[0036] Compare the term sets of all semantic nodes pairwise: compare the term sets of "JS-001" and "JS-002", find that both contain "steel member welding process", record the semantic node pair (JS-001, JS-002); compare "JS-004" (term set {steel member welding process, weld repair process}) and "JS-001", also contain the same term, record the node pair (JS-004, JS-001).
[0037] Arrange all node pairs containing the same term to form a list of semantic node combinations containing the same field term, and mark the common field term for each combination.
[0038] Step S1143: Analyze the syntactic role of the same field term in the text content of the two semantic nodes in the semantic node combination containing the same field term. If the same field term in one semantic node is a subject expression, and the same field term in the other semantic node is a pointer expression, it is determined that there is a term pointing relationship between the two semantic nodes.
[0039] Take node pair (JS-005, JS-006) as an example, the text of JS-005 is "The main steel column installation adopts a double-machine lifting scheme, and the lifting point of the main steel column is set at 1 / 3 of the column height", and the term set is {main steel column installation}; the text of JS-006 is "Its hoisting process needs to monitor the column verticality throughout the process, and immediately adjust when the deviation exceeds the allowed value", and the term set is {main steel column installation} ("it" refers to "main steel column installation").
[0040] Analysis of syntactic role shows that "main steel column installation" in JS-005 is a subject expression, and "it" in JS-006 is a pointer expression, so it is determined that there is a term pointing relationship between them.
[0041] Step S1144: Analyze the range of expression content corresponding to the same field term in the semantic node combination containing the same field term. If the expression content of the same field term in one semantic node is a whole concept, and the expression content of the same field term in the other semantic node is a partial concept, it is determined that there is a term subordinate relationship between the two semantic nodes.
[0042] Take node pair (JS-003, JS-007) as an example, the text of JS-003 is "High-strength bolt connection standard requires bolt pre-tightening force to reach design value, and the connecting surface needs to be sandblasted for rust removal treatment, and the torque value needs to be reviewed after installation", which expresses the whole requirement of "high-strength bolt connection standard"; the text of JS-007 is "The pre-tightening force control in the high-strength bolt connection standard needs to use a torque wrench, and it needs to be tightened to the design value in three times", which expresses the local requirement of "pre-tightening force control" in "high-strength bolt connection standard".
[0043] Since there is a whole and partial relationship between the expressions, it is determined that there is a term subordination relationship between the two semantic nodes.
[0044] Step S1145: In the combination of semantic nodes containing the same field terms, the completeness of the expression corresponding to the same field terms is analyzed. If the expressions of the same field terms in the two semantic nodes cover different aspects and jointly constitute a complete concept, it is determined that there is a term supplement relationship between the two semantic nodes.
[0045] Taking the node pair (JS-001, JS-004) as an example, the text of JS-001 describes the welding method, material and detection requirements of “steel member welding process”; the text of JS-004 describes the weld quality grade and deformation control requirements of “steel member welding process”.
[0046] Both of them cover different aspects of the welding process, and form a complete description of the welding process after combination, so it is determined that there is a term supplement relationship between the two semantic nodes.
[0047] Step S1146: Add corresponding term association relationship type labels to each semantic node combination containing the same field terms to form a semantic node association relationship list.
[0048] For each semantic node combination, add “reference relationship”, “subordination relationship” or “supplement relationship” labels according to the above analysis results. For example, the node pair (JS-001, JS-002) adds a “supplement relationship” label, the node pair (JS-003, JS-007) adds a “subordination relationship” label, and the node pair (JS-005, JS-006) adds a “reference relationship” label.
[0049] Arrange all the labeled node pairs into a semantic node association relationship list, which contains three fields of node identification pairs, common terms and association types.
[0050] Step S115: Arrange the semantic nodes as chain nodes, and arrange the term association relationships between the semantic nodes as chain connection edges, according to the business expression order of the bid document unit, to form an initial bid document field semantic chain.
[0051] Based on the business logic order of the bid document unit (project requirements-technical scheme-commercial offer-construction organization), arrange the semantic nodes as chain nodes in turn. For example, arrange the project requirement type semantic nodes (such as XQ-001, XQ-002) first, then arrange the technical scheme type semantic nodes (such as JS-001, JS-002, JS-003), then arrange the commercial offer type nodes, and finally arrange the construction organization type nodes.
[0052] According to the semantic node association relationship list, chain connection edges are added between nodes with an association relationship, and the connection edges are marked with the association relationship type. For example, a connection edge marked with a "complementary relationship" is added between JS-001 and JS-002, and a connection edge marked with a "subordinate relationship" is added between JS-003 and JS-007, forming an initial tender document field semantic chain, which is stored in a graph structure data format.
[0053] Step S116: Check whether there are isolated chain nodes without chain connection edges in the initial tender document field semantic chain. If there are, analyze the potential association of the field terms in the isolated chain nodes with the field terms in other chain nodes, add chain connection edges based on the potential association, and form a complete tender document field semantic chain.
[0054] All chain nodes in the initial semantic chain are traversed to check whether there are isolated nodes that are not connected to any other node. For example, the construction organization semantic node "SG-005" (term set {construction safety protection measures}) is not connected to other nodes, and is determined to be an isolated node.
[0055] The potential association of the field term "construction safety protection measures" of the isolated node with the terms of other nodes is analyzed: it is found that the description content of the technical scheme class node "JS-008" (term set {high-altitude welding operation scheme}) involves high-altitude operation safety, and there is a potential "complementary relationship" (construction safety protection measures complement the safety requirements of high-altitude welding operation).
[0056] Based on the potential association, a chain connection edge marked with a "complementary relationship" is added between SG-005 and JS-008. The above process is repeated until all nodes have at least one connection edge, forming a complete tender document field semantic chain without isolated nodes.
[0057] Step S120: Input the tender document field term library into a pre-trained large language model, adjust the semantic parsing weight of the pre-trained large language model, adapt the pre-trained large language model to the semantic parsing requirements of the tender document field, and obtain a field-adapted semantic parsing model.
[0058] A general pre-trained large language model is selected as the base model, which is adapted to the semantic parsing requirements of the industrial plant steel structure installation tender document through field term injection and weight adjustment. The specific process is as follows.
[0059] Step S121: Select core terms from the tender document field term library, and match corresponding tender document field standard expression texts for each core term to form a field-adapted corpus pair composed of core terms and standard expression texts.
[0060] Core terms are selected from the tender document field terminology database, with the screening criteria being "terms that cover the four categories and are core expressions in tender documents", such as "steel member welding process", "high-strength bolt connection standard", "construction period arrangement", "steel member unit price", etc.
[0061] A standard expression text is matched for each core term, and the expression text needs to conform to the standard expression habits of industrial steel structure installation bidding. For example, the standard expression text matched for "steel member welding process" is "The steel member welding process should be determined according to the material quality and stress requirements of the member, and automatic submerged arc welding or gas shielded welding should be preferred. The welding material should match the base material, the weld quality grade should meet the design requirements, and non-destructive testing should be performed after welding"; the standard expression text matched for "high-strength bolt connection standard" is "High-strength bolt connection should comply with relevant industry standards, the connection surface treatment should meet the design roughness requirements, the bolt pre-tightening force should be applied through a torque wrench according to the specified procedure, and torque rechecking should be performed after installation.
[0062] Each core term is combined with its matched standard expression text to form a field-adapted corpus pair, each corpus pair containing two fields of "core term-standard expression".
[0063] Step S122: Convert the field-adapted corpus pair according to the input format requirements of the pre-trained large language model, and batch input to the corpus input module of the pre-trained large language model.
[0064] The input format requirements of the pre-trained large language model are "instruction-content" structure, so the field-adapted corpus pair is converted to this structure: the instruction field is unified as "analyze tender document field term: [core term]", and the content field is the corresponding standard expression text. For example, a certain corpus pair is converted to "instruction: analyze tender document field term: steel member welding process; content: steel member welding process should be determined according to the material quality and stress requirements of the member……".
[0065] Group all converted corpus pairs according to the batch size required by the model, and batch input to the corpus input module through the model's corpus input interface, ensuring that the input data format conforms to the model's parsing requirements.
[0066] Step S123: Call the parameter viewing interface of the pre-trained large language model to obtain the weight parameter set related to term recognition and semantic understanding in the semantic parsing layer.
[0067] Through the API interface provided by the pre-trained large language model, the parameter viewing function is called to locate the neuron cluster responsible for term recognition and semantic understanding in the semantic parsing layer of the model. The weight parameters corresponding to these neurons are extracted to form a weight parameter set, which contains parameter identification, current parameter value, and corresponding semantic parsing function description, etc.
[0068] For example, extract the neuron weight parameters responsible for "process terminology recognition", the neuron weight parameters responsible for "connection relationship semantic understanding", etc., to ensure that the obtained parameters cover the core links of terminology recognition and semantic understanding.
[0069] Step S124: Based on the core terminology in the field adaptation corpus pair, increase the weight parameter value of the semantic analysis layer for identifying the field terminology of the bidding document, and reduce the weight parameter value of the semantic analysis layer for identifying the general field terminology, so that the semantic analysis layer preferentially identifies the field terminology of the bidding document.
[0070] For the core terminology type in the field adaptation corpus pair, locate the terminology recognition weight parameter of the corresponding category in the semantic analysis layer. For example, for technical terms such as "steel member welding process" and "high-strength bolt connection standard", find the weight parameter responsible for technical terminology recognition; for commercial terms such as "steel member unit price" and "payment cycle agreement", find the weight parameter responsible for commercial terminology recognition.
[0071] Increase the value of these weight parameters, while reducing the value of the weight parameters responsible for the general field terminology (such as "ordinary bolt", "general welding", and other non-bidding core terminology) recognition. For example, adjust the technical terminology recognition parameter value to be a certain percentage higher than the general terminology recognition parameter value, to ensure that the model preferentially captures the core terminology of the bidding document when analyzing the text.
[0072] Step S125: Encapsulate the adjusted semantic analysis weight parameters into an independent field adaptation layer module, which includes a terminology weight adjustment submodule, a field semantic correction submodule, and a expression specification submodule.
[0073] The field adaptation layer module is constructed through the following sub-steps:
[0074] Step S1251: Write the adjusted identification weight parameters of the bidding document field terminology into the program of the terminology weight adjustment submodule, so that when the terminology weight adjustment submodule receives input text, it can strengthen the identification strength of the bidding document field terminology based on the identification weight parameters, and output text containing field terminology markers.
[0075] The program of the terminology weight adjustment submodule is written in Python, and the adjusted weight parameters are imported into the program in the form of a configuration file. After the program receives input text, it performs word segmentation processing on the text through a word segmentation algorithm, and then performs terminology recognition scoring on each word based on the imported weight parameters. The word segmentation that scores higher than the set threshold is marked as the bidding document field terminology, and is annotated in the output text through special markers (such as square brackets). For example, the input text "steel member welding adopts submerged arc automatic welding" is output as "[steel member welding] adopts submerged arc automatic welding".
[0076] Step S1252: import the bidding document field semantic rule library, which contains common semantic expression logic rules in the bidding document field, so that the field semantic correction submodule can compare the semantic expression of the input text with the rules in the bidding document field semantic rule library and correct the expression content that does not conform to the field semantic rules.
[0077] The bidding document field semantic rule library contains expression logic rules of industrial steel structure installation bidding documents, such as “the technical scheme expression should first specify the process type, then specify the parameter requirement, and finally specify the detection standard” “the business clause expression should first specify the cost item, then specify the calculation method, and finally specify the payment method” “the construction plan expression should first specify the stage division, then specify the duration of each stage, and finally specify the node acceptance requirement” and the like.
[0078] The semantic rule library is imported into the field semantic correction submodule in XML format, and the submodule program compares the expression logic of the input text with the corresponding rules in the rule library. For example, the input text is “welding seam needs to be detected by ultrasonic wave, steel member welding process adopts submerged arc automatic welding, and welding current is controlled within a reasonable range”. It is found that the expression order does not conform to the rule of “first specifying the process type, then specifying the parameter requirement, and finally specifying the detection standard”. The submodule will automatically adjust the expression order to “[steel member welding] adopts submerged arc automatic welding, welding current is controlled within a reasonable range, and welding seam needs to be detected by ultrasonic wave”, so as to ensure compliance with the semantic expression specification of the bidding document.
[0079] Step S1253: store the bidding document field standard expression template, which covers the specification expression format of project requirements, technical scheme and business clause, so that the expression specification submodule can adjust the corrected semantic expression to the format conforming to the bidding document field standard expression template.
[0080] The bidding document field standard expression template is classified according to categories. The project requirement category template contains the format of “[requirement term]: meet [specific standard], corresponding [performance index] requirement”. The technical scheme category template contains the format of “[work art term]: adopt [process method], [parameter requirement], [detection standard]”. The business clause category template contains the format of “[cost term]: [unit price standard], [calculation basis], [payment method]”.
[0081] The above template is stored in the template library of the expression standardization submodule. After the submodule receives the text output by the field semantic correction submodule, the corresponding standard template is matched according to the field term category contained in the text, and the expression format is adjusted. For example, the corrected text is "[steel member welding] adopts submerged arc automatic welding, and the welding current is controlled in a reasonable range. The weld needs to be detected by ultrasonic wave". After matching the technical scheme class template, it is adjusted to "[steel member welding process] : adopt submerged arc automatic welding, control the welding current in a reasonable range, and the weld detection standard is ultrasonic wave detection". The expression is more in line with the standard format of the bidding document.
[0082] Step S1254: After the field adaptation layer module receives the input text, the term weight adjustment submodule is first called for processing, and the processing result is input into the field semantic correction submodule. Finally, the correction result is input into the expression standardization submodule.
[0083] The flow control logic is set in the main program of the field adaptation layer module to determine the calling order of the submodules: when the module receives external input text, the processing flow of the term weight adjustment submodule is first triggered to complete the field term marking; the marked text is automatically transmitted to the field semantic correction submodule for expression logic correction; the corrected text is transmitted to the expression standardization submodule for format adjustment.
[0084] This flow is realized through the function call relationship in the program. For example, the main program calls the term weight adjustment submodule through the "term_weight_adjust(input_text)" function, and returns the result as a parameter to the "semantic_correction(marked_text)" function. The semantic correction result is transmitted to the "format_standardization(corrected_text)" function to ensure that the submodules work in sequence.
[0085] Step S1255: After the program codes of the term weight adjustment submodule, the field semantic correction submodule and the expression standardization submodule are integrated, the bidding document text segment containing non-standard expressions is input into the field adaptation layer module. Check whether the output text processed by the term weight adjustment submodule, the field semantic correction submodule and the expression standardization submodule in turn meets the requirements of field term marking, semantic rules and expression standardization. If it meets the requirements, the field adaptation layer module is completed, and the test of the collaborative effect of the submodules is completed.
[0086] The Python codes of the three sub-modules are integrated into a complete domain adaptation layer module program to generate an executable file. A text segment of the bidding file containing non-standard expressions is selected as a test case, for example, the test text is "welding steel components with submerged arc welding, the current should be appropriate, and the welds should be checked". The text has problems such as unclear terminology, disordered expression order and non-standard format.
[0087] The test text is input into the domain adaptation layer module: the term weight adjustment submodule outputs "[welding steel components] with submerged arc welding, the current should be appropriate, and the welds should be checked"; the domain semantic correction submodule adjusts the expression order to "[welding steel components] adopts submerged arc welding, the welding current is controlled within a reasonable range, and the welds need to be detected"; and the expression specification submodule outputs "[steel component welding process]: adopts submerged arc welding, the welding current is controlled within a reasonable range, and the weld detection standard needs to meet the design requirements" after matching the template.
[0088] The output text is checked: the term label is accurate, the expression logic conforms to the semantic rules, the format conforms to the standard template, and the coordination effect of the judgment submodule meets the standard, completing the construction of the domain adaptation layer module.
[0089] Step S126: The domain adaptation layer module is embedded between the input layer and the semantic analysis layer of the pre-trained large language model, so that the pre-trained large language model processes the input text through the domain adaptation layer module first when receiving the input text.
[0090] The constructed domain adaptation layer module is embedded into the model architecture through the plugin interface of the pre-trained large language model, and the embedding position of the module is between the input layer and the semantic analysis layer. When the model receives external input text, the input layer first transmits the text to the domain adaptation layer module, and then transmits the processed text to the semantic analysis layer for subsequent semantic understanding and feature extraction after processing the text through term labeling, semantic correction and format specification.
[0091] The calling flow code of the model is modified, and the calling instruction of the domain adaptation layer module is added in the text output function of the input layer to ensure that the text flow path is "input layer—domain adaptation layer module—semantic analysis layer", realizing the seamless integration of the domain adaptation function and the original model.
[0092] Step S127: Select a test bidding file text and input it into the pre-trained large language model embedded with the domain adaptation layer module, and compare the consistency of the output semantic analysis result with the standard analysis result of the bidding file domain. If the consistency meets the preset adaptation standard, a semantic analysis model after domain adaptation is determined.
[0093] The historical bidding file text of the steel structure installation project of the industrial plant is selected as the test text, which contains the contents of four parts of project requirements, technical scheme, commercial terms and construction organization. The test text is input into the pre-trained large language model of the embedded field adaptation layer module, and the semantic analysis result output by the pre-trained large language model is obtained, which includes field term recognition result, semantic relationship analysis result and core content extraction result.
[0094] At the same time, invite bidding field experts to manually analyze the same test text and generate standard analysis results of the bidding file field. Compare the consistency of the model analysis results and the standard analysis results, evaluate from three dimensions of term recognition accuracy, semantic relationship judgment accuracy and core content extraction completeness.
[0095] The preset adaptation standard is set as "the accuracy of the three dimensions is not less than a certain proportion", if the term recognition accuracy, semantic relationship judgment accuracy and core content extraction completeness of the model analysis result all meet the standard, it is determined that the model adaptation meets the standard, and the field adapted semantic analysis model is determined; if it does not meet the standard, return to step S124 to adjust the semantic analysis weight parameter, repeat the adaptation and test process until the preset standard is met.
[0096] Step S130: input each semantic node in the bidding file field semantic chain into the field adapted semantic analysis model, extract the core expression content of each semantic node through the field adapted semantic analysis model, form a dynamic abstract segment corresponding to each semantic node, and integrate all dynamic abstract segments to obtain a dynamic abstract segment set.
[0097] With the complete bidding file field semantic chain as input, each semantic node is processed one by one by the field adapted semantic analysis model to generate dynamic abstract segments and integrate them. The specific process is as follows.
[0098] Step S131: according to the arrangement order of the chain nodes in the bidding file field semantic chain, read each chain node corresponding semantic node in turn to form a semantic node sequence, and complete the extraction of the semantic node sequence of the bidding file field semantic chain.
[0099] Call the semantic chain analysis program to read the graph structure data of the bidding file field semantic chain, extract the semantic node corresponding to each chain node according to the arrangement order of the chain nodes (project requirements-technical scheme-commercial offer-construction organization), and record the identification, text content and belonging category of each semantic node.
[0100] The extracted semantic nodes are arranged in sequence to form a semantic node sequence, for example, the sequence is "XQ-001, XQ-002, JS-001, JS-002, SW-001, SW-002, SG-001, SG-002,...", and the position of each node in the sequence is consistent with the arrangement order in the semantic chain, ensuring that the processing order conforms to the business logic of the bidding document.
[0101] Step S132: The text content of each semantic node in the semantic node sequence is sequentially input to the text input port of the domain-adapted semantic parsing model, completing the individual input of the semantic nodes to the domain-adapted semantic parsing model.
[0102] Through the text input interface of the domain-adapted semantic parsing model, the text content of each semantic node is input to the model in the order of the semantic node sequence. During the input process, a unique input identifier is added to each semantic node, which is consistent with the original identifier of the semantic node (such as "JS-001" and "SG-001"), ensuring that the model output result can be associated with the corresponding semantic node.
[0103] For example, the text content of XQ-001 is first input, which is "Steel structure bearing capacity level needs to meet Class A requirements, column span is designed according to 12 meters, and roof load standard is 0.7 kN / m²", and the input identifier is "XQ-001"; then the text content of XQ-002 is input, which is "Steel structure seismic grade needs to reach Class VIII, and fire resistance grade needs to meet Class I standard", and the input identifier is "XQ-002", and the input of all semantic nodes is completed in turn.
[0104] Step S133: Enable the core information extraction mode in the domain-adapted semantic parsing model, which is used to filter out sentences reflecting the core meaning of the semantic nodes from the input semantic node text content based on the domain semantic rules of the bidding document.
[0105] The core information extraction mode is enabled and the core sentences are filtered out through the following sub-steps:
[0106] Step S1331: In the function selection interface of the domain-adapted semantic parsing model, check the core information extraction option, and set the running parameters of the core information extraction mode, including the number of core sentence filtering and the retention proportion of core terms.
[0107] Log in to the management interface of the domain-adapted semantic parsing model, find the "core information extraction" option in the function module list and check it, enter the parameter setting page. Set the core sentence screening quantity to "screen 1-3 core sentences per semantic node", and the specific number is adjusted according to the length of the semantic node text; set the core term retention proportion to "100%", that is, ensure that all domain terms are retained in the core sentence.
[0108] Save the parameter settings, and the model automatically enables the core information extraction mode. The subsequent input semantic node text will be processed according to this mode.
[0109] Step S1332: From the built-in rule library of the domain-adapted semantic parsing model, call the target rules related to the core information identification of the bid document field, which include the core term sentence priority rule and the business logic keyword guided sentence priority rule.
[0110] After the model enables the core information extraction mode, it automatically calls the core information identification rules of the bid document field from the built-in rule library: the core term sentence priority rule means that "the sentence containing the core term in the domain term library is preferred to be judged as a core sentence"; the business logic keyword guided sentence priority rule means "the sentence containing the business logic keywords such as 'need to meet', 'adopt','require', 'agree', 'arrange' is preferred to be judged as a core sentence".
[0111] For example, the text of semantic node JS-001 contains four sentences: "[steel member welding process] adopts submerged arc automatic welding", "welding material selects low alloy steel welding wire", "welding current is controlled within a reasonable range", and "welding seam needs to be detected by ultrasonic". The first and fourth sentences contain the core term "steel member welding process" and the logic keywords "adopt" and "need to perform", which will be preferentially screened.
[0112] Step S1333: Divide the input semantic node text content into independent sentence units according to punctuation marks, assign a unique sentence identifier to each sentence unit, and compare each sentence unit with the loaded bid document field semantic rules to count the number of rules that each sentence unit meets.
[0113] After the model receives the semantic node text, it divides the text into independent sentence units according to punctuation marks such as periods, question marks, and exclamation marks. Each sentence unit is assigned a unique identifier in the format of "semantic node identifier-sentence number", such as JS-001-01, JS-001-02, etc.
[0114] Compare each sentence unit with the sentence priority rule of the core term and the sentence priority rule guided by the business logic key word, and count the number of rules that each sentence unit meets. For example, JS-001-01 “[steel member welding process] adopts submerged arc automatic welding” meets 2 rules, JS-001-04 “welds need to be ultrasonic tested” meets 1 rule, JS-001-02 “welding material selects low alloy steel welding wire” meets 1 rule, and JS-001-03 “welding current is controlled within a reasonable range” meets 0 rules.
[0115] Step S1334: Sort the number of rules met by each sentence unit, and select the sentence unit with the most rules met as the candidate core sentence.
[0116] Sort the number of rules met by each sentence unit from most to least, and preferentially select the sentence unit with the most rules met as the candidate core sentence. For example, the sentence units of JS-001 are sorted as JS-001-01 (2), JS-001-02 (1), JS-001-04 (1), and JS-001-03 (0), and JS-001-01 is selected as the primary candidate core sentence.
[0117] Step S1335: Check whether the candidate core sentence contains the core field term of the semantic node and whether it can independently express a complete semantic meaning. If it meets the requirements, it is determined as the sentence that embodies the core meaning of the semantic node.
[0118] Check JS-001-01 “[steel member welding process] adopts submerged arc automatic welding”: It contains the core field term “steel member welding process” and can independently express the complete semantic “steel member welding process adopted”. Therefore, it is determined as the core sentence.
[0119] If the candidate core sentence does not contain the core term or the semantic is incomplete, the next sentence unit is selected from the sorted sentence units for checking in turn until a core sentence that meets the requirements is found. If a single sentence cannot completely express the semantic, multiple sentences are selected to form a core sentence set to ensure the completeness of the core meaning.
[0120] Step S134: The selected core meaning sentence is simplified in expression by the semantic analysis model after field adaptation, forming a dynamic abstract segment corresponding to the semantic node.
[0121] The model simplifies the selected core sentence in expression: delete redundant modifiers (such as adverbs, adjectives, etc.), retain core terms and key action expressions; simplify complex sentence patterns into simple sentence patterns to ensure concise and clear expression; unify the expression style to make the abstract segment conform to the formal expression habits of the bid document.
[0122] For example, the core sentence "[steel member welding process] adopts submerged arc automatic welding, and low alloy steel welding wire is selected for welding material, and ultrasonic testing is required for the weld" is simplified to "steel member welding process adopts submerged arc automatic welding, selects low alloy steel welding wire, and weld requires ultrasonic testing", forming a dynamic abstract segment corresponding to JS-001.
[0123] If there are multiple core sentences, they are combined into a coherent abstract segment in logical order after simplification, ensuring semantic coherence and no repetitive expression.
[0124] Step S135: Add corresponding semantic node identifier, position identifier in the bidding document field semantic chain, and field term list to each dynamic abstract segment. Integrate all generated dynamic abstract segments in the order of their corresponding semantic nodes in the semantic node sequence, associate the corresponding association information, and form a dynamic abstract segment collection.
[0125] Add metadata information to each dynamic abstract segment: the semantic node identifier is consistent with the corresponding semantic node (such as "JS-001"); the position identifier is the serial number of the semantic node in the semantic node sequence (such as "3", indicating the third node in the sequence); the field term list is all field terms contained in the semantic node (such as "{steel member welding process}").
[0126] Arrange all dynamic abstract segments in the order of the semantic node sequence, for example, integrate them in the order of "XQ-001 abstract, XQ-002 abstract, JS-001 abstract, JS-002 abstract……", and each abstract segment is associated with its metadata information.
[0127] Store the integrated abstract segments and metadata information as a dynamic abstract segment collection in JSON format. Each entry in the collection contains four fields: "abstract text", "semantic node identifier", "position identifier", and "field term list", which facilitates subsequent association order determination and semantic connection processing.
[0128] Step S140: According to the connection relationship of the semantic nodes in the bidding document field semantic chain, determine the association order of each dynamic abstract segment in the dynamic abstract segment collection, and perform semantic connection processing on each dynamic abstract segment according to the association order to obtain a preliminary fusion abstract text.
[0129] Based on the node connection relationship of the bidding document field semantic chain, the association logic of the abstract segment is sorted, and a coherent preliminary fusion abstract is formed through connection processing. The specific process is as follows.
[0130] Step S141: Extract the chain connection edge information between each semantic node from the structure data of the bidding document field semantic chain to form a connection relationship table containing semantic node identifier, associated node identifier, and association relationship type.
[0131] Parse the graph structure data of the field semantic chain of the bid document, and extract the information of all chain connection edges: each connection edge contains the starting semantic node identifier, the ending semantic node identifier, and the association relationship type (reference relationship, subordinate relationship, supplementary relationship).
[0132] Organize the above information into a connection relationship table, where each record corresponds to a connection edge, for example, record 1 is "starting point: JS-001, ending point: JS-002, association type: supplementary relationship"; record 2 is "starting point: JS-003, ending point: JS-007, association type: subordinate relationship"; record 3 is "starting point: JS-005, ending point: JS-006, association type: reference relationship", etc., fully reflecting the association logic between semantic nodes.
[0133] Step S142: Associate each dynamic summary fragment with the corresponding semantic node entry in the connection relationship table by the semantic node identifier of the dynamic summary fragment, determine the association fragment identifier and the association relationship type of each dynamic summary fragment.
[0134] Traverse the dynamic summary fragment set, extract the semantic node identifier of each summary fragment, find the record with the identifier as the starting point in the connection relationship table, and get the corresponding ending semantic node identifier (i.e. association fragment identifier) and association relationship type.
[0135] For example, the semantic node identifier of summary fragment JS-001 is "JS-001", the corresponding record in the connection relationship table is found, the ending identifier is "JS-002", and the association type is "supplementary relationship", so the association fragment identifier of JS-001 summary is determined as "JS-002", and the association relationship type is "supplementary relationship"; the association fragment identifier of summary fragment JS-003 is "JS-007", and the association relationship type is "subordinate relationship".
[0136] For the records in the connection relationship table with a certain identifier as the ending point, add reverse association information to the corresponding summary fragment to ensure the bidirectional nature of the association relationship.
[0137] Step S143: Take the dynamic summary fragment corresponding to the starting semantic node of the bid document field semantic chain as the starting point, and arrange the subsequent dynamic summary fragments in the order of the association nodes in the connection relationship table to form the association order of the dynamic summary fragments.
[0138] The starting semantic node of the bid document field semantic chain is usually the first node of the project requirement type (such as XQ-001), and its corresponding dynamic summary fragment is the starting point of the association order.
[0139] According to the connection relationship table, the associated node identifier (such as XQ-002) with XQ-001 as the starting point is found, the summary fragments corresponding to XQ-002 are arranged after the summary of XQ-001; the associated node identifier (such as JS-001) with XQ-002 as the starting point is found, the summary fragments corresponding to JS-001 are arranged after the summary of XQ-002; the associated node identifier (such as JS-002) with JS-001 as the starting point is found, the summary of JS-002 is arranged after the summary of JS-001, and so on.
[0140] If there are multiple associated nodes (such as JS-003 associated with JS-007 and SW-001) for a semantic node, the arrangement order is determined according to the principles of “technical logic is prior to business logic” and “main process is prior to auxiliary process”, for example, the summaries of the technical associated nodes JS-007 are arranged first, and then the summaries of the business associated nodes SW-001 are arranged.
[0141] In the above manner, a complete dynamic summary fragment association order is formed, the arrangement logic of each fragment in the order is consistent with the node connection relationship of the semantic chain in the field of the bid document, and it is ensured that the association of the summary fragments conforms to the business expression logic of the industrial steel structure installation bid document.
[0142] Step S144: According to the association relationship type of adjacent dynamic summary fragments, a bid document field connection sentence template library is called, the bid document field connection sentence template library includes connection sentence templates corresponding to different association relationship types, and a connection sentence connecting adjacent dynamic summary fragments is generated based on the connection sentence template.
[0143] The connection sentence is generated through the following sub-steps:
[0144] Step S1441: According to the default term association relationship type in the semantic chain in the field of the bid document, the connection sentence template is divided into a reference relationship connection sentence template, a subordinate relationship connection sentence template and a supplementary relationship connection sentence template, and is respectively stored in different directories of the bid document field connection sentence template library.
[0145] The bid document field connection sentence template library stores templates according to the association relationship type: the reference relationship connection sentence template directory includes templates such as “for [term], its corresponding [expression content] is …” and “the specific execution requirement of the above [term] is …”; the subordinate relationship connection sentence template directory includes templates such as “[overall term] covers [partial term] and needs to meet …” and “under the framework of [overall term], the execution standard of [partial term] is …”; and the supplementary relationship connection sentence template directory includes templates such as “[term] needs to meet the above requirements in addition to …” and “for [term], the supplementary description is as follows …”.
[0146] Each template is provided with a term placeholder (such as "[term]" and "[overall term]"), so as to facilitate subsequent substitution of specific terms to generate a cohesive sentence.
[0147] Step S1442: The type of the association relationship between the adjacent dynamic summary segments is read. If the type is a term reference relationship, a reference relationship cohesive sentence template is called. If the type is a term subordinate relationship, a subordinate relationship cohesive sentence template is called. If the type is a term supplement relationship, a supplement relationship cohesive sentence template is called.
[0148] The type of the association relationship between the adjacent dynamic summary segments is extracted from the connection relationship table: for example, the association relationship between the JS-005 summary and the JS-006 summary is a reference relationship, so a reference relationship cohesive sentence template is called; the association relationship between the JS-003 summary and the JS-007 summary is a subordinate relationship, so a subordinate relationship cohesive sentence template is called; the association relationship between the JS-001 summary and the JS-002 summary is a supplement relationship, so a supplement relationship cohesive sentence template is called.
[0149] Step S1443: The core field terms contained in each of the adjacent two dynamic summary segments are obtained from the association information of the two dynamic summary segments, and the core field terms common to the two dynamic summary segments are determined.
[0150] The metadata information of the adjacent summary segments is analyzed, and the field term list therein is extracted: for example, the field term list of the JS-005 summary is {main steel column installation}, the field term list of the JS-006 summary is {main steel column installation}, and the core field term common to the two is determined to be "main steel column installation"; the field term list of the JS-003 summary is {high-strength bolt connection standard}, the field term list of the JS-007 summary is {high-strength bolt connection standard}, and the core field term common to the two is "high-strength bolt connection standard".
[0151] Step S1444: The core field term common to the two is substituted into the term placeholder in the called cohesive sentence template, and the general expression in the cohesive sentence template is replaced with a specific expression related to the core term, to generate a preliminary cohesive sentence.
[0152] "Main steel column installation" is substituted into the reference relationship cohesive sentence template "for [term], its corresponding [expression content] is …", to generate the preliminary cohesive sentence "for main steel column installation, its corresponding hoisting process needs …"; "High-strength bolt connection standard" is substituted into the subordinate relationship cohesive sentence template "[overall term] covers [partial term] needs to meet …", combined with the partial expression "pre-tightening force control" of the JS-007 summary, to generate the preliminary cohesive sentence "high-strength bolt connection standard covers pre-tightening force control needs to meet …"; "Steel member welding process" is substituted into the supplement relationship cohesive sentence template "[term] needs to meet the above requirements in addition to …", to generate the preliminary cohesive sentence "steel member welding process needs to meet the above requirements in addition to …".
[0153] Step S1445: According to the overall expression style of the preliminary fused summary text, adjust the language style of the preliminary linking sentence to be consistent with the overall style, check whether the linking sentence accurately reflects the correlation between adjacent dynamic summary segments, and whether there is semantic ambiguity, if there is, modify the expression content until the semantics is accurate.
[0154] The overall expression style of the preliminary fused summary text is formal and concise bidding document expression style, so the preliminary linking sentence "For the installation of main steel columns, the corresponding hoisting process needs to be…" is adjusted to "For the installation of main steel columns, the hoisting process needs to be…" to make it more consistent with the formal expression habit; check "The pre-tightening force control covered by the high-strength bolt connection standard needs to meet…" to confirm that the sentence accurately reflects the subordinate relationship between the whole and the part, and there is no semantic ambiguity; if a linking sentence such as "Steel member welding process in addition to the above requirements, should also…" has the problem of ambiguous reference, modify it to "Steel member welding process in addition to the above requirements, should also…" to make the reference clear and ensure the semantic accuracy.
[0155] Step S145: According to the correlation order, insert the corresponding semantic linking sentence between the adjacent two dynamic summary segments to form a continuous text sequence, analyze the length distribution of the sentences in the text sequence, split the excessively long sentences into short sentences that conform to the reading habit, adjust the order of the sentences, format the adjusted text sequence, remove the redundant linking sentences, and obtain the preliminary fused summary text.
[0156] According to the correlation order of the dynamic summary segments, the summary segments and the corresponding linking sentences are spliced in sequence: for example, XQ-001 summary + linking sentence 1 + XQ-002 summary + linking sentence 2 + JS-001 summary + linking sentence 3 + JS-002 summary…, to form a continuous text sequence.
[0157] Analyze the length of the sentences in the text sequence, if a sentence contains multiple clauses and the total length exceeds a preset threshold (such as more than 30 words), split it into short sentences, for example, "Steel member welding process adopts submerged arc automatic welding, selects low alloy steel welding wire, and welds need to be detected by ultrasonic and the quality level needs to reach level two" is split into "Steel member welding process adopts submerged arc automatic welding, selects low alloy steel welding wire. Welds need to be detected by ultrasonic and the quality level needs to reach level two".
[0158] Adjust the order of the sentences to make the expression more consistent with the logic, for example, "Welds need to be detected by ultrasonic, steel member welding process adopts submerged arc automatic welding" is adjusted to "Steel member welding process adopts submerged arc automatic welding, welds need to be detected by ultrasonic".
[0159] Format the text sequence: unify the use of punctuation, delete repetitive conjunctions (such as two "in addition" conjunctions appearing in succession, keep one and adjust the expression), ensure that the text format is unified and the expression is concise and coherent, and finally form a preliminary fusion abstract text.
[0160] Step S150: Receive semantic feedback information of the tender document reader on the preliminary fusion abstract text, adjust the expression content and association order of each dynamic abstract segment based on the semantic feedback information, update the tender document domain semantic chain, generate the final tender document abstract and output to the target receiving end.
[0161] For example, by optimizing the abstract text through reader feedback, ensure that the abstract meets the actual use requirements, the specific process is as follows.
[0162] Step S151: Add feedback input control in the display interface of the preliminary fusion abstract text, the feedback input control supports the tender document reader to input semantic feedback information in the form of text, and the semantic feedback information includes semantic missing feedback, expression redundancy feedback and order confusion feedback.
[0163] In the display platform (such as web page, client) of the preliminary fusion abstract text, add feedback input control for each text area corresponding to the dynamic abstract segment, including three single selection buttons of "semantic missing", "expression redundancy" and "order confusion" and a text input box.
[0164] If the reader finds that a certain abstract segment misses core information, he can check "semantic missing" and input the specific content missing in the text box (such as "welding process of steel member does not mention welding current requirement"); if he finds that the expression is repeated or redundant, he can check "expression redundancy" and input the redundant content (such as "welding seam detection" is mentioned twice); if he finds that the order of the segment does not conform to the logic, he can check "order confusion" and input the adjustment suggestion (such as "commercial offer segment should be after technical solution segment").
[0165] Step S152: Obtain the semantic feedback information input by the reader through the feedback receiving interface, extract the keywords of the feedback information, and determine the dynamic abstract segment identifier and feedback type corresponding to the feedback information.
[0166] The feedback receiving interface receives the feedback information submitted by the reader in real time, and stores the feedback information into the feedback database. The keyword extraction algorithm is called to process the feedback text: for example, the feedback information "JS-001 abstract does not mention the welding current requirement of steel member welding process", the keywords "JS-001", "steel member welding process" and "welding current requirement" are extracted, it is determined that the corresponding dynamic abstract segment identifier is "JS-001", and the feedback type is "semantic missing"; the feedback information "XQ-002 and JS-001 are in reverse order", the keywords "XQ-002", "JS-001" and "order reversal" are extracted, and it is determined that the corresponding segment identifiers are "XQ-002" and "JS-001", and the feedback type is "order confusion".
[0167] Step S153: For semantic missing feedback, the missing core terms and expression contents are supplemented in the corresponding dynamic abstract segment; for expression redundancy feedback, the redundant modification components in the corresponding dynamic abstract segment are deleted.
[0168] For the semantic missing feedback "does not mention the welding current requirement" identified as JS-001, the text content of the original semantic node JS-001 "steel member welding process adopts submerged arc automatic welding, and low alloy steel welding wire is selected for welding material, and the weld needs to be detected by ultrasonic wave" is called, and "welding current control in a reasonable range" is supplemented in the corresponding dynamic abstract segment, and the supplemented abstract segment is "steel member welding process adopts submerged arc automatic welding, selects low alloy steel welding wire, controls welding current in a reasonable range, and weld needs ultrasonic detection".
[0169] For the expression redundancy feedback "SW-001 abstract repeatedly mentions 'performance guarantee letter'", the SW-001 abstract "the performance guarantee letter ratio is a certain ratio, and the performance guarantee letter needs to be submitted after signing" is called, the repeated "performance guarantee letter" is deleted, and it is modified to "the performance guarantee letter ratio is a certain ratio, and needs to be submitted after signing", and the redundant components are removed.
[0170] Step S154: For order confusion feedback, the position of the corresponding dynamic abstract segment in the text sequence is rearranged by referring to the connection relationship table of the semantic chain in the bidding document field, and the semantic connection sentence of the adjacent dynamic abstract segment is updated.
[0171] For the feedback "XQ-002 and JS-001 are in reverse order", the associated node of XQ-002 in the connection relationship table is JS-001, and the business logic should be "project demand-technical scheme", so the JS-001 abstract is adjusted to be after the XQ-002 abstract.
[0172] Reproduce the linking sentence of the adjacent section: the linking sentence between the original XQ-002 and JS-001 is "Based on the above project requirements, the corresponding technical solution is …", after adjusting the order, still use this linking sentence, ensure that the linking logic matches the new arrangement order; if the adjustment involves multiple sections (such as the order of JS-002, SW-001, and SG-001), refer to the business logic order in the connection relationship table (technical solution - business quotation - construction organization) to rearrange and update the corresponding linking sentence, for example, adjust the linking sentence between JS-002 and SW-001 to "After the technical solution is determined, the corresponding business quotation is as follows …".
[0173] Step S155: According to the adjusted dynamic summary section association order and expression content, correct the position and text content of the corresponding semantic node in the field semantic chain of the tender document, supplement or delete the corresponding chain connection edge.
[0174] For the JS-001 summary of the supplementary content in step S153, find its corresponding semantic node JS-001, supplement "welding current control in a reasonable range" in the text content of the node, and update the field term list of the semantic node (if the supplementary content involves new terms, it needs to be added to the term list simultaneously).
[0175] For the XQ-002 and JS-001 nodes adjusted in step S154, correct the arrangement position of the two nodes in the tender document field semantic chain, ensure that the chain node order is consistent with the adjusted summary section order; if the adjustment causes the original chain connection edge to be invalid (such as the change of the associated node of a certain node), delete the invalid connection edge, add a new chain connection edge according to the new association relationship, for example, if the associated node of JS-001 changes from JS-003 to JS-002, delete the connection edge between JS-001 and JS-003, and add the connection edge between JS-001 and JS-002 and mark the association relationship type.
[0176] Step S156: Reassemble the adjusted dynamic summary sections and semantic linking sentences to form complete text content, check the semantic coherence and expression completeness of the text content, and obtain the final tender document summary.
[0177] According to the adjusted dynamic summary section association order, reassemble each section with the updated linking sentence to form complete text content. Check the semantic coherence of the text: read the entire text to confirm that the linking between sections is natural and there is no logical break (such as whether the linking between the technical solution section and the business quotation section conforms to the "technology - business" logic); check the expression completeness: confirm that all the content mentioned in the semantic missing feedback has been supplemented, the redundant expression has been deleted, and there is no key information missing or duplication.
[0178] For example, the check found that the pre-tightening force requirement of the "high-strength bolt connection standard" has been supplemented, the redundant expression of "weld detection" has been deleted, the sequence of each segment conforms to the business logic, and the judgment text meets the requirements, and the final bid document abstract is obtained.
[0179] Step S157: According to the format requirements of the target receiving end, the final bid document abstract is converted into the corresponding file format, and is sent to the target receiving end through the data transmission interface, and the final bid document abstract and the corresponding updated bid document field semantic chain are stored in the database.
[0180] If the target receiving end is an enterprise internal project management system, the format requirement is PDF format, then the final bid document abstract is converted into PDF format through the format conversion tool, to ensure that the format meets the system document upload specification (such as page size A4, font Songti); if the target receiving end is a mail recipient, the format requirement is Word format, then it is converted into Word format and set uniform header and footer (including project name, document type).
[0181] The converted file is sent to the target receiving end through the data transmission interface (such as HTTP interface, FTP interface): if it is a project management system, the PDF file is uploaded through the file upload interface of the system; if it is a mail recipient, the Word file is sent as an attachment through the mail sending interface.
[0182] At the same time, the final bid document abstract (text format and converted format file) and the updated bid document field semantic chain (graph structure data) are stored in the specified directory of the database, and the directory is named according to "project name-generation time" (such as "industrial plant steel structure installation-XXXX"), which is convenient for subsequent retrieval and reuse. The database automatically records the storage time, file size, associated bid document set and other metadata information, to ensure data traceability.
[0183] Figure 2 A schematic diagram of exemplary hardware and software components of a bid document abstract generation system 100 based on a large language model that can implement the idea of the present application is shown. For example, the processor 120 can be used in the bid document abstract generation system 100 based on a large language model, and is used to execute the functions in the present application.
[0184] The bid document abstract generation system 100 based on a large language model can be a general server or a special-purpose server, both of which can be used to implement the bid document abstract generation method based on a large language model of the present application. Although only one server is shown in the present application, for the sake of convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0185] For example, the large language model based bid document abstract generation system 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. Illustratively, the large language model based bid document abstract generation system 100 can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The methods of the present application can be implemented according to these program instructions. The large language model based bid document abstract generation system 100 also includes an I / O interface 150 between the computer and other input and output devices.
[0186] For ease of illustration, only one processor is described in the large language model based bid document abstract generation system 100. However, it should be noted that the large language model based bid document abstract generation system 100 in the present application can also include multiple processors, so the steps performed by one processor described in the present application can also be jointly performed or separately performed by multiple processors. For example, if the processor of the large language model based bid document abstract generation system 100 performs steps A and B, it should be understood that steps A and B can also be jointly performed by two different processors or separately performed in one processor. For example, a first processor performs step A, a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0187] In addition, the embodiment of the present application also provides a readable storage medium, wherein computer executable instructions are pre-set in the readable storage medium, and when a processor executes the computer executable instructions, the large language model based bid document abstract generation method is realized.
[0188] It should be noted that, in order to simplify the description of the present application and to help understand one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A method for generating bid document summaries based on a large language model, characterized in that, The method includes: Domain terms are extracted from the input set of tender documents to be processed to form a domain terminology library for tender documents. Based on the terminology associations in each tender document unit, each tender document unit is split into semantic nodes. The semantic nodes are connected through terminology associations to form a domain semantic chain for tender documents. The domain terminology database of the tender document is input into the pre-trained large language model, and the semantic parsing weights of the pre-trained large language model are adjusted to adapt the pre-trained large language model to the semantic parsing requirements of the tender document domain, thus obtaining a domain-adapted semantic parsing model. Each semantic node in the semantic chain of the tender document domain is input into the domain-adapted semantic parsing model. The core expression content of each semantic node is extracted through the domain-adapted semantic parsing model to form a dynamic summary fragment corresponding to each semantic node. All dynamic summary fragments are integrated to obtain a dynamic summary fragment set. Based on the connection relationship of semantic nodes in the semantic chain of the tender document domain, the association order of each dynamic summary fragment in the dynamic summary fragment set is determined, and the semantic connection processing of each dynamic summary fragment is performed according to the association order to obtain the preliminary fused summary text; The system receives semantic feedback from the tender document reader regarding the preliminary fused summary text, adjusts the content and order of each dynamic summary segment based on the semantic feedback, updates the semantic chain of the tender document domain, generates the final tender document summary, and outputs it to the target receiving end.
2. The method for generating bid document summaries based on a large language model according to claim 1, characterized in that, The process involves extracting domain terms from the input set of tender documents to form a domain terminology library for tender documents. Based on the terminology relationships within each tender document unit, each tender document unit is divided into semantic nodes. These semantic nodes are then connected through terminology relationships to form a domain semantic chain for tender documents. This includes: The text of each unit of the tender documents in the set of tender documents to be processed is scanned sentence by sentence to identify words with the domain attributes of tender documents. The words with the domain attributes of tender documents include words describing project requirements, words describing technical solutions, words describing commercial terms, and words describing implementation plans, thereby completing the extraction of domain terms from the set of tender documents to be processed. Frequency statistics are performed on the identified words with the domain of bidding documents. Words whose frequency of occurrence meets the core terminology standard of bidding document domain are retained, and duplicate words are deleted to form a terminology database of bidding document domain. Using terms from the domain terminology library of the tender documents as splitting markers, each tender document unit is divided into multiple text segments. Each text segment contains at least one domain term and its corresponding description. Each text segment serves as a semantic node. By comparing domain terms in different semantic nodes, semantic node combinations containing the same domain terms are identified, and term association relationships between semantic node combinations containing the same domain terms are determined. The term association relationships include term reference relationships, term subordination relationships, and term supplementation relationships. Using semantic nodes as chain nodes and the terminological relationships between semantic nodes as chain connecting edges, the chain nodes are arranged in the order of business description of the tender document units to form the initial tender document domain semantic chain. Check if there are isolated chain nodes without connecting edges in the domain semantic chain of the initial bid document. If so, analyze the potential associations between the domain terms in the isolated chain node and the domain terms in other chain nodes, and add connecting edges based on the potential associations to form a complete domain semantic chain of the bid document.
3. The method for generating bid document summaries based on a large language model according to claim 2, characterized in that, The step of comparing domain terms in different semantic nodes, identifying combinations of semantic nodes containing the same domain terms, and determining the terminology associations between combinations of semantic nodes containing the same domain terms includes: Traverse the text content of each semantic node, filter out words belonging to the domain terminology library of the tender document from the text content, and form a domain terminology set corresponding to each semantic node; Compare the domain term sets of any two semantic nodes item by item, record the semantic node pairs that have the same domain terms, and form a semantic node combination containing the same domain terms. In analyzing semantic node combinations containing terms from the same domain, the grammatical roles of the terms in the text content of the two semantic nodes are determined. If the term in one semantic node is the main description and the term in the other semantic node is the referential description, then a term reference relationship is determined between the two semantic nodes. Analyze the scope of the content to be expressed by the same domain terms in the combination of semantic nodes. If the content to be expressed by the same domain terms in one semantic node is a whole concept and the content to be expressed by the same domain terms in another semantic node is a local concept, then it is determined that there is a term dependency relationship between the two semantic nodes. Analyze the completeness of the expression content corresponding to the same domain term in the semantic node combination. If the expression content of the same domain term in two semantic nodes covers different aspects and together constitutes a complete concept, then it is determined that there is a term supplement relationship between the two semantic nodes. Add a corresponding term association type label to each semantic node combination containing terms from the same domain to form a list of semantic node associations.
4. The method for generating bid document summaries based on a large language model according to claim 1, characterized in that, The terminology database of the bidding document domain is input into a pre-trained large language model. The semantic parsing weights of the pre-trained large language model are adjusted to adapt the pre-trained large language model to the semantic parsing requirements of the bidding document domain, resulting in a domain-adapted semantic parsing model, including: Core terms are selected from the domain terminology database of the tender documents, and corresponding standard expression texts of the tender document domain are matched for each core term to form a domain-adapted corpus pair consisting of core terms and standard expression texts; The domain-adapted corpus pairs are converted according to the input format requirements of the pre-trained large language model and then input in batches into the corpus input module of the pre-trained large language model. Call the parameter viewing interface of the pre-trained large language model to obtain the set of weight parameters related to term recognition and semantic understanding in the semantic parsing layer; Based on the core terms in the domain-adapted corpus, the weight parameter value for identifying domain terms in the bidding documents is increased in the semantic parsing layer, while the weight parameter value for identifying general domain terms is decreased, so that the semantic parsing layer prioritizes the identification of domain terms in the bidding documents. The adjusted semantic parsing weight parameters are encapsulated into an independent domain adaptation layer module, which includes a term weight adjustment submodule, a domain semantic correction submodule, and a representation standardization submodule. The domain adaptation layer module is embedded between the input layer and the semantic parsing layer of the pre-trained large language model, so that the pre-trained large language model processes the input text first through the domain adaptation layer module. The test tender document text is input into a pre-trained large language model embedded in the domain adaptation layer module. The consistency of the output semantic parsing result is compared with the standard parsing result of the tender document domain. If the consistency meets the preset adaptation standard, the domain-adapted semantic parsing model is determined.
5. The method for generating bid document summaries based on a large language model according to claim 4, characterized in that, The step of encapsulating the adjusted semantic parsing weight parameters into an independent domain adaptation layer module includes: The program of the terminology weight adjustment submodule is written into the identification weight parameters of the domain terms in the tender document. When the terminology weight adjustment submodule receives input text, it can enhance the identification strength of the domain terms in the tender document based on the identification weight parameters and output text containing domain term tags. Import the semantic rule library of the bidding document domain. The semantic rule library of the bidding document domain contains common semantic expression logic rules in the bidding document domain. This enables the semantic correction submodule of the domain to compare the semantic expression of the input text with the rules in the semantic rule library of the bidding document domain and correct the expression content that does not conform to the domain semantic rules. The system stores standard expression templates for the bidding document field. These templates cover the standardized expression formats for project requirements, technical solutions, and commercial terms. This allows the expression specification submodule to adjust the corrected semantic expressions to conform to the format of the standard expression templates for the bidding document field. After receiving the input text, the constraint domain adaptation layer module first calls the term weight adjustment submodule to process it, and then inputs the processing result into the domain semantic correction submodule. Finally, it inputs the correction result into the expression standardization submodule. After integrating the program code of the terminology weight adjustment submodule, the domain semantic correction submodule, and the expression standardization submodule, input the tender document text fragment containing non-standard expressions into the domain adaptation layer module. Check whether the output text processed by the terminology weight adjustment submodule, the domain semantic correction submodule, and the expression standardization submodule in sequence meets the requirements of tender document domain terminology marking, semantic rules, and expression standardization. If they meet the requirements, the construction of the domain adaptation layer module is completed, and the test of the submodule collaboration effect is completed.
6. The method for generating bid document summaries based on a large language model according to claim 1, characterized in that, Each semantic node in the semantic chain of the tender document domain is input into the domain-adapted semantic parsing model. The core expression content of each semantic node is extracted through the domain-adapted semantic parsing model to form a dynamic summary fragment corresponding to each semantic node. All dynamic summary fragments are integrated to obtain a dynamic summary fragment set, including: According to the arrangement order of the chain nodes in the semantic chain of the tender document domain, the semantic nodes corresponding to each chain node are read in sequence to form a semantic node sequence, thus completing the extraction of the semantic node sequence of the semantic chain of the tender document domain. The text content of each semantic node in the semantic node sequence is sequentially input into the text input port of the domain-adapted semantic parsing model, thus completing the sequential input of semantic nodes into the domain-adapted semantic parsing model. In the domain-adapted semantic parsing model, a core information extraction mode is enabled. This core information extraction mode is used to filter out sentences that embody the core meaning of semantic nodes from the input semantic node text content based on the domain semantic rules of the tender document. The core meaning sentences selected by the domain-adapted semantic parsing model are simplified and expressed to form dynamic summary fragments corresponding to the semantic nodes. Add a corresponding semantic node identifier, its position identifier in the semantic chain of the tender document, and a list of domain terms to each dynamic summary fragment; All generated dynamic summary fragments are arranged according to the order of their corresponding semantic nodes in the semantic node sequence, and associated with the corresponding related information to form a dynamic summary fragment set.
7. The method for generating bid document summaries based on a large language model according to claim 6, characterized in that, The aforementioned core information extraction mode is enabled in the domain-adapted semantic parsing model. This core information extraction mode is used to filter sentences that embody the core meaning of semantic nodes from the input semantic node text content based on the domain semantic rules of the tender document, including: In the function selection interface of the domain-adapted semantic parsing model, check the core information extraction option and set the running parameters of the core information extraction mode. The running parameters include the number of core sentences to be filtered and the proportion of core terms to be retained. From the built-in rule base of the domain-adapted semantic parsing model, target rules related to the identification of core information in the tender document domain are called. These target rules include sentence priority rules for core terms and sentence priority rules for business logic keywords. The input semantic node text content is divided into independent sentence units according to punctuation marks. A unique sentence identifier is assigned to each sentence unit. Each sentence unit is compared with the loaded bid document domain semantic rules, and the number of rules that each sentence unit conforms to is counted. The number of rules met by each sentence unit is sorted, and the sentence unit with the most rules met is selected as the candidate core sentence. Check whether the candidate core sentence contains the core domain terminology of the semantic node and whether it can independently express a complete semantic meaning. If it meets the requirements, it is determined to be a sentence that embodies the core meaning of the semantic node.
8. The method for generating bid document summaries based on a large language model according to claim 1, characterized in that, The process involves determining the association order of dynamic summary fragments in the dynamic summary fragment set based on the connection relationships of semantic nodes in the semantic chain of the tender document domain, and performing semantic connection processing on each dynamic summary fragment according to the association order to obtain a preliminary fused summary text, including: From the structural data of the semantic chain in the tender document domain, extract the link connection information between each semantic node to form a connection table containing semantic node identifiers, associated node identifiers, and association relationship types; By using the semantic node identifier of the dynamic summary fragment, each dynamic summary fragment is associated with the corresponding semantic node entry in the connection relationship table, thereby determining the associated fragment identifier and association relationship type of each dynamic summary fragment; Starting with the dynamic summary fragment corresponding to the initial semantic node of the domain semantic chain of the tender document, the subsequent dynamic summary fragments are arranged in sequence according to the order of the associated nodes in the connection relationship table to form the dynamic summary fragment association order; Based on the relationship type between adjacent dynamic summary fragments, the bid document domain connection statement template library is invoked. The bid document domain connection statement template library contains connection statement templates corresponding to different relationship types. Connection statements connecting adjacent dynamic summary fragments are generated based on the connection statement templates. Following the order of association, corresponding semantic connecting statements are inserted between two adjacent dynamic summary segments to form a continuous text sequence. The length distribution of sentences in the text sequence is analyzed, and excessively long sentences are broken down into short sentences that conform to reading habits. The word order between sentences is adjusted, and the adjusted text sequence is formatted and repetitive connecting statements are removed to obtain a preliminary fused summary text.
9. The method for generating bid document summaries based on a large language model according to claim 8, characterized in that, The step involves calling a domain-specific connector template library for tender documents based on the association type of adjacent dynamic summary fragments. This library contains connector templates corresponding to different association types. Connecting statements linking adjacent dynamic summary fragments are generated based on these templates, including: Based on the default term association type in the semantic chain of the tender document domain, the connector statement templates are divided into referential connector statement templates, subordinate connector statement templates, and supplementary connector statement templates, which are stored in different directories of the tender document domain connector statement template library. Read the relationship type of adjacent dynamic summary fragments. If it is a term reference relationship, call the reference relationship connection statement template. If it is a term subordination relationship, call the subordination relationship connection statement template. If it is a term supplement relationship, call the supplement relationship connection statement template. From the association information of two adjacent dynamic summary segments, obtain the core domain terms contained in each segment and determine the core domain terms common to the two dynamic summary segments; Substitute common core domain terms into the terminators in the called connector template, replace the general expressions in the connector template with specific expressions related to the core terms, and generate preliminary connector statements. Based on the overall style of the preliminary integrated abstract text, the language style of the preliminary connecting sentences is adjusted to be consistent with the overall style. The connecting sentences are checked to see if they accurately reflect the relationship between adjacent dynamic abstract segments and whether there is any semantic ambiguity. If so, the content is modified until the semantics are accurate.
10. A tender document summary generation system based on a large language model, characterized in that, The method includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the bid document summary generation method based on a large language model as described in any one of claims 1-9.
Citation Information
Patent Citations
Dialogue long context precision optimization method based on abstract and index
CN120561285A
Purchase review document abstract generation method and system in combination with AI large model
CN120930649A