Intelligent construction special scheme paragraph intelligent compilation method based on information network
Through the intelligent compilation method of information network, the complex and time-consuming problem of traditional construction plans is solved, efficient and standardized construction plans are achieved, and the quality and applicability of the plans are improved.
Patent Information
- Application Number
- CN202510839037.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The preparation methods of traditional construction technical solutions are complex, time-consuming and labor-intensive, unstable in quality, lack of unified standards, and it is difficult to effectively guide actual construction.
The intelligent compilation method of intelligent construction special plan paragraphs based on the information network is adopted, and the construction plan that meets the standards is generated through word segmentation, vector embedding, index construction, similarity search and large-scale model generation, combined with unsupervised and supervised fine-tuning.
The efficiency and quality of construction plan preparation have been improved, the plan complies with industry standards, reduce manual writing time, and generate customized and tailor-made plans to meet actual needs.
Smart Images

Figure CN120336334A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of construction technical plan compilation, and particularly to an intelligent compilation method for paragraphs of a special construction plan based on an information network and intelligent construction. Background Art
[0002] The compilation of a construction technical plan is a key link to ensure the smooth implementation of a project. However, traditional construction technical plan compilation methods face difficulties such as a large compilation workload, low reuse level, failure to fully play the supporting role of the backend, and uneven quality.
[0003] The original plan compilation process is complex and mainly relies on manual writing. From collecting materials, organizing ideas to writing word by word, it consumes a large amount of time and manpower.
[0004] Previously, the plan content lacked unified standards in terms of format, expression, etc. The styles and habits of different compilation personnel varied greatly, resulting in inconsistent plan content styles. Some plans were not detailed or accurate enough in key technical parameters, construction process descriptions, etc. At the same time, it also increased the difficulty of review and modification, easily caused misunderstandings and disputes during the construction process, and posed potential threats to project quality and safety.
[0005] Currently, the compilation of technical plans mainly relies on manual compilation, and the quality of the plan depends on the professional knowledge and experience of the plan compilation personnel. The low quality of some plan content is mainly reflected in the lack of in-depth understanding of construction techniques, failure to fully combine the actual characteristics and requirements of the project, and applying past experience or templates. As a result, the plan content lacks pertinence and operability, and the construction techniques cannot effectively guide actual construction. Therefore, an intelligent compilation method for paragraphs of a special construction plan based on an information network and intelligent construction is needed. Summary of the Invention
[0006] Based on the existing technical problems, the present invention proposes an intelligent compilation method for paragraphs of a special construction plan based on an information network and intelligent construction.
[0007] An intelligent compilation method for paragraphs of a special construction plan based on an information network and intelligent construction proposed by the present invention includes the following steps: Step 1: According to different input contents, the content processing is divided into two categories: project general content and project characteristic content;
[0008] Step 2: Vector embedding;
[0009] Step 3: Index construction;
[0010] Step 4: Index storage;
[0011] Step 5: Similarity retrieval;
[0012] Step 6: Output the retrieval result;
[0013] Step 7: Context construction;
[0014] Step 8: Input the large model;
[0015] Step 9: Use RAG and large model algorithms to generate intelligent compilation results that meet the compilation requirements. Among them, for the adaptation of the large model in a specific field, unsupervised fine-tuning and supervised fine-tuning of the large model are adopted.
[0016] Preferably, the general project content in step 1 is the general project compilation principle, compilation basis, and general content according to historical plans and standard requirements;
[0017] Project characteristic content is the content according to the project engineering overview, natural conditions, construction key points and difficulties, and construction plan characteristic requirements;
[0018] After receiving the input instruction, the two types of input content are tokenized, expressed as:
[0019] ; where means splitting the input text by characters to form sequence, and then converting it into the index sequence in the corpus.
[0020] Preferably, the vector embedding in step 2 adopts the formula:
[0021] ; where represents the embedding matrix, represents the vector of the input project general content, represents the vector of the input project characteristic content.
[0022] Preferably, the index construction in step 3 adopts the formula:
[0023] ; where represents the historical plan document, represents the segmented document set, represents a function, represents the sequence after tokenizing the document set, represents the set of vectorized index vectors.
[0024] Preferably, the formula adopted for index storage in step 4 is: ; where represents a function for storing into the database, represents the constructed index structure stored in the vector database.
[0025] Preferably, the formula used for similarity retrieval in step five is: ; where represents calculating the similarity between the input vector and the index set.
[0026] Preferably, the formula used for outputting the retrieval result in step six is: ; where, represents a function for sorting the expression within the parentheses, represents the retrieval result after sorting by similarity, and the retrieval results with the highest similarity are taken as the output of the RAG algorithm.
[0027] Preferably, the formula for context construction in step seven is: ; where, represents the context input information formed by concatenating the general content vector of the input item , the project feature content vector and the RAG retrieval result vector set .
[0028] Preferably, the formula for inputting the large model in step eight is: ; where, represents the compilation result output by the large model .
[0029] Preferably, the unsupervised fine-tuning process adopted in step nine includes the following implementation methods: data collection, data cleaning, data transformation, and fine-tuning training;
[0030] The data collection process includes screening and separating historical solutions, screening and organizing the basis description documents, screening and organizing standard specification data, summarizing project technical experience, and screening and analyzing industry experts' opinions and cases;
[0031] The data cleaning process includes:
[0032] a1. Removing noise, deleting irrelevant information such as comments and duplicate content;
[0033] a2. Text preprocessing, performing steps of paragraphing and stemming;
[0034] a3. Handling missing values, filling in and deleting missing data;
[0035] a4. Standardization, ensuring that the data format and feature formula meet the requirements of model input; the cleaned data is converted into the format of model training, including word segmentation and vectorization, which can be expressed as: ;
[0036] Fine-tune the model using self-supervised learning and unlabeled data, and optimize the model parameters to adapt to specific tasks; adopt the next-word prediction method of GPT, and the training loss function is expressed as: ; where represents the loss value during the training process, represents the sum from the first to the th . represents the th , represents the first th , represents the probability that the model predicts the next ;
[0037] In the supervised fine-tuning process implementation method adopted in Step 9, the data annotation process is added compared with the unsupervised process, and the language model is made to match the user's approach to various tasks by fine-tuning based on human feedback;
[0038] The data annotation and fine-tuning process adopt method. To ensure the accuracy and professionalism of the annotated data, experts and experienced compilers are selected for data annotation to compare and rank multiple answers for the collected data; the fine-tuning training process adopts the reward model training as follows: ; where is a constant, is the combination number, the combination number of selecting two elements from elements, is the prompt with parameter and the scalar output of the reward model for completion and completion , and are the generated results of the winner and the loser respectively, is the dataset containing the comparison results after annotation, is function, used to map the difference to the interval.
[0039] The beneficial effects of the present invention are:
[0040] Using the RAG technology, the general content of the input project enhances the generation ability of the large model by retrieving historical solution information and compilation principles. The algorithm process includes three core steps: index construction, retrieval, and generation. Index construction divides document paragraphs into text chunks, uses the Transformer encoder model to convert the text chunks into vector representations, stores them in the vector database Chroma, and constructs a vector index. During the retrieval process, the query text input by the user is converted into a vector, the similarity between the query vector and the vectors in the database is calculated, and the top k most relevant text chunks are retrieved. The retrieved text chunks are used as context information and input into the large model together with the project feature content to generate the final compilation result; the compilation efficiency and quality are improved through automated and intelligent means. Description of the Drawings
[0041] Figure 1 It is a schematic flow chart of an intelligent compilation method for paragraphs of a special construction plan based on an information network;
[0042] Figure 2 It is a schematic RAG algorithm flow chart of an intelligent compilation method for paragraphs of a special construction plan based on an information network. Detailed Implementation Modes
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0044] Refer to Figure 1 - Figure 2 , an intelligent compilation method for paragraphs of a special construction plan based on an information network, includes the following steps: Step 1: According to different input contents, the content processing is divided into two categories: general project content and project feature content; the general project content in Step 1 is the general project compilation principles, compilation basis, and general content according to historical solutions and standard requirements.
[0045] Project feature content is the content according to the project engineering overview, natural conditions, construction key points and difficulties, and construction plan feature requirements.
[0046] After receiving the input instruction, the two types of input contents are tokenized, expressed as:
[0047] ; among them, means that the input text is segmented according to characters to form sequence, and then converted into an index sequence in the corpus.
[0048] Step 2: Vector embedding; the vector embedding in Step 2 uses the formula:
[0049] ; Among them, represents the embedding matrix, is represented as the general content vector of the input item, is represented as the characteristic content vector of the input item.
[0050] Step 3: Index construction; The formula used for index construction in Step 3 is:
[0051] ; Among them, represents the historical solution document, is represented as the document set after partitioning, represents a function, is represented as after tokenizing the document set sequence, is represented as the set of vectorized index vectors.
[0052] Step 4: Index storage; The formula used for index storage in Step 4 is: ; Among them, represents a function for storing into the database, is represented as the constructed index structure stored in the vector database.
[0053] Step 5: Similarity retrieval; The formula used for similarity retrieval in Step 5 is: ; Among them represents calculating the similarity between the input vector and the index set.
[0054] Step 6: Output retrieval results; The formula used for outputting retrieval results in Step 6 is: ; Among them, represents a function for sorting the expression within the parentheses, is represented as the retrieval results after sorting by similarity, taking the retrieval results with the highest similarity as the output of the RAG algorithm.
[0055] Step 7: Context construction; The formula for context construction in Step 7: ; Among them, represents concatenating the general content vector of the input item , the characteristic content vector of the item and the RAG retrieval result vector set to form the context input information.
[0056] Step 8: Input into the large model; The formula used for inputting into the large model in Step 8 is: ; Among them, Represented as a large model The compiled result output.
[0057] Based on the use of a large model, the RAG algorithm adds the retrieval and matching of the historical knowledge base, making the generated compiled result conform to the standardized template and better meet the compilation requirements. It can efficiently and intelligently compile general content to ensure that the theme and content generated by the large model meet the requirements of the solution compilation.
[0058] Use a large model to automatically annotate the historical solutions and standardized solutions of our bureau to form a global solution element knowledge base, and then compare according to the characteristics of the project solution to find the most similar solution for overall recommendation.
[0059] Step 9: Use the RAG and large model algorithms to generate an intelligent compilation result that meets the compilation requirements. To adapt the large model to a specific field, unsupervised fine-tuning and supervised fine-tuning of the large model are adopted. The unsupervised fine-tuning process adopted in Step 9 includes the following implementation methods: data collection, data cleaning, data transformation, and fine-tuning training.
[0060] The data collection process includes the screening and separation of historical solutions, the screening and sorting of compilation basis description documents, the screening and sorting of standard specification data, the summary of project technical experience, and the screening and sorting of industry expert insights and case analyses.
[0061] The data cleaning process includes:
[0062] a1. Remove noise, delete comments and duplicate and irrelevant information;
[0063] a2. Text preprocessing, performing paragraphing and stemming steps;
[0064] a3. Handle missing values, fill in and delete missing data;
[0065] a4. Standardization to ensure that the data format and features meet the requirements of model input; the cleaned data is converted into the format of model training, including tokenization and vectorization, which can be expressed as: .
[0066] Use self-supervised learning and unlabeled data to fine-tune the model, optimize the model parameters to adapt to specific tasks; adopt the next-word prediction method of GPT, and the training loss function is expressed as: ; where represents the loss value during the training process, represents from the first to the th sum, represents the th , represents the previous one , represents the probability that the model predicts the next .
[0067] The unsupervised fine-tuning process makes the generated results of the large model prefer to be closer to the project compilation requirements and compilation habits. Fine-tuning and adapting according to special domain data provides results that are more accurate and more in line with the requirements of scheme compilation.
[0068] The implementation method of the supervised fine-tuning process adopted in step nine adds a data annotation process compared with the unsupervised process, and fine-tunes through human feedback to align the language model with the user's approach to various tasks.
[0069] The data annotation and fine-tuning process adopts the following method. To ensure the accuracy and professionalism of the annotated data, experts and experienced compilers are selected for data annotation to compare and rank multiple answers for the collected data; the fine-tuning training process adopts the reward model training as follows: ; where is a constant, is the combination number, the combination number of selecting two elements from elements, is the prompt with parameter and and the scalar output of the reward model for completion, and are the generated results of the winner and loser respectively, is the dataset containing the comparison results after annotation, is a function used to map the difference to the interval.
[0070] The supervised fine-tuning process can make the output of the large model follow the intentions and preferences of experts, reducing the generation of outputs that are untrue, incorrect, or simply useless to users.
[0071] Using the RAG technology, the general content of the input project enhances the generation ability of the large model by retrieving historical solution information and compilation principles. The algorithm process includes three core steps: index construction, retrieval, and generation. Index construction divides document paragraphs into text chunks, uses the Transformer encoder model to convert the text chunks into vector representations, stores them in the vector database Chroma, and constructs a vector index. The retrieval process converts the query text input by the user into a vector, calculates the similarity between the query vector and the vectors in the database, and retrieves the top k most relevant text chunks. The retrieved text chunks are used as context information and combined with the project characteristic content input information to be input into the large model to generate the final compilation result; the compilation efficiency and quality are improved through automated and intelligent means.
[0072] Through automated word segmentation, vector embedding, and index construction, the efficiency of compiling construction special plans has been greatly improved, and the time required for manual writing has been reduced; by using historical plans and standard requirements, it is ensured that the generated plan content meets industry standards and specifications, and the standardization degree of the plan has been improved; by distinguishing between project general content and project characteristic content, customized plans can be generated according to the characteristics and requirements of specific projects; through the retrieval and matching of the RAG algorithm and the historical knowledge base, historical experience and professional knowledge have been effectively integrated, improving the quality and practicality of the plan; the unsupervised and supervised fine-tuning processes enable the large model to more accurately adapt to the needs of specific fields and generate results that are more in line with actual compilation requirements; the combination of automated annotation and expert annotation reduces human errors and improves the accuracy and professionalism of plan compilation; the established vector database and knowledge base are easy to update and maintain and can be continuously optimized as new projects and new experiences accumulate.
[0073] The generated plan is more in line with compilation requirements and habits, improving the overall quality of the plan; the automated process significantly shortens the plan compilation cycle and improves work efficiency; through similarity retrieval and integration of expert experience, more powerful and accurate data support is provided for decision-makers; the manpower and time costs required for manual compilation are reduced, and the overall project cost is lowered; the generated plan is more in line with actual needs, improving the satisfaction and trust of project stakeholders; through the digital storage and utilization of historical plans and expert experience, the inheritance and sharing of knowledge and experience are promoted; this method can adapt to different types and scales of construction projects and has wide applicability.
[0074] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for intelligent compilation of paragraphs in a special construction plan based on an information network, characterized in that: It includes the following steps: Step 1, classify the content processing into two major categories according to different input content: project general content and project characteristic content; Step 2, vector embedding; Step 3, index construction; Step 4, index storage; Step 5, similarity retrieval; Step 6, output the retrieval result; Step 7, context construction; Step 8, input into the large model; Step 9, use RAG and large model algorithms to generate intelligent compilation results that meet the compilation requirements. Among them, in order to adapt the large model to a specific field, unsupervised fine-tuning and supervised fine-tuning of the large model are adopted.
2. The intelligent compilation method for paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The general content of the project in Step 1 is the general content of the general project compilation principles, compilation basis according to historical plans and standard requirements; Project feature content It is the content required according to the general situation of the project engineering, natural conditions, construction key points and difficulties, and construction plan features; After receiving the input instruction, segment the two types of input content, which is expressed as: ; wherein, means splitting the input text by characters to form a sequence, and then converting it into an index sequence in the corpus.
3. A method for intelligent compilation of paragraph of a special construction plan based on an information network according to claim 1, characterized in that: The vector embedding in Step 2 adopts the formula: ; wherein, represents an embedding matrix, is represented as a general content vector of the input item, is represented as a feature content vector of the input item.
4. A method for intelligent compilation of paragraphs in a special construction plan for intelligent construction based on an information network according to claim 1, characterized in that: The index construction in Step 3 adopts the formula: ; Among them, is represented as a historical solution document, is represented as a set of chunked documents, is represented as a function, is represented as after word segmentation of the document set sequence, is represented as a set of vectorized index vectors.
5. A method for intelligent compilation of paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The formula used for index storage in the fourth step is as follows: ; where is expressed as a function for storing into the database, is expressed as the constructed index structure stored in the vector database.
6. The intelligent compilation method for paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The formula used for similarity retrieval in the fifth step is as follows: ; where represents the calculation of the similarity between the input vector and the index set.
7. A method for intelligent compilation of paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The formula used to output the retrieval results in Step 6 is as follows: ; where represents a function for sorting the expression within the parentheses, represents the retrieval results after sorting by similarity, and the top retrieval results with the highest similarity are taken as the output of the RAG algorithm.
8. The intelligent compilation method for paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The formula used for context construction in Step 7 is: ; where is expressed as the context input information formed by concatenating the general content vector of the input item , the characteristic content vector of the item and the RAG retrieval result vector set .
9. A method for intelligent compilation of paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The formula adopted by the large model input in Step 8 is as follows: ; where represents the compilation result output by the large model .
10. A method for intelligent compilation of paragraphs of a special construction plan based on an information network according to claim 1, characterized in that: The unsupervised fine-tuning process adopted in Step 9 includes the following implementation methods: data collection, data cleaning, data transformation, and fine-tuning training; The data collection process includes the screening and separation of historical solutions, the screening and collation of compilation basis description documents, the screening and collation of standard specification data, the summary of project technical experience, and the screening and collation of industry expert opinions and case analyses; The data cleaning process includes: a1. Remove noise, delete comments and duplicate irrelevant information; a2. Text preprocessing, perform paragraphing and stemming processing steps; a3. Process missing values, fill in and delete missing data; a4. Standardize to ensure that the data format and feature formula meet the requirements of model input; convert the cleaned data into the format for model training, including word segmentation and vectorization, which can be expressed as: ; Fine-tune the model using self-supervised learning and unlabeled data, and optimize the model parameters to adapt to specific tasks; adopt the next-word prediction method of GPT, and the training loss function is expressed as: ; where represents the loss value during the training process, represents the sum from the first to the th ; represents the th , represents the first th , represents the probability that the model predicts the next ; The implementation method of the supervised fine-tuning process adopted in Step 9 adds a data annotation process compared with the unsupervised process, and fine-tunes through human feedback to enable the language model to match the user's approach to various tasks for various tasks. The data annotation and fine-tuning process adopts the following method. To ensure the accuracy and professionalism of the annotated data, experts and experienced compilers are selected for data annotation to compare and rank multiple answers for the collected data; the fine-tuning training process adopts the reward model training as follows: wherein, is a constant, is the combination number, which is the combination number of selecting two elements from elements, is a prompt with parameter , and is the scalar output of the reward model for completion, and are the generation results of the winner and the loser respectively, is the data set containing the comparison results after annotation, is a function used to map the difference to the interval.
Citation Information
Patent Citations
Building industry construction scheme intelligent compilation method and system
CN113298435A
Intelligent generation method and system for engineering construction scheme compilation basis
CN115982329A
Construction scheme multi-modal retrieval enhancement generation method and system based on large model
CN119576987A
Intelligent compilation system and method of building construction scheme
CN119646131A
Data knowledge extraction method, system and equipment based on large language model and storage medium
CN119646134A
Cited By
Intelligent engineering construction scheme production method based on knowledge base and multivariate data analysis
CN121579615A
Risk adaptation type intelligent compilation method for construction full-content type
CN121639121A
A risk-adaptive intelligent compiling method for construction full content type
CN121639121B