Building construction document intelligent identification method and device and electronic equipment
By performing completeness checks and intention recognition on building construction documents and calling temporary knowledge bases for document recognition, the problem of static knowledge base adaptation is solved, and the recognition accuracy and accuracy of question-and-answer results are improved.
Patent Information
- Application Number
- CN202510457512.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing building construction document recognition system, it is difficult for the static knowledge base to dynamically adapt to changes in industry regulations, resulting in a decrease in the accuracy of identification results, lack of a special Q&A interactive decision-making mechanism, and it is difficult to provide accurate Q&A results.
By receiving user problem information and construction documents, after performing integrity verification, we identify the user's intentions, determine whether to call the temporary knowledge base, use the temporary knowledge base for document identification, and generate special question and answer results.
It improves the accuracy of document recognition, provides more accurate Q&A results, and realizes a special Q&A interactive decision-making mechanism for user questions.
Smart Images

Figure CN120296134A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the fields of document recognition and human-computer interaction, and particularly to an intelligent recognition method, device, and electronic device for construction project documents. Background Art
[0002] With the gradual penetration of artificial intelligence technology into the construction industry, the standardized review of construction project documents has become a key link in ensuring project quality, safety, and compliance. However, the knowledge base set in the existing recognition systems is a static knowledge base, which is difficult to dynamically adapt to changes in industry regulations. At the same time, there is a lack of a special Q&A interaction decision-making mechanism for assisting document review. As a result, the accuracy of document recognition results is reduced, and it is difficult to provide accurate Q&A results for users.
[0003] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] This section of the present disclosure is used to introduce the inventive concept in a brief form, which will be described in detail in the following detailed implementation section. This section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Some embodiments of the present disclosure provide an intelligent recognition method, device, and electronic device for construction project documents to solve the technical problems mentioned in the above background art section.
[0006] In a first aspect, some embodiments of the present disclosure provide an intelligent recognition method for construction project documents, the method comprising: in response to receiving user question information uploaded by a user through a user terminal and receiving a synchronously uploaded construction project document, where after receiving, integrity verification is performed on the above construction project document; performing user intent recognition on the above user question information to generate user intent information, where the above user intent information includes: an intent tag set; determining whether it is necessary to call a temporary knowledge base according to the above user intent information; in response to determining to call, calling a temporary knowledge base corresponding to the above construction project document; using the above temporary knowledge base and the above user intent information to perform document recognition on the above construction project document to generate a document recognition result corresponding to the above user question information; and sending the above document recognition result as a user Q&A result to the above user terminal.
[0007] Second aspect, some embodiments of the present disclosure provide an intelligent recognition device for construction documents. The device includes: a receiving unit configured to respond to receiving user question information uploaded by a user through a user terminal and receiving a construction document uploaded synchronously, wherein after receiving, integrity verification is performed on the above-mentioned construction document; an intention recognition unit configured to perform user intention recognition on the above-mentioned user question information to generate user intention information, wherein the above-mentioned user intention information includes: an intention tag set; a determination unit configured to determine whether to call a temporary knowledge base according to the above-mentioned user intention information; a calling unit configured to, in response to determining to call, call a temporary knowledge base corresponding to the above-mentioned construction document; a document recognition unit configured to use the above-mentioned temporary knowledge base and the above-mentioned user intention information to perform document recognition on the above-mentioned construction document to generate a document recognition result corresponding to the above-mentioned user question information; a sending unit configured to send the above-mentioned document recognition result as a user Q&A result to the above-mentioned user terminal.
[0008] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having stored thereon one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the above first aspect.
[0009] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, wherein when the program is executed by a processor, the method described in any implementation manner of the above first aspect is implemented.
[0010] The above embodiments of the present disclosure have the following beneficial effects: Through the intelligent recognition method of construction document in some embodiments of the present disclosure, the accuracy of document recognition can be improved, and more accurate Q&A results can be provided for users. Specifically, the reasons for the reduction in the accuracy of document recognition results and the difficulty in providing accurate Q&A results for users are as follows: The knowledge base set in the existing recognition system is a static knowledge base, which is difficult to dynamically adapt to the changes in industry regulations. At the same time, there is a lack of a special Q&A interaction decision-making mechanism for assisting document review. Based on this, in the intelligent recognition method of construction document in some embodiments of the present disclosure, first, in response to receiving user question information uploaded by the user through the user terminal and receiving the synchronously uploaded construction document, where the integrity verification of the above construction document is performed after receiving. Here, through document integrity verification, the recognition of incomplete documents can be avoided, and the recognition accuracy can be improved. Then, the user intention recognition is performed on the above user question information to generate user intention information, where the above user intention information includes: an intention tag set. Here, through user intention recognition, the intention of the user can be determined, and targeted document recognition can be performed accordingly, so as to provide more accurate Q&A results for users. After that, according to the above user intention information, it is determined whether it is necessary to call the temporary knowledge base. Here, by establishing a temporary knowledge base, special Q&A results can be generated. Then, in response to determining to call, the temporary knowledge base corresponding to the above construction document is called. Next, using the above temporary knowledge base and the above user intention information, the above construction document is recognized to generate a document recognition result corresponding to the above user question information. Thus, a special Q&A interaction decision-making mechanism for user questions is realized. Thereby, the accuracy of generating document recognition results can be improved. Finally, the above document recognition result is sent to the above user terminal as the user Q&A result. Furthermore, more accurate Q&A results can be provided for users. Brief Description of the Drawings
[0011] Combined with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of each embodiment of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.
[0012] Figure 1 is an application scenario diagram according to some embodiments of the intelligent recognition method of construction document of the present disclosure;
[0013] Figure 2 is a flowchart according to some embodiments of the intelligent recognition method of construction document of the present disclosure;
[0014] Figure 3 is a schematic diagram of intelligent recognition;
[0015] Figure 4 is a schematic structural diagram of some embodiments of an intelligent recognition device for construction documents according to the present disclosure;
[0016] Figure 5 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] In addition, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0023] Figure 1 is a schematic diagram of an application scenario of an intelligent recognition method for construction documents according to some embodiments of the present disclosure.
[0024] In Figure 1In the application scenario, first, the computing device 101 can respond to receiving user question information (i.e., user input) uploaded by the user through the user terminal, and receive a synchronously uploaded construction document (not shown in the figure). Among them, after receiving, the integrity verification of the above construction document is performed. Then, the computing device 101 can perform user intent recognition on the above user question information to generate user intent information, where the above user intent information includes: an intent tag set. For example, intent recognition can be performed through a large model. After that, the computing device 101 can determine whether it is necessary to call the temporary knowledge base according to the above user intent information. For example, the temporary knowledge base can include a knowledge base or a vector database. Next, the computing device 101 can respond to the determination to call and call the temporary knowledge base corresponding to the above construction document. Then, the computing device 101 can use the above temporary knowledge base and the above user intent information to perform document recognition on the above construction document to generate a document recognition result corresponding to the above user question information. For example, Figure 1 In the step of recall information integration, through text summarization and structural verification, the final document recognition result is obtained. Finally, the computing device 101 can send the above document recognition result to the above user terminal as a user Q&A result.
[0025] It should be noted that the above computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here. It should be understood that, Figure 1 the number of computing devices in can have any number according to implementation needs.
[0026] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0027] Figure 2 The flow 200 of some embodiments of the intelligent recognition method for construction documents according to the present disclosure is shown. The intelligent recognition of construction documents includes the following steps:
[0028] Step 201, in response to receiving user question information uploaded by the user through the user terminal, and receiving a synchronously uploaded construction document.
[0029] In some embodiments, the execution subject of the intelligent recognition method for construction documents (e.g., a computing device) can respond to receiving user question information uploaded by a user through a user terminal and a synchronously uploaded construction document in a wired or wireless manner. Among them, after receiving, integrity verification is performed on the above-mentioned construction document. Secondly, the user question information can be a question raised by the user regarding the construction document. For example, "Please correct the typos in the document", "Please adjust the document layout", "Please give the document theme", etc. It can also be a question raised by the user alone. For example, "How to measure the floor slab thickness". The construction document can be a document organized for the process of a certain construction work or the usage method of a certain intelligent device. Here, the integrity verification can be performed through the following steps: First, detect the page numbers of the construction document to determine whether there are missing page numbers in the document. Then, it can be detected whether the last sentence on the last page of the construction document is missing. Specifically, the integrity of the last sentence can be determined through a preset template matching algorithm. In this way, it can be determined whether there is an incomplete last sentence. Thus, integrity verification is achieved. In addition, for a construction document that fails the integrity verification, the user can be instructed to upload it again.
[0030] Step 202, perform user intention recognition on the user question information to generate user intention information.
[0031] In some embodiments, the above-mentioned execution subject can perform user intention recognition on the above-mentioned user question information to generate user intention information. Among them, the above-mentioned user intention information can include: an intention tag set. Here, the above-mentioned user question information can be subjected to user intention recognition in various ways to generate user intention information. Each intention tag can represent a kind of intention of the user.
[0032] As an example, the intention tag can be a text tag, such as: "measure", "detect typos", "check", etc.
[0033] In some optional implementation manners of some embodiments, the above-mentioned execution subject performs user intention recognition on the above-mentioned user question information to generate user intention information, including:
[0034] The first step is to extract keywords from the above-mentioned user question information to obtain a group of question keywords. Among them, the stop words corresponding to the preset stop word group in the above-mentioned user question information can be determined for removal. Thus, each word after removing the stop words can be determined as a question keyword, and a group of question keywords is obtained.
[0035] In the second step, extract the problem keywords corresponding to the preset keyword groups from the above problem keyword groups to obtain the target keyword groups, and determine the intention labels corresponding to each target keyword in the above target keyword groups to obtain the first intention label groups. Among them, in response to not receiving the construction document, determine the above first intention label groups as the user intention information. In practice, the preset keyword groups can include keywords preset for the user's questions. For example, the preset basic keywords can include: "measurement instrument", "range finder", etc. Thus, the target keywords that can clearly reflect the user's intention in the target keyword groups can be screened out. Thus, in the case where the user does not upload the construction document, corresponding Q&A results can be given for other intentions of the user.
[0036] In the third step, in response to receiving the construction document and passing the integrity check, determine each of the remaining problem keywords in the above problem keyword groups as the current keyword groups. Among them, each of the remaining problem keywords in the above problem keyword groups can represent keywords related to the construction document. Thus, after the construction document passes the check, the user's question can be answered in combination with the current keywords and the construction document.
[0037] In the fourth step, perform matching processing on each current keyword in the above current keyword groups with a pre-constructed set of local knowledge graphs to obtain a set of matched local knowledge graphs. Among them, each local knowledge graph in the above set of local knowledge graphs corresponds to a preset main entity, and there is no association relationship between the entities corresponding to the local knowledge graphs. Here, for each current keyword, the matching processing can be performed through the following steps: First, the category of the above current keyword can be determined through a classification algorithm (for example, support vector machine). Then, at least one local knowledge graph of the same category can be determined. After that, through a text similarity algorithm (for example, bag-of-words model, Word2Vec word embedding model, etc.), the vector similarity between the preset main entity in each local knowledge graph and the above current keyword can be determined to obtain a set of vector similarities. Here, the local knowledge graph corresponding to the vector similarity that is greater than the preset similarity threshold and is the largest in the set of vector similarities can be determined as the matched local knowledge graph.
[0038] As an example, the preset main entities can be: crane, excavator, warehouse, tunnel, etc.
[0039] In practice, first, considering that constructing a fully quantified knowledge graph not only requires dealing with a large amount of complex relationship extraction and verification work, but also is difficult to maintain and update. For example, each update requires adjusting a large number of associated entity relationships, so it consumes more computing power and time. Therefore, a local knowledge graph is introduced. The local knowledge graph can be constructed based on each preset main entity in the preset main entity group. There is no need to define the association relationship between different preset main entities (i.e., different local knowledge graphs). Thus, when constructing the local knowledge graph, it can be simpler and faster, and at the same time, it can meet the usage requirements. Compared with the fully quantified knowledge graph, each local knowledge graph of the present application can focus on a specific preset main entity, significantly reducing the complexity of construction. Then, the matching of different keywords can be more rapid and convenient. Specifically, the lowest correlation similarity corresponding to each local knowledge graph (i.e., the preset similarity threshold) is determined by the minimum similarity value between each preset main entity and each secondary entity. This ensures that the local knowledge graph can be matched by only matching the preset main entity, without repeatedly matching the secondary entity. Thus, not only can computing power be saved, but also a suitable local knowledge graph can be quickly matched. In addition, although the local knowledge graph of the present application cannot be associated with a large-scale knowledge system, in the step of identifying the user's intention in the present application, it is not necessary to directly identify all the association relationships, that is, there is no need to perform cross-graph queries, so it does not affect the actual use. At the same time, when knowledge graph maintenance is required, only the individual changed entities and their corresponding entity relationships need to be adjusted, without global mobilization. Thus, the pressure of dynamic maintenance can be reduced and computing power can be saved.
[0040] As an example, taking the intelligent tape measure as the preset main entity, the secondary entities may include: laser tape measure, laser distance measurement, length measurement, and construction. The entity relationship between the intelligent tape measure and the laser tape measure: The intelligent tape measure is a type of laser tape measure. The entity relationship between the intelligent tape measure and laser distance measurement: The measurement method used by the intelligent tape measure is laser distance measurement. The entity relationship between the intelligent tape measure and length measurement: The intelligent tape measure has the function of length measurement. The entity relationship between the intelligent tape measure and construction: The intelligent tape measure can be applied to the construction scenario.
[0041] In the fifth step, according to the above-mentioned matched local knowledge graph group, label regression is performed on each current keyword in the above-mentioned current keyword group to generate a second intention label group. Among them, label regression can be performed on each current keyword in the above-mentioned current keyword group through a preset label regression algorithm to generate a second intention label group.
[0042] As an example, the label regression algorithm may include, but is not limited to, at least one of the following: breadth-first search algorithm, depth-first search algorithm, etc. For example, taking "intelligent tape measure" in the local knowledge graph as an example, it is connected to "laser ranging technology" through the "adoption" relationship, and "laser ranging technology". It is also connected to "higher precision" through the "improvement direction" relationship. Using the depth-first search algorithm to search along this path, when the keyword "intelligent tape measure" is encountered, the intent label can be determined as "the direction of improving the laser ranging technology of the intelligent tape measure towards higher precision" according to the searched path information.
[0043] Step 6: Determine the above first intent label group and the above second intent label group as user intent information.
[0044] Step 203: Determine whether it is necessary to call the temporary knowledge base according to the user intent information.
[0045] In some embodiments, the above execution subject may determine whether it is necessary to call the temporary knowledge base according to the above user intent information. Among them, if the intent label representing specification check or the intent label representing document error correction is included in the user intent information, it is determined that it is necessary to call the temporary knowledge base.
[0046] Optionally, if there is an intent label representing modification of construction documents in the user intent information, it is determined that it is necessary to call the temporary knowledge base. Thus, the document can be quickly modified according to the preset logic through the temporary knowledge base.
[0047] Optionally, the above temporary knowledge base is constructed through the following steps:
[0048] Step 1: Extract structured text from each target document in the target document set to generate a structured semantic data set. Among them, the structured text can be extracted through the pdfplumber function and the python-docx library, and the table data is escaped into Markdown format and the semantic relationship is retained to obtain the structured semantic data set.
[0049] Step 2: Divide each structured semantic data in the above structured semantic data set into text blocks to generate a text block set. Among them, the semantic-aware segmentation algorithm can be used: predict the paragraph boundary based on the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture) algorithm, dynamically adjust the block size (128 - 512 tokens), and cooperate with the overlapping window (32 - 128 tokens) to divide each structured semantic data in the above structured semantic data set into text blocks to generate a text block set.
[0050] In practice, the model can be loaded and accelerated through the following code:
[0051] model = AutoModel.from_pretrained("BAAI / bge-large-zh-v1.5",
[0052] device_map = "auto",
[0053] torch_dtype = torch.bfloat16)
[0054] In addition, the model can process 512 text blocks at a time, with a GPU utilization rate of over 90%. By enabling flash_attention optimization, the inference speed can be increased by about 3 times. At the same time, a hybrid index can also be set:
[0055] index = faiss.index_factory(1024,
[0056] "IVF2048_HNSW32,Flat",
[0057] faiss.METRIC_INNER_PRODUCT)
[0058] Thus, the speed of text partitioning by the model is improved.
[0059] In the third step, the above text block set is embedded into a pre-constructed vector database to obtain an added vector database. Among them, the text blocks can be converted into embedding vectors through the OpenAI Embeddings vector function in the langchain library. Then, the embedding vectors can be stored in a preset vector database to obtain an added vector database.
[0060] In the fourth step, a retriever, a retrieval model, and a retrieval Q&A chain set associated with the above added vector database are created to obtain a temporary knowledge base. Among them, through a preset creation code, a retriever, a retrieval model, and a retrieval Q&A chain set associated with the above added vector database are created to obtain a temporary knowledge base.
[0061] As an example, taking the langchain library as an example, through the following code, a retriever, a retrieval model, and a retrieval Q&A chain set associated with the vector database are created and stored in the Chroma vector database to obtain a temporary knowledge base (temporary_knowledge_base):
[0062]
[0063] Step 204, in response to determining the invocation, invoke the temporary knowledge base corresponding to the construction document.
[0064] In some embodiments, the above-mentioned execution entity may, in response to determining the invocation, invoke the temporary knowledge base corresponding to the above-mentioned construction document. Wherein, each construction document may correspond to a document identifier, which characterizes the document type. Thus, the temporary knowledge base with the same document identifier can be invoked.
[0065] Optionally, before using the above-mentioned temporary knowledge base and the above-mentioned user intention information to perform document recognition on the above-mentioned construction document, the above-mentioned execution entity may further include the following steps:
[0066] First step, perform document format conversion on the above-mentioned construction document to obtain a converted construction document. Among them, the file type of the construction document can be determined. For example, PDF (Portable Document Format), txt (Trusted Execution Technology) text format, word format, etc. Then, the word format can be converted to PDF format to obtain a converted construction document. In practice, considering that there are some word documents that cannot be recognized, it is easier to recognize after converting to pdf.
[0067] Second step, perform document segmentation on the above-mentioned converted construction document to generate segmented text segments, and perform text vectorization on the segmented text segments to obtain a text vector set. Among them, the adaptive parameters can be set according to the file size (for example, "chunk_size = the segmentation length of the text in the document, chunk_overlap = the overlapping length of each segment of text after segmentation") to perform text segmentation to obtain segmented text segments. Then, the segmented text segments can be text vectorized through the above-mentioned vector function to obtain a text vector set. Or use the bge-large-zhv1.5 vector model to perform text vectorization on the segmented text segments to obtain a text vector set. In addition, the text vectors can also be stored in the faiss library according to the text storage vector type of the faiss library.
[0068] Third step, in response to determining that the above-mentioned user intention information includes an intention label indicating document error correction, perform typo recognition on each text vector in the above-mentioned text vector set to obtain a document typo identification set. Among them, the optical character recognition (OCR) technology can be used to perform typo recognition on each text vector in the above-mentioned text vector set to obtain a document typo identification set.
[0069] Step 4: Using the above temporary knowledge base, correct the typos corresponding to the document typo identifications in the document typo identification set to obtain a corrected text vector set. Among them, the standard word corresponding to each document typo identification can be matched from the above temporary knowledge base. Here, the word with the highest similarity in the temporary knowledge base can be matched as the standard word through a text similarity algorithm. Then, the text vector of the standard word can be replaced with the text vector corresponding to the document typo identification to obtain a corrected text vector set.
[0070] Step 5: Identify word order errors in the above corrected text vector set to generate a word order error identification set. Among them, the word order errors in the above corrected text vector set can be identified through a preset word order rule set to generate a word order error identification set. Specifically, the part-of-speech order in the corrected text vector can be identified, and the part-of-speech order identification is used as the order sequence. If the order sequence does not conform to the preset word order rules, it is determined that there is a word order error in the corrected text vector, and the identification of the corrected text vector is determined as a word order error identification.
[0071] Here, the word order rules can be preset part-of-speech sorting rules. For example, a word order rule is: the sequential relationship of subject - predicate - object.
[0072] Step 6: Using the above temporary knowledge base, adjust the sentences of the corrected text vectors corresponding to the word order error identifications in the above word order error identification set to obtain an adjusted text vector set. Among them, for each corrected text vector with a word order error, the word order can be adjusted according to the matched word order rules to obtain an adjusted text vector.
[0073] Step 205: Use the temporary knowledge base and user intention information to perform document recognition on the construction document to generate a document recognition result corresponding to the user question information.
[0074] In some embodiments, the above execution subject can use the above temporary knowledge base and the above user intention information to perform document recognition on the above construction document to generate a document recognition result corresponding to the above user question information.
[0075] In some optional implementation manners of some embodiments, the above execution subject uses the above temporary knowledge base and the above user intention information to perform document recognition on the above construction document to generate a document recognition result corresponding to the above user question information, including:
[0076] In response to determining that the above user intention information includes an intention tag indicating specification inspection, perform the following processing steps:
[0077] First step, adjust the format of the above construction document to obtain the adjusted construction document. Among them, the document can be converted from other formats to the txt text format to obtain the adjusted construction document.
[0078] Second step, retrieve the text specification information that matches each adjusted text vector in the above adjusted text vector set from the above temporary knowledge base to obtain the first text specification information set. Among them, the text specification information can include special information representing the standard format of the text. For example, the special information can be instrument operation step information arranged in a preset field order. Here, the standard format can include the format of text typesetting and the restrictive conditions on the corresponding numerical values or descriptions of the text. Secondly, the conditional query code of a preset regular expression can be used to retrieve the text specification information that matches each adjusted text vector in the above adjusted text vector set from the above temporary knowledge base according to the Q&A logic of the temporary database to obtain the first text specification information set. In addition, duplicate specifications can also be removed.
[0079] Third step, extract text specification features from the above adjusted construction document to generate a second text specification information set. Among them, each second text specification information after duplicate removal is included in the above second text specification information set. Here, the above adjusted construction document can be input into a preset large language model for text specification feature extraction to generate a second text specification information set.
[0080] For example, the large language model can include but is not limited to at least one of the following: deepseek model, Zhipu Qingyan (chatglm-4-plus) model, etc.
[0081] Fourth step, proofread the above first text specification information set and the above second text specification information set to generate a proofread text specification information set. Among them, the first text specification information and the second text specification information corresponding to the same specification in the first text specification information set and the second text specification information set can be proofread to determine whether there are errors. If there are no errors, the first text specification information can be determined as the proofread text specification information. If there are errors, it is sent to the terminal for manual review to generate the final proofread text specification information.
[0082] Step 5: Batch review the above-mentioned text specification information set after proofreading to generate a text specification information set after review. Among them, batch review is used to select the text specification information after proofreading with an unexpired specification period as the text specification information after review. Among them, each piece of text specification information after proofreading can correspond to a time period range. For example, 2023 - 2024. In addition, during the batch review process, the document information that is different from the content corresponding to each piece of text specification information after proofreading in the construction document can also be marked. Thus, the text specification information after review generated can include: document paragraph identifier (or attribute node, such as the construction sequence attribute node), whether there is a different anomaly identifier, and the corresponding specification document information, etc.
[0083] In practice, batch review can group all specifications (i.e., text specification information after proofreading) into groups of 7 - 8, and query whether each specification has expired one by one. Specifically, this form is adopted because during the Q&A process, the review of each specification is equivalent to asking a question. When conducting the review, the local temporary knowledge base will provide several relevant pieces of content according to the content of each specification. Too many specifications at one time may lead to more similar content, thus resulting in errors.
[0084] Step 6: Integrate each piece of text specification information after review in the above-mentioned text specification information set after review to generate a document recognition result corresponding to the above-mentioned user question information. Among them, first, the content information of each piece of text specification information after review can be extracted. Then, according to the order of the documents or the page number order of each piece of text specification information after review, the content information can be used as the document recognition result. Thus, it can be used to feedback to the user the document problems existing in the construction document (corresponding to the part where the user raises the question) and the answers to the user's questions.
[0085] In addition, the extracted content information can also be input into a preset large language model for information integration to obtain a document recognition result corresponding to the above-mentioned user question information.
[0086] Optionally, if the text specification information corresponding to a certain adjusted text vector cannot be extracted from the temporary knowledge base, the corresponding text specification information can be retrieved through a preset large language model. Finally, the retrieved text specification information can be stored in the library according to the logic of the temporary knowledge base, so as to realize the dynamic update of the temporary knowledge base.
[0087] In some optional implementation manners of some embodiments, the above-mentioned execution subject uses the above-mentioned temporary knowledge base and the above-mentioned user intention information to perform document recognition on the above-mentioned construction document to generate a document recognition result corresponding to the above-mentioned user question information, and further includes:
[0088] First step, in response to the undetected intent tag, perform document recognition on the above construction document to generate a first recognition result. Among them, in the case of undetected intent tags, it can indicate that the user's needs are not clear. Therefore, the pre-set document structure detection algorithm can be used to extract the document structure of the construction document to generate document structure information. Here, the document structure information can include multiple document titles and corresponding page numbers. Then, the document titles can be used to match the user intent information to obtain the document titles that match the user intent information. Here, the matching can be performed through natural language processing technology. Thus, it can be used to match the document title and corresponding page number with the highest degree of relevance as the first recognition result in the case of unclear user intent.
[0089] Second step, use the above first recognition result and the above user intent information as the current intent information, and use the above temporary knowledge base to generate a second recognition result. Among them, in response to the failure to generate the second recognition result, the above first recognition result is determined as the document recognition result. Here, considering that the scenario frequency of construction document recognition is relatively low, the recognition accuracy can be improved by taking a longer recognition time. Therefore, in order to further improve the accuracy of the feedback to the user's question, for the first recognition result obtained through matching, the above current intent information can be further recognized through the document recognition algorithm pre-deployed in the temporary knowledge base to generate a second recognition result.
[0090] As an example, the above document recognition algorithm can include but is not limited to at least one of the following: EAST (Efficient and Accurate Scene Text Detection) efficient and accurate scene text detection algorithm, CTPN (Connectionist Text Proposal Network) connectionist text extraction network, Faster R-CNN (Faster Region-based Convolutional Neural Network) region-based convolutional neural network.
[0091] In practice, for the questions of users, it is difficult to effectively improve the logic of the local knowledge base if only using large models to answer. In addition, for different industries, there are also differences in the capabilities that large models can recognize. Therefore, in order to better answer users' questions, a temporary knowledge base is set up specifically. Here, in the field of construction, different temporary knowledge bases can be set up for different projects, different equipment, different departments, etc. Thus, the operating logic in the temporary knowledge base can be adjusted specifically to make it more suitable for special Q&A. Therefore, it can be used as the first means for identifying and answering users' question information. Thus, there is no need to distinguish the field where the user's question lies, and the identification logic can be reduced accordingly. Moreover, the special information stored in the temporary knowledge base can be used to analyze users' questions in a more granular way. At the same time, because users have uploaded construction documents, associations can be established between users' questions and the documents through special terms. Thus, the accuracy of identifying users' intentions can be further improved. After that, in the case where the temporary knowledge base fails to identify, the large model can be introduced for final identification, which provides a priori guidance for the adjustment of the temporary knowledge base. This facilitates the dynamic adjustment of the logic and recognition ability of the temporary knowledge base. Thus, the accuracy of the temporary knowledge base in identifying users' question information can be further improved.
[0092] Optionally, refer to Figure 3 , from Figure 3 It can be seen that the use of the knowledge base can be selective. First, it is necessary to determine whether a file is uploaded simultaneously during the user's Q&A. If a file is uploaded synchronously, it can be determined whether the user's Q&A intention is biased towards specification (i.e., specification review, used to adjust problems such as document format and data errors) or towards correcting typos in the document. At the same time, it is determined whether the (temporary) knowledge base needs to be called. If the knowledge base is needed, the corresponding result can be queried from the knowledge base. If no result is found, further answers can be obtained through the large model. Thus, it is ensured that the user's Q&A results can be output. When the knowledge base does not need to be called, typo correction and specification checking can be directly performed to obtain the final result. For example, the final result can be the corrected document or the Q&A feedback information to the user, etc.
[0093] Step 206, send the document recognition result to the user terminal as the user's Q&A result.
[0094] In some embodiments, the above-mentioned execution entity may send the above-mentioned document recognition result to the above-mentioned user terminal as the user's Q&A result.
[0095] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the intelligent recognition method of construction document of some embodiments of the present disclosure, the accuracy of document recognition can be improved, and more accurate question-and-answer results can be provided for users. Specifically, the reasons for the reduction of the accuracy of document recognition results and the difficulty in providing accurate question-and-answer results for users are as follows: The knowledge base set in the existing recognition system is a static knowledge base, which is difficult to dynamically adapt to the changes of industry regulations. At the same time, there is a lack of a special question-and-answer interaction decision-making mechanism for assisting document review. Based on this, in the intelligent recognition method of construction document of some embodiments of the present disclosure, first, in response to receiving user question information uploaded by the user through the user terminal and receiving the synchronously uploaded construction document, wherein the integrity verification of the above-mentioned construction document is performed after receiving. Here, through document integrity verification, the recognition of incomplete documents can be avoided, and the recognition accuracy can be improved. Then, the user intention of the above-mentioned user question information is recognized to generate user intention information, wherein the above-mentioned user intention information includes: an intention tag set. Here, through user intention recognition, the intention of the user can be determined, so as to perform targeted document recognition, so as to provide more accurate question-and-answer results for users. After that, according to the above-mentioned user intention information, it is determined whether it is necessary to call the temporary knowledge base. Here, by establishing a temporary knowledge base, special question-and-answer results can be generated. Then, in response to determining to call, the temporary knowledge base corresponding to the above-mentioned construction document is called. Next, the above-mentioned construction document is recognized by using the above-mentioned temporary knowledge base and the above-mentioned user intention information to generate a document recognition result corresponding to the above-mentioned user question information. Thus, a special question-and-answer interaction decision-making mechanism for user questions is realized. Thereby, the accuracy of generating the document recognition result can be improved. Finally, the above-mentioned document recognition result is sent to the above-mentioned user terminal as the user question-and-answer result. Furthermore, more accurate question-and-answer results can be provided for users.
[0096] Further referring to Figure 4 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an intelligent recognition device for construction documents. These device embodiments correspond to Figure 2 the method embodiments shown, and the intelligent recognition device for construction documents can be specifically applied to various electronic devices.
[0097] As Figure 4As shown, the intelligent identification device 400 for construction documents in some embodiments includes: a receiving unit 401, an intention recognition unit 402, a determination unit 403, a calling unit 404, a document recognition unit 405, and a sending unit 406. Among them, the receiving unit 401 is configured to respond to receiving user question information uploaded by the user through the user terminal and the construction document uploaded synchronously, and perform integrity verification on the above construction document after receiving it; the intention recognition unit 402 is configured to perform user intention recognition on the above user question information to generate user intention information, where the above user intention information includes: an intention tag set; the determination unit 403 is configured to determine whether it is necessary to call the temporary knowledge base according to the above user intention information; the calling unit 404 is configured to respond to the determination to call and call the temporary knowledge base corresponding to the above construction document; the document recognition unit 405 is configured to use the above temporary knowledge base and the above user intention information to perform document recognition on the above construction document to generate a document recognition result corresponding to the above user question information; the sending unit 406 is configured to send the above document recognition result as a user Q&A result to the above user terminal.
[0098] It can be understood that the various units described in the intelligent identification device 400 for construction documents correspond to the respective steps in the method described in the reference Figure 2 Therefore, the operations, features, and beneficial effects described above for the method also apply to the intelligent identification device 400 for construction documents and the units included therein, and will not be repeated here.
[0099] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device (such as the computing device 101 shown in Figure 1 ) suitable for implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present disclosure. As shown in Figure 5 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the intelligent identification methods for construction documents. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the intelligent identification methods for construction documents. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some structures related to the present disclosure solution, and does not constitute a limitation on the computer device to which the present disclosure solution is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0100] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0101] Among them, in one embodiment, the above-mentioned processor is used to run a computer program stored in the memory to implement the following steps: in response to receiving user question information uploaded by the user through the user terminal and receiving the synchronously uploaded construction document, where the integrity verification of the above-mentioned construction document is performed after receiving; perform user intent recognition on the above-mentioned user question information to generate user intent information, where the above-mentioned user intent information includes: an intent tag set; determine whether it is necessary to call the temporary knowledge base according to the above-mentioned user intent information; in response to determining to call, call the temporary knowledge base corresponding to the above-mentioned construction document; use the above-mentioned temporary knowledge base and the above-mentioned user intent information to perform document recognition on the above-mentioned construction document to generate a document recognition result corresponding to the above-mentioned user question information; send the above-mentioned document recognition result as a user question-and-answer result to the above-mentioned user terminal.
[0102] The embodiments of the present disclosure also provide a computer-readable storage medium. A computer program is stored on the above-mentioned computer-readable storage medium, and the computer program includes program instructions. The method implemented when the above-mentioned program instructions are executed may refer to the various embodiments of the intelligent recognition method for construction documents of the present disclosure.
[0103] Among them, the above computer-readable storage medium may be an internal storage unit of the computer device in the foregoing embodiment, such as the hard disk or memory of the computer device. The above computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.
[0104] It should be noted that in this document, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0105] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.
Claims
1. An intelligent recognition method for construction documents, characterized in that, Including: In response to receiving user question information uploaded by a user through a user terminal and receiving a synchronously uploaded construction document, wherein, after receiving, integrity verification is performed on the construction document; Perform user intention recognition on the user question information to generate user intention information, wherein the user intention information includes: an intention tag set; Determine whether it is necessary to call a temporary knowledge base according to the user intention information; In response to determining to call, call the temporary knowledge base corresponding to the construction document; Use the temporary knowledge base and the user intention information to perform document recognition on the construction document to generate a document recognition result corresponding to the user question information; Send the document recognition result as a user Q&A result to the user terminal.
2. The method according to claim 1, characterized in that The method further includes: In response to receiving user question information uploaded by a user through a user terminal and not receiving a construction document, call a preset fixed knowledge base according to the intention tag set; Answer questions for the intention tag set through the fixed knowledge base to generate a user Q&A result; In response to determining that the question answering fails, call a preset language model for secondary answering to obtain a user Q&A result.
3. The method according to claim 1, wherein The temporary knowledge base is constructed through the following steps: Extract structured text from each target document in the target document set to generate a structured semantic data set; Divide each structured semantic data in the structured semantic data set into text blocks to generate a text block set; Embed the text block set into a pre-constructed vector database to obtain an added vector database; Create a retriever, a retrieval model, and a retrieval Q&A chain set associated with the added vector database to obtain a temporary knowledge base.
4. The method according to claim 1, wherein Before using the temporary knowledge base and the user intention information to perform document recognition on the construction document, the method further includes: Convert the document format of the construction document to obtain a converted construction document; Perform document segmentation on the converted construction document to generate segmented text segments, and perform text vectorization on the segmented text segments to obtain a text vector set; In response to determining that the user intention information includes an intention tag representing document error correction, perform misspelling recognition on each text vector in the text vector set to obtain a document misspelling identification set; Use the temporary knowledge base to correct the misspelled words corresponding to the document misspelling identifications in the document misspelling identification set to obtain a corrected text vector set; Perform word order error recognition on the corrected text vector set to generate a word order error identification set; Use the temporary knowledge base to adjust the sentences of the corrected text vectors corresponding to the word order error identifications in the word order error identification set to obtain an adjusted text vector set.
5. The method according to claim 4, wherein The using the temporary knowledge base and the user intention information to perform document recognition on the construction document to generate a document recognition result corresponding to the user question information includes: In response to determining that the user intention information includes an intention tag representing specification inspection, perform the following processing steps: Adjust the document format of the construction document to obtain an adjusted construction document; Retrieve text specification information that matches each adjusted text vector in the adjusted text vector set from the temporary knowledge base to obtain a first text specification information set; Extract text specification features from the adjusted construction document to generate a second text specification information set, where each second text specification information in the second text specification information set is de-duplicated; Perform specification proofreading on the first text specification information set and the second text specification information set to generate a proofread text specification information set; Perform batch review on the proofread text specification information set to generate a reviewed text specification information set, where batch review is used to select proofread text specification information with an unexpired specification period as the reviewed text specification information; Integrate each reviewed text specification information in the reviewed text specification information set to generate a document recognition result corresponding to the user question information.
6. The method according to claim 5, wherein The using the temporary knowledge base and the user intention information to perform document recognition on the construction document to generate a document recognition result corresponding to the user question information further includes: In response to not detecting an intention tag, perform document recognition on the construction document to generate a first recognition result; Use the first recognition result and the user intention information as the current intention information, and use the temporary knowledge base to generate a second recognition result, where in response to not generating a second recognition result, determine the first recognition result as the document recognition result.
7. The method according to claim 1, characterized in that The performing user intention recognition on the user question information to generate user intention information includes: Extract keywords from the user question information to obtain a question keyword group; Extract the question keywords corresponding to the preset keyword group from the question keyword group to obtain a target keyword group, and determine the intention tags corresponding to each target keyword in the target keyword group to obtain a first intention tag group, where in response to not receiving a construction document, determine the first intention tag group as the user intention information; In response to receiving a construction document and the document passing the verification, determine each remaining question keyword in the question keyword group as the current keyword group; Perform matching processing on each current keyword in the current keyword group with a pre-constructed set of local knowledge graphs to obtain a set of matched local knowledge graphs, where each local knowledge graph in the set of local knowledge graphs corresponds to a preset main entity, and there is no association relationship between the entities corresponding to the local knowledge graphs; Perform label regression on each current keyword in the current keyword group according to the set of matched local knowledge graphs to generate a second intention tag group; Determine the first intention tag group and the second intention tag group as the user intention information.
8. An intelligent recognition device for construction documents, comprising: A receiving unit, configured to, in response to receiving user question information uploaded by a user through a user terminal and receiving a synchronously uploaded construction document, perform integrity verification on the construction document after receiving it; An intention recognition unit, configured to perform user intention recognition on the user question information to generate user intention information, where the user intention information includes: an intention tag set; A determination unit, configured to determine whether it is necessary to call a temporary knowledge base according to the user intention information; A calling unit, configured to, in response to determining to call, call the temporary knowledge base corresponding to the construction document; A document recognition unit, configured to use the temporary knowledge base and the user intention information to perform document recognition on the construction document to generate a document recognition result corresponding to the user question information; A sending unit, configured to send the document recognition result as a user question and answer result to the user terminal.
9. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Automatic graph auditing system based on BIM and knowledge graph
CN115687649A
Bridge field construction scheme examination method based on large model and knowledge graph
CN118411016A
Multi-level domain knowledge question-answering method and device based on large model
CN119202213A
Customer service interaction method, interaction device, equipment, storage medium and program product
CN119646137A