Building document information extraction method and system and medium

By dividing architectural documents into text and table areas, extracting and managing data separately, and determining consistency status based on user input, the complexity and consistency issues of architectural document information extraction in existing technologies are resolved, thereby improving the accuracy and reliability of information extraction.

CN121786007APending Publication Date: 2026-04-03TECHNOLOGY (CHENGDU) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-13
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing architectural document information extraction technologies are ill-suited to complex and ever-changing text structures and lack the ability to process information across chapters and pages. In particular, they lack the ability to uniformly associate and verify the consistency of content in different formats, especially in scenarios where technical and business documents coexist.

Method used

By dividing architectural documents into text and table regions, extracting and storing text and table data respectively, and retrieving corresponding data from the text and table databases based on user input, the consistency status is determined, and the target output is generated.

Benefits of technology

It enables separate management of unstructured text data and structured tabular data, improving the accuracy and efficiency of information extraction, reducing the cost of error identification and compliance risks, and enhancing the credibility of the output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786007A_ABST
    Figure CN121786007A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information extraction, and provides a building document information extraction method and system and a medium, and the method comprises the steps: obtaining an original file of a building document; scanning an original file, and dividing the original file into a text region and a table region; extracting text data of the text area, and storing the text data into a text database; extracting table data of the table area, and storing the table data into a table database; based on the user input, target text data and target table data corresponding to the user input are retrieved in a text database and a table database; and determining a consistency state of the target text data and the target table data, and generating target output. According to the method and the device, separated management of the unstructured text data and the structured table data is realized, so that the accuracy and the processing efficiency of information extraction are improved, the text data and the table data are compared and verified, and the credibility of target output is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information extraction technology, and in particular to a method, system and medium for extracting information from architectural documents. Background Technology

[0002] With the continuous advancement of digital transformation in the construction engineering field, intelligent parsing and key information extraction of construction documents (such as bidding documents and construction documents) have become an important foundation for project cost management, contract management, and risk control. Construction documents typically include various content formats such as technical documents, commercial documents, and bills of quantities, containing both a large amount of unstructured text and structured tabular data. Accurately extracting and annotating key information such as construction techniques, material specifications, and technical parameters from these different formats is a crucial step in achieving subsequent automated processing and decision support.

[0003] Existing architectural document information extraction technologies mainly include rule-based or template-based extraction methods and retrieval enhancement generation methods based on general large language models. The former relies on fixed formats or predefined rules, making it difficult to adapt to the complex and varied text structures in architectural documents, and its ability to handle information association across chapters and pages is limited; while the latter can improve the information retrieval effect of unstructured text, when faced with scenarios where technical and business documents coexist, it usually only extracts information from a single content format, lacking the ability to uniformly associate and verify the consistency of content from different formats.

[0004] Therefore, there is an urgent need for a method, system, and medium for extracting information from architectural documents, so as to extract, store, and output content in different formats from architectural documents separately, and to establish a corresponding mapping relationship between different formats of content to ensure the consistency and reliability of the extraction results in engineering logic. Summary of the Invention

[0005] To effectively extract complex and ever-changing information from bidding documents, this invention provides a method, system, and medium for extracting information from architectural documents.

[0006] The invention includes a method for extracting information from architectural documents. The method includes: acquiring the original file of the architectural document, the original file including text content and table content; scanning the original file and dividing it into text regions and table regions; extracting text data from the text regions and storing the text data in a text database; extracting table data from the table regions and storing the table data in a table database; based on user input, retrieving target text data and target table data corresponding to the user input from the text database and the table database; determining the consistency status of the target text data and the target table data, and generating a target output.

[0007] The invention includes an information extraction system for architectural documents. The system comprises: an acquisition module configured to acquire the original file of the architectural document, the original file including text content and table content; and a processing module configured to: scan the original file, dividing it into text regions and table regions; extract text data from the text regions and store the text data in a text database; extract table data from the table regions and store the table data in a table database; based on user input, retrieve target text data and target table data corresponding to the user input from the text database and the table database; determine the consistency status of the target text data and the target table data, and generate a target output.

[0008] The invention includes a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the information extraction method for building documents described in the above embodiments.

[0009] The beneficial effects of this invention include, but are not limited to: (1) By using appropriate methods to extract content from different types of content in the original file, separate management of unstructured text data and structured tabular data is achieved, thereby improving the accuracy and efficiency of information extraction; (2) Based on user input, corresponding text data and tabular data are retrieved from the text database and tabular database respectively, and the text data and tabular data are compared and verified, which effectively improves the credibility of the target output and reduces the cost of error identification and compliance risk in the information extraction process of building documents; (3) Different target outputs are generated according to the consistency status and user selection, which can reduce the risk of erroneous target outputs. By reusing historical user selection results in subsequent conflict scenarios, user operations can be reduced, the efficiency of target output generation can be improved, and the system can gradually form an output decision mode that conforms to specific business scenarios or user preferences. Attached Figure Description

[0010] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0011] Figure 1 This is a schematic diagram illustrating an application scenario of the information extraction method for architectural documents according to some embodiments of this specification; Figure 2 This is an exemplary system block diagram of an information extraction system for architectural documents according to some embodiments of this specification; Figure 3This is an exemplary flowchart of an information extraction method for architectural documents according to some embodiments of this specification; Figure 4 This is an exemplary flowchart illustrating the retrieval of target text data and target tabular data according to some embodiments of this specification; Figure 5 These are exemplary schematic diagrams illustrating the generation of the target output according to some embodiments of this specification; Figure 6 This is an exemplary flowchart illustrating the generation of target confidence scores for target outputs according to some embodiments of this specification. Detailed Implementation

[0012] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0013] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0014] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0015] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0016] Figure 1 This is a schematic diagram illustrating an application scenario of the information extraction method for architectural documents according to some embodiments of this specification.

[0017] In some embodiments, such as Figure 1As shown, the application scenario 100 of the information extraction method for architectural documents may include a network 110, a user terminal 120, a processor 130, a database 140, and a third-party platform 150. In some embodiments, architectural documents may be documents related to the construction industry, such as tender documents, construction documents, industry standard documents, etc.

[0018] Network 110 may include any suitable network capable of facilitating information and / or data exchange. In some embodiments, at least one component of the application scenario 100 of the building document information extraction method (e.g., user terminal 120, processor 130, database 140, and third-party platform 150, etc.) can exchange information and / or data with at least one other component in the application scenario 100 of the building document information extraction method via network 110. For example, processor 130 can send text data, tabular data, etc., to database 140 for storage via network 110.

[0019] In some embodiments, network 110 can be any one or more of a wired network or a wireless network. Network 110 may include one or more network access points.

[0020] User terminal 120 can be a terminal device used by the user (such as at least one of mobile phone 121, tablet 122, and computer 123). The user can be a person who needs to extract information from the building documents, such as the bidding personnel, cost estimators, and project managers of the bidding party (construction unit, etc.), or other relevant users.

[0021] Processor 130 can process data and / or information obtained from at least one component of the application scenario 100 of the information extraction method for architectural documents. Processor 130 can execute program instructions based on this data, information, and / or processing results to perform one or more functions described herein. In some embodiments, processor 130 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 130 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), a microprocessor, or any combination thereof.

[0022] Database 140 may store data, instructions, and / or any other information related to the extraction of information from building documents. In some embodiments, database 140 may store data and / or information sent by processor 130.

[0023] In some embodiments, database 140 may include one or more storage units, each of which may be a separate device or part of another device. In some embodiments, database 140 may be implemented on a cloud platform. In some embodiments, database 140 may be part of processor 130.

[0024] In some embodiments, database 140 may include text databases and tabular databases, etc., to store text data and tabular data respectively.

[0025] Third-party platform 150 may be a platform for obtaining the original files of architectural documents. In some embodiments, third-party platform 150 may be a platform for publishing architectural documents. Third-party platform 150 may be a server, etc.

[0026] In some embodiments, the application process of the building document information extraction method includes: the processor 130 obtains and scans the original building document from the third-party platform 150 via the network 110, then extracts text data and table data from the original document, and sends the text data and table data to the database 140 for storage via the network 110. After receiving user input from the user terminal 120, the processor 130 retrieves the data corresponding to the user input from the database 140 and checks the consistency of the data, and outputs the retrieved data to the user terminal 120.

[0027] For a detailed explanation of the above content, please refer to [link / reference]. Figures 2-6 And its related descriptions.

[0028] Figure 2 This is an exemplary system block diagram of an information extraction system for architectural documents according to some embodiments of this specification.

[0029] In some embodiments, the building document information extraction system 200 may include an acquisition module 210 and a processing module 220.

[0030] The acquisition module 210 is configured to acquire the original files of the building documents.

[0031] The processing module 220 is configured to: scan the original file and divide it into text regions and table regions; extract text data from the text regions and store the text data in a text database; extract table data from the table regions and store the table data in a table database; based on user input, retrieve target text data and target table data corresponding to the user input from the text database and the table database; determine the consistency status of the target text data and target table data, and generate the target output.

[0032] In some embodiments, the acquisition module and the processing module may be configured in a processor and / or server. The processor and / or server may process the acquired data and / or information, and execute program instructions based on the data, information, and / or processing results to perform one or more functions described herein. In some embodiments, the processor and / or server may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction processor (ASIP), or any combination thereof.

[0033] For a detailed explanation of the above content, please refer to [link / reference]. Figures 3-6 And its related descriptions.

[0034] Figure 3 This is an exemplary flowchart of an information extraction method for architectural documents according to some embodiments of this specification. Figure 3 As shown, process 300 includes steps 310-360 as described below. In some embodiments, process 300 may be executed by a processor (such as processor 130).

[0035] Step 310: Obtain the original files of the architectural documents.

[0036] Original documents refer to the initial, unprocessed, original content of architectural documents, such as technical documents and commercial documents. For example, when architectural documents include architectural tenders, the original document can refer to the unprocessed content of the tender issued by the tendering party (construction unit, etc.). In some embodiments, the processor can obtain the original documents of architectural documents through a third-party platform (the platform for publishing architectural documents). For example, the processor can obtain the original documents of architectural tenders through the platform where the tendering party publishes architectural tenders. As another example, the processor can obtain the original documents of industry standard documents through a national publication platform.

[0037] Technical documents can be documents that explain technical aspects such as project implementation plans, technical routes, construction techniques, and quality assurance measures. Commercial documents can be documents that explain commercial aspects such as bills of quantities, contract terms, qualifications, and performance capabilities.

[0038] In some embodiments, the original document includes at least two of the following: text content, table content, and drawing content. It is understood that technical documents may primarily consist of text content and drawing content, while business documents may primarily consist of table content.

[0039] Text content refers to data or information presented in natural language. Examples include construction plan descriptions, construction techniques, and material specifications in technical documents.

[0040] Table content refers to data or information presented in tabular form. Examples include bills of quantities and detailed price lists in business documents.

[0041] Drawing content refers to data or information presented in the form of drawings, such as construction drawings in technical documents.

[0042] Step 320: Scan the original document and divide it into text areas and table areas.

[0043] A text area can be a region in the original file that contains text content. A drawing area can be a region in the original file that contains drawing content. A table area can be a region in the original file that contains table content.

[0044] In some embodiments, the processor can scan the original document and divide it into text areas, table areas, and drawing areas in various ways. Examples include coordinate positioning with regular expressions, template matching, and layout analysis algorithms.

[0045] In some embodiments, the processor scans the original document using a first model to obtain a first feature corresponding to the text content and a second feature corresponding to the table content, and determines the text region based on the first feature and the table region based on the second feature. When the original document also includes drawing content, the processor obtains a third feature corresponding to the drawing content using the first model and determines the drawing region based on the third feature.

[0046] In some embodiments, the first model may include one or more combinations of a deep learning-based Large Language Model (LLM), such as the GPT model, the BERT model, or other custom models.

[0047] The first feature refers to the feature that characterizes the region containing text content. In some embodiments, the first feature includes the coordinates of the region's extent, such as the text block. The text block may be a region that covers text content in the original file.

[0048] The second feature refers to the characteristics that characterize the region containing the table content. In some embodiments, the second feature includes the coordinates of the area bounded by the table borders, etc. The table borders can be the borders that constitute the table content in the original file. If a text block falls within the area enclosed by the table borders, the processor includes that text block in the table area.

[0049] The third feature refers to the feature that characterizes the area containing the drawing content. In some embodiments, the third feature includes the coordinates of the area range of an image block, etc. The image block may be an area that covers the drawing content in the original file.

[0050] In some embodiments, the processor may determine the region corresponding to the first feature as a text region, the region corresponding to the second feature as a table region, and the region corresponding to the third feature as a drawing region.

[0051] In some embodiments, the processor may use the coordinates of the text block's area range as the coordinate information of the text region, and the coordinates of the table frame's area range as the coordinate information of the table region.

[0052] In some embodiments of this specification, the processor scans the original file using a first model, dividing the original file into different regions according to content type, so that the content in each region can be extracted and queried separately.

[0053] Step 330: Extract the text data from the text region and store the text data in the text database.

[0054] Text data refers to the specific text data or information within a text area.

[0055] A text database is a database used to store and manage text data.

[0056] In some embodiments, the processor segments the text content within a text region into multiple independent text data based on semantic integrity or semantic delimiters (such as periods, paragraph ends, etc.) and stores them in a text database.

[0057] In some embodiments, the processor can extract drawing data from the drawing area using technologies such as Optical Character Recognition (OCR). The drawing data includes legends and corresponding images. The processor can store the drawing data in a text database.

[0058] In some embodiments, the processor can store text data and the coordinate information of the text region together in a text database.

[0059] Step 340: Extract the table data from the table area and store the table data in the table database.

[0060] Table data refers to the specific table data or information within a table area.

[0061] A tabular database is a database used to store and manage tabular data.

[0062] In some embodiments, the processor may determine one or more data names based on one or more header rows in a table area; determine data values ​​corresponding to one or more data names based on the data rows corresponding to one or more header rows; determine structured or semi-structured table data based on one or more data names and their corresponding data values; and store the structured or semi-structured table data in a table database.

[0063] The header row can be a row used to define the attributes or meaning of the data in each column of the table.

[0064] In some embodiments, for each table within a table area, the processor may treat the first row of the table as the header row and the remaining rows as data rows.

[0065] In some embodiments, the table area includes a first table and a second table spanning multiple pages.

[0066] The first table refers to a standalone table within the current page of the table region. The second table refers to a standalone table within the next page of the current page. It is understandable that the table region determined by the processor may cover two or more pages of the original document. In this case, the processor can analyze page by page within the table region to determine whether the standalone tables on adjacent pages are the same table.

[0067] In some embodiments, the processor identifies the header row of a first table (denoted as the first header row) and the header row of a second table (denoted as the second header row); determines the similarity between the header rows of the first table and the second table; when the similarity meets the similarity condition, the first table and the second table are determined to be the same table; when the similarity does not meet the similarity condition, the first table and the second table are determined to be different tables. The processor extracts data from the same table into one set of table data, and extracts data from different tables into two sets of table data respectively.

[0068] In some embodiments, the processor may construct a first vector based on the content of a first header row, construct a second vector based on the content of a second header row, and calculate the similarity (such as cosine similarity) between the first vector and the second vector as the similarity between the first header row and the second header row.

[0069] The similarity condition refers to the condition used to determine whether the first table and the second table are the same table. In some embodiments, the similarity condition includes that the similarity between vectors is not less than a first similarity threshold, which can be set by those skilled in the art based on experience.

[0070] In some embodiments, if the similarity between the header rows does not meet the similarity condition, the processor may calculate whether the data row content corresponding to each unit in the first header row is similar to the data row content corresponding to the unit in the same position order in the second header row. If the similarity between the data row content corresponding to each unit in the first header row and the data row content corresponding to the unit in the same position order in the second header row both meet the similarity condition, the processor will regard the first table and the second table as the same table; otherwise, the processor will regard the first table and the second table as different tables.

[0071] For example, for each cell in the first header row, the processor can construct a third vector based on the data row content corresponding to that cell, and construct a fourth vector based on the data row content of cells in the second header row that have the same positional order as that cell. The similarity between the third vector and the fourth vector is then calculated as the similarity between the two data row contents. Through this method, the processor determines whether the data row content corresponding to each cell in the first header row is similar to the data row content corresponding to cells in the second header row with the same positional order.

[0072] It is understandable that if the first table and the second table are the same table, then the header row of the second table may be a continuation of the data row content corresponding to each cell in the header row of the first table. Therefore, if the data row content corresponding to each cell in the header row of the first table is similar to the data row content corresponding to the cells in the same position in the header row of the second table, it means that the header row of the second table is a continuation of the data row content corresponding to each cell in the header row of the first table, and the first table and the second table are the same table.

[0073] In some embodiments of this specification, the processor determines whether two tables are the same table based on the similarity of the header row content and the similarity of the data row content in a cross-page scenario. This ensures the continuity and integrity of cross-page table data and improves the reliability of table data extraction and storage.

[0074] A data name is a name used to identify the meaning of the data in a data row below the header row.

[0075] In some embodiments, the processor can identify the text content of each cell in each header row as a data name. For example, if a header row contains cells with text content of “Project Name”, “Specification”, and “Unit” in sequence, the processor will identify “Project Name”, “Specification”, and “Unit” as three data names respectively.

[0076] Data values ​​can be specific data or information within a data row.

[0077] In some embodiments, the processor uses the text content corresponding to a cell in the header row as the data value corresponding to the target data name (the data name corresponding to that cell) based on the data row corresponding to each header row. For example, if there is a cell in the header row with the text content "Project Name", then the data name corresponding to that cell is "Project Name". If the text content corresponding to that cell in multiple data rows is "Exterior Paint" and "Interior Paint" respectively, then the data value corresponding to "Project Name" is "Exterior Paint" and "Interior Paint".

[0078] Structured tabular data can be tabular data with a fixed data storage structure, such as tabular data where data names and corresponding data values ​​are stored in a fixed order and fixed position.

[0079] Semi-structured tabular data can be tabular data with only partial structured features. In some embodiments, the data names and corresponding data values ​​in semi-structured tabular data can be organized through key-value relationships or hierarchical relationships, such as JSON style.

[0080] For example, semi-structured tabular data can be represented as: ("Project Name": "Exterior Paint", "Material Type": "Stone Paint"), where there are keys K1 (Project Name) and K2 (Material Type), where the key value of key K1 is "Exterior Paint" and the key value of key K2 is "Stone Paint".

[0081] In some embodiments, the processor may store data names and corresponding data values ​​in a structured or semi-structured tabular data format as needed to obtain structured or semi-structured tabular data, and store the structured or semi-structured tabular data in a tabular database.

[0082] In some embodiments, the processor can store both table data and coordinate information of the table area in a table database.

[0083] In some embodiments of this specification, structured or semi-structured tabular data is generated and stored based on data names and corresponding data values, thereby achieving standardized storage of tabular data, enhancing the understandability and retrieval of tabular data, and reducing subsequent data processing costs.

[0084] Step 350: Based on user input, retrieve the target text data and target table data corresponding to the user input from the text database and the table database.

[0085] User input refers to text information provided by the user for querying information. In some embodiments, the user can input text information through a user terminal, which then sends the text information to the processor.

[0086] Target text data refers to text data retrieved from a text database based on user input.

[0087] Target table data refers to table data retrieved from a table database based on user input.

[0088] In some embodiments, the processor can determine the target text data and target table data based on user input through various methods. For example, based on user input, the processor can retrieve the target text data and target table data corresponding to the user input from text databases and table databases through methods such as character matching.

[0089] In some embodiments, the processor can also determine keywords based on user input, and determine target text data and target table data based on the keywords. See also: [link to related content] Figure 4 And its related descriptions.

[0090] Step 360: Determine the consistency status of the target text data and the target table data, and generate the target output.

[0091] A consistency state refers to the degree of semantic consistency between target text data and target table data. In some embodiments, consistency states include consistency, conflict, and missing. A missing state can mean that the processor retrieves only the target text data or the target table data.

[0092] In some embodiments, the processor can determine the consistency status of the target text data and the target table data in various ways. For example, the processor constructs feature vectors based on the target text data and the target table data respectively, and calculates the similarity between the feature vectors. If the similarity is not less than a second similarity threshold, the processor determines the consistency status as consistent. If the similarity is less than the second similarity threshold, the processor determines the consistency status as conflicting. The second similarity threshold can be set by those skilled in the art based on experience.

[0093] The target output refers to the output result used to respond to user input. The processor can send the target output to the user terminal.

[0094] In some embodiments, when the consistency state is consistent, the processor uses both target text data and target table data as target output. When the consistency state is conflicting, the processor can determine either target text data or target table data as target output based on user selection. When the consistency state is missing, the processor can use either the retrieved target text data or target table data as target output.

[0095] In some embodiments, the processor may also generate the target output based on text objects, table objects, and the consistency state between text objects and table objects. See also: [link to relevant content] Figure 5 And its related descriptions.

[0096] In some embodiments, the processor binds the target output to the coordinate information of the text area and the table area. Users can directly view the original text within the text area and table area by performing interactive operations such as clicking on the target output.

[0097] In some embodiments of this specification, the generated target output is bound to the coordinate information of the corresponding text area and table area, thereby retaining the corresponding original text position index while generating the target output, which makes it easier for users to quickly locate the original text in the text area and table area, and improves the traceability of information verification and the credibility of the results.

[0098] In some embodiments of this specification, by employing appropriate methods to extract content from different types of content in the original file, separate management of unstructured text data and structured tabular data is achieved, thereby improving the accuracy and efficiency of information extraction. Furthermore, based on user input, corresponding text and tabular data are retrieved from the text and tabular databases respectively, and the text and tabular data are compared and verified, effectively enhancing the credibility of the target output and reducing the cost of error identification and compliance risks in the information extraction process of architectural documents.

[0099] Figure 4 This is an exemplary flowchart illustrating the retrieval of target text data and target tabular data according to some embodiments of this specification. Figure 4 As shown, process 400 includes the following steps 410-440.

[0100] Step 410: Determine keywords based on user input.

[0101] For more information on user input, please see [link / reference]. Figure 3 And its related descriptions.

[0102] Keywords can be words or phrases in the user's input that represent the user's core concerns. For example, if the user inputs "paint the exterior walls", the keyword could be "exterior wall paint".

[0103] In some embodiments, the processor may, based on a knowledge graph, retrieve phrases from user input that are identical or similar to standard fields corresponding to nodes in the knowledge graph as keywords.

[0104] A knowledge graph is a graph structure that uses knowledge related to the field of architecture as its objects and represents them in a structured way through the relationships between these objects.

[0105] In some embodiments, the processor uses architectural terminology as a core node (e.g., Figure 5 Node 571 in the document sets attribute information related to technical terms as attribute nodes (e.g., ...). Figure 5 Node 572 in the middle), and through the edge (such as Figure 5 Edge 573 in the graph represents the semantic relationship between core nodes and attribute nodes, thus constructing a knowledge graph. Each core node and attribute node corresponds to a standard field.

[0106] Standard fields can be standard terms related to the construction industry.

[0107] For example, the knowledge graph contains a core node with the standard field "waterproof coating". This core node is connected to the attribute node with the standard field "waterproof" via an edge representing "functional attribute", and to the attribute node with the standard field "coating" via an edge representing "material type". If the user input is "use waterproof paint for painting", the processor can use "waterproof paint", which is semantically similar to "waterproof coating", as a keyword in the user input.

[0108] Step 420: Based on keywords, determine text-related terms and table-related terms.

[0109] Textual conjunctions are words related to keywords and used to retrieve target text data. For example, textual conjunctions could be "functional attributes" or "material types".

[0110] Table-related terms are words that are related to keywords and used to retrieve target table data. For example, table-related terms could be "total material quantity" or "material unit price".

[0111] In some embodiments, the processor can determine text-related terms and table-related terms based on a local database or knowledge graph. For example, the processor finds nodes in the knowledge graph that correspond to standard fields with semantically similar keywords, and determines the standard fields corresponding to the nodes connected to those nodes as text-related terms or table-related terms. Text-related terms are mainly related to implementation plans and construction techniques, while table-related terms are mainly related to quantities of work and detailed pricing.

[0112] Step 430: Determine the text query statement based on text association words, and retrieve the target text data from the text database based on the text query statement.

[0113] For more information on text databases and target text data, see [link to relevant information]. Figure 3 And its related descriptions.

[0114] A text query statement is a statement used to retrieve target text data. In some embodiments, a text query statement may consist of keywords and textual conjunctions. For example, a text query statement may be "functional attributes and material types of exterior wall coatings".

[0115] In some embodiments, the processor determines the target feature vector based on the text query statement; determines the similarity between multiple text feature vectors and the target feature vector based on multiple text feature vectors in the text database; and selects a preset number of text data as target text data based on the similarity.

[0116] In some embodiments, the processor determines the target feature vector based on a text query statement through various methods. For example, the processor may determine the target feature vector by querying an existing word vector library in the architecture domain based on the text query statement. Alternatively, the processor may determine the target feature vector based on the text query statement using a pre-trained word-to-vector model (such as the Word2Vec model).

[0117] In some embodiments, the processor can determine the text feature vector corresponding to each text data based on multiple text data in a text database. The method for determining the text feature vector is similar to the method for determining the target feature vector, and will not be described in detail here.

[0118] In some embodiments, the processor may calculate the cosine similarity between a single text feature vector and a target feature vector as the similarity between the single text feature vector and the target feature vector.

[0119] In some embodiments, the processor sorts the similarity between multiple text feature vectors and target feature vectors in descending order, selects a preset number (e.g., 5) of text feature vectors sequentially from the top of the sorting results, and uses the text data corresponding to the selected multiple text feature vectors as the target text data.

[0120] In some embodiments of this specification, text data is filtered based on the similarity between vectors, making the text data retrieval process based on semantic features rather than just keyword matching. This enables more accurate identification of text information related to content of interest to the user, further improving the accuracy and relevance of text data retrieval.

[0121] Step 440: Determine the table query statement based on the table related terms, and retrieve the target table data from the table database based on the table query statement.

[0122] For more information on tabular databases and target tabular data, please see [link to relevant documentation]. Figure 3 And its related descriptions.

[0123] A table query statement is a statement used to retrieve data from a target table. In some embodiments, a table query statement may consist of keywords and table related terms. For example, a table query statement could be "total amount of exterior wall coating materials, unit price of materials".

[0124] In some embodiments, the processor may retrieve target table data from the table database based on the table query statement, using methods such as exact matching or regular expression matching.

[0125] For example, for semi-structured tabular data, the processor performs a search and match for "exterior wall coating", "total material quantity", and "unit price of material" to determine the target tabular data as ("material type": "real stone paint", "total material quantity": "45 tons", "unit price of material": "45 yuan / square meter").

[0126] For example, for structured tabular data, the processor first determines the corresponding data row based on the keyword "exterior wall coating," and then determines the corresponding column based on the table association terms "total material" and "unit price of material." The processor then determines the cell content at the intersection of the corresponding data row and the corresponding column as the target tabular data. Alternatively, the processor can treat the entire corresponding data row as the target tabular data.

[0127] For example, the processor can determine the target table data based on the similarity between the data name of the table data and the keywords or table-related words. When the similarity is greater than a preset threshold, the table data is determined to be the target table data.

[0128] In some implementations of this specification, user input is parsed to determine keywords, and text-related terms and table-related terms are determined based on the keywords, thereby generating corresponding text query statements and table query statements. This allows the query statements to match the expressive characteristics of different types of data, avoiding matching bias caused by using a uniform retrieval strategy.

[0129] Figure 5 This is an exemplary schematic diagram illustrating the generation of the target output according to some embodiments of this specification.

[0130] In some embodiments, the processor determines text core parameters 520 related to keywords based on the original text 510 corresponding to the target text data; determines a text object 530 based on the text core parameters 520; determines table core parameters 550 related to keywords based on the original table text 540 corresponding to the target table data; determines a table object 560 based on the table core parameters 550; determines the consistency state 580 between the text object 530 and the table object 560 based on the knowledge graph 570; and generates the target output 590 based on the text object 530, the table object 560, and the consistency state 580.

[0131] For more information on keywords, target text data, target table data, consistency status, and target output, please refer to [link to relevant documentation]. Figure 3 and Figure 4 And related content.

[0132] The original text corresponding to the target text data refers to the original text content in the original file that corresponds to the target text data.

[0133] Core parameters of a text refer to information related to keywords in the original text. Examples include material type, functional attributes, performance indicators, and grading standards.

[0134] For example, when the original text is "The exterior wall surface layer adopts high-elastic waterproof coating with a thickness of not less than 2mm and a crack resistance level of Grade 1", the corresponding core parameters of the text for the keyword "exterior wall coating" are "material type", "functional attribute", "thickness" and "crack resistance level standard".

[0135] In some embodiments, the processor determines the node corresponding to the keyword in the knowledge graph, and among other nodes that are associated with the node, filters out nodes whose standard fields are the same or similar to the semantics of the original text, and determines the standard fields corresponding to the filtered nodes as the core parameters of the text.

[0136] A text object can be the original text corresponding to the target text data presented in a semi-structured data format.

[0137] The original table text corresponding to the target table data refers to the original table content in the original file that corresponds to the target table data.

[0138] The core parameters of a table refer to information related to keywords in the original table text. For example, material type, total material quantity, and unit price of material.

[0139] In some embodiments, the processor determines the node corresponding to the keyword in the knowledge graph, and among other nodes that are associated with the node, filters out nodes whose standard fields are semantically the same or similar to the header row of the original table text, and determines the standard fields corresponding to the filtered nodes as the core parameters of the table.

[0140] A table object can be the original table text corresponding to the target table data presented in a semi-structured data format.

[0141] In some embodiments, the processor can determine the text parameter content corresponding to the text core parameters from the original text based on the text core parameters; and determine the text object based on the text core parameters and the text parameter content. Similarly, the processor can determine the table parameter content corresponding to the table core parameters from the original table text based on keywords and table core parameters; and determine the table object based on the table core parameters and the table parameter content.

[0142] Text parameter content refers to the parameter content corresponding to the core text parameter in the original text. In some embodiments, the processor combines the core text parameter with the corresponding text parameter content to generate a text object. For example, when the original text is "The exterior wall surface layer uses high-elasticity waterproof coating with a thickness of not less than 2mm and a crack resistance level of Grade 1", the text object can be ("Material Type": "Waterproof Coating", "Functional Attribute": "High Elasticity", "Thickness": "Not Less Than 2mm", "Crack Resistance Level Standard": "Grade 1").

[0143] The table parameter content refers to the parameter content corresponding to the core parameters in the original table text.

[0144] In some embodiments, the processor can perform column-level filtering on the original table text, deleting data columns unrelated to the core table parameters and retaining only the data columns corresponding to the core table parameters. The processor can determine the data row corresponding to the keyword based on the keyword, and use the content of multiple cells in that data row as the table parameter content.

[0145] In some embodiments, the processor can combine the core table parameters with the corresponding table parameter content to generate a table object. For example, if the original table text after retaining the data columns corresponding to the core table parameters is "Material Type": "Stone Paint", "Total Material": "45 tons", "Unit Price of Material": "45 yuan / square meter", the processor can determine the data row corresponding to the keyword "Stone Paint" and use the content of multiple cells (total material and unit price of material) in that data row as the table parameter content. Then the table object can be ("Material Type": "Stone Paint", "Total Material": "45 tons", "Unit Price of Material": "45 yuan / square meter").

[0146] In some embodiments of this specification, text objects and table objects are further extracted from the original file to achieve structural uniformity of the target text data and target table data, providing a data foundation for subsequent consistency verification.

[0147] In some embodiments, the processor determines a mapping table based on a knowledge graph, text core parameters, and table core parameters; based on the mapping table, it determines the standard fields corresponding to the text core parameters and table core parameters respectively, and classifies the standard fields; based on the classified standard fields and the corresponding text parameter content and / or table parameter content, it determines one or more entity pairs; based on the knowledge graph, it determines the consistency state of one or more entity pairs, and based on the consistency state of one or more entity pairs, it determines the consistency state of the text object and the table object.

[0148] A mapping table is a table that contains the correspondence between text core parameters, table core parameters, and standard fields. Specifically, a single text core parameter corresponds to a single standard field, and a single table core parameter corresponds to a single standard field.

[0149] In some embodiments, for a single text core parameter, the processor can filter the nodes corresponding to the text core parameter based on a knowledge graph, and establish a correspondence between the standard fields corresponding to the filtered nodes and the text core parameter. For a single table core parameter, the processor can filter the nodes corresponding to the text core parameter based on a knowledge graph, and establish a correspondence between the standard fields corresponding to the filtered nodes and the table core parameter. The processor determines the mapping table using the above method.

[0150] In some embodiments, the processor determines the standard fields corresponding to the text core parameters and the table core parameters respectively according to the mapping table, and classifies the standard fields in the mapping table according to whether there are standard fields that simultaneously map the text core parameters and the table core parameters.

[0151] In some embodiments, the categorized standard fields include at least a first standard field and a second standard field.

[0152] The first standard field simultaneously maps to a set of text core parameters and table core parameters, and the first standard field simultaneously corresponds to a set of text parameter content and table parameter content.

[0153] In some embodiments, the second standard field maps to a text core parameter or a table core parameter, and the second standard field corresponds to a text parameter content or a table parameter content.

[0154] An entity pair is a semi-structured data object generated from a categorized standard field and its corresponding text and / or table parameter content. In some embodiments, an entity pair may be represented as (standard field, text parameter content, table parameter content), etc.

[0155] In some embodiments, for a single first standard field, the processor combines the first standard field, a set of corresponding text parameter contents, and table parameter contents to form a single entity pair. For a single second standard field, the processor combines the second standard field with its corresponding text parameter contents or table parameter contents to form a single entity pair, where any text parameter contents or table parameter contents for which the second standard field does not correspond are set to 0.

[0156] The consistency status of entity pairs is used to characterize the degree of semantic consistency between the text parameter content and the table parameter content in the entity pair.

[0157] In some embodiments, the consistency state of an entity pair includes consistent, conflicting, and missing.

[0158] In some embodiments, the processor performs a consistency comparison between the text parameter content and the table parameter content in the entity pair corresponding to the first standard field based on the knowledge graph; if the text parameter content and the table parameter content in the entity pair corresponding to the first standard field are consistent or have an inclusive relationship, the consistency state of the entity pair corresponding to the first standard field is determined to be consistent; if the text parameter content and the table parameter content in the entity pair corresponding to the first standard field are mutually exclusive, the consistency state of the entity pair corresponding to the first standard field is determined to be conflicted.

[0159] In some embodiments, for a single entity pair, the processor retrieves the nodes corresponding to the text parameter content and the table parameter content in the knowledge graph, calculates the path distance and similarity between the two nodes, and performs a consistency comparison between the text parameter content and the table parameter content based on the path distance and similarity.

[0160] Path distance refers to the length of the shortest path between two nodes. For example, if the shortest path from node A to node D is ABCD, then the path distance is 3. Similarity can be represented by the similarity of the standard fields corresponding to two nodes (such as the cosine similarity between feature vectors constructed based on the standard fields).

[0161] In some embodiments, if the path distance and similarity between the nodes corresponding to the text parameter content and the table parameter content respectively meet the consistency condition, the processor determines that the text parameter content and the table parameter content in the entity pair corresponding to the first standard field are consistent. The consistency condition may include a path distance less than a distance threshold and a similarity not less than a third similarity threshold. The distance threshold and the third similarity threshold can be set by those skilled in the art based on experience.

[0162] An inclusion relationship can occur when, in an entity pair corresponding to the first standard field, the text parameter content or table parameter content is contained within the content of another parameter. For example, if the text parameter content is "waterproof coating" and the table parameter content is "polymer waterproof coating", then the table parameter content is contained within the text parameter content.

[0163] In some embodiments, if the path distance and similarity between the nodes corresponding to the text parameter content and the table parameter content do not meet the consistency condition, the processor determines that the text parameter content and the table parameter content in the entity pair corresponding to the first standard field are mutually exclusive.

[0164] In some embodiments, by performing a consistency comparison between the text parameter content and the table parameter content in the entity pair corresponding to the first standard field based on the knowledge graph, misjudgments caused by naming or expression differences between the text parameter content and the table parameter content can be effectively eliminated, thereby improving data credibility.

[0165] In some embodiments, the processor determines the consistency status of the entity pair corresponding to the second standard field as missing.

[0166] In some embodiments, the processor determines the consistency state of an entity pair as the consistency state of the text object and the table object corresponding to the text core parameter and the table core parameter in the entity pair.

[0167] In some embodiments of this specification, by constructing a mapping table between core text parameters and core table parameters based on a knowledge graph, and determining the standard fields corresponding to the core text parameters and core table parameters respectively, a stable semantic alignment relationship can be established between text objects and table objects. Based on this, by classifying the standard fields and constructing entity pairs, the consistency status of the text parameter content and the table parameter content is analyzed accordingly. This provides a clear basis for judging the consistency between text content and table content, helps to accurately identify parameter conflicts or missing information in the text and table, and improves the credibility of the consistency verification results.

[0168] In some embodiments, the processor can also determine prompt words based on text objects, table objects, and target-related content in the knowledge graph; input the prompt words into the second model to obtain the consistency state of text objects and table objects.

[0169] The target-related content can be content in the knowledge graph that is related to keywords, core text parameters, core table parameters, text parameter content, and table parameter content.

[0170] In some embodiments, the processor can construct prompt words based on text objects, table objects, and target-related content in the knowledge graph. For example, the processor can construct prompt words based on text objects, table objects, and target-related content in the knowledge graph, using prompt word templates or similar methods.

[0171] For example, the prompt could be: "You are a construction cost auditing expert. The text object is ("Material Type": "Waterproof Coating", "Crack Resistance Standard": "Level 1"), and the table object is ("Material Type": "Stone Paint", "Total Material": "45 tons"). The target-related content is that waterproof coating may have the functional attribute of "high elasticity", and stone paint may be a "three-coat, two-brush" construction process. Please determine the consistency between the text object and the table object and provide the reason for your judgment." In some embodiments, the second model is similar to the first model; for a description of the first model, please refer to [link to relevant documentation]. Figure 3 And its related descriptions.

[0172] In some embodiments, the second model may also be a machine learning model. For example, the second model may include any one or a combination of neural network (NN) models, graph neural networks (GNN) or other custom model structures.

[0173] In some embodiments, the processor can train a second model based on a large number of labeled training samples using methods such as gradient descent. The training samples may include sample text objects, sample table objects, and sample knowledge graphs. The first label can be the actual consistency state of the sample text objects and sample table objects. The training samples can be obtained based on historical data, and the labels can be determined based on manual annotation.

[0174] In some embodiments, the second model can be trained as follows: multiple labeled training samples are input into an initial second model; a loss function is constructed using the labels and the prediction results of the initial second model; the initial second model is iteratively updated based on the loss function; and the training of the second model is complete when the loss function of the initial second model satisfies a preset condition. The preset condition may be, for example, the loss function converging or the number of iterations reaching a set value.

[0175] In some embodiments of this specification, prompt words are constructed using text objects, table objects, and target-related content in a knowledge graph. These prompt words are then input into a large language model to determine the consistency status between the text objects and the table objects. This introduces domain knowledge as semantic constraints into the consistency judgment process, thereby avoiding reliance solely on surface text matching for comparison and improving the accuracy of consistency status determination.

[0176] In some embodiments, when the consistency state of the text object and the table object is consistent, the table object is used as the first output (i.e., the primary output) in the target output, and the text object is used as the second output (i.e., the supplementary output) in the target output.

[0177] In some embodiments, when the consistency status of a text object and a table object is in conflict, the text object or the table object is determined as the target output based on the user's selection.

[0178] In some embodiments, the processor records the user's selection; when the consistency state of a future text object and a future table object conflicts, the processor determines the future text object or the future table object as the target output based on the user's selection.

[0179] A future text object refers to a semi-structured data object generated from the core parameters of the text determined during future processing and their corresponding parameter content in the original text.

[0180] A future table object refers to a semi-structured data object generated from the core parameters of the table determined during future processing and their corresponding parameter content in the original table text.

[0181] In some embodiments of this specification, generating different target outputs based on consistency status and user selection can reduce the risk of erroneous target outputs. By reusing historical user selection results in subsequent conflict scenarios, user operations can be reduced, target output generation efficiency can be improved, and the system can gradually form an output decision-making pattern that conforms to specific business scenarios or user preferences.

[0182] In some embodiments of this specification, by constructing text objects and table objects in a unified format and introducing knowledge graphs to perform semantic-level association and comparison of text objects and table objects, the accuracy and reliability of consistency judgment between unstructured text and structured tables are improved, and unified understanding and efficient output of information across data formats are realized.

[0183] In some embodiments, the target text data has a first weight, and the target table data has a second weight. The second weight is greater than the first weight. The values ​​of the first and second weights can be preset manually, such as the first weight being 0.4 and the second weight being 0.6.

[0184] Figure 6 This is an exemplary flowchart illustrating the generation of target confidence scores for target outputs according to some embodiments of this specification. For example... Figure 6 As shown, process 600 includes the following steps 610-630.

[0185] Step 610: Determine the first similarity between the target text data and the user input.

[0186] For more information on target text data and user input, see [link to relevant documentation]. Figure 3 And its related descriptions.

[0187] The first similarity refers to the numerical value that characterizes the semantic similarity between the target text data and the user input.

[0188] In some embodiments, the first similarity can be the cosine similarity between the feature vector corresponding to the target text data and the feature vector corresponding to the user input.

[0189] Step 620: Determine the second similarity between the target table data and the user input.

[0190] For more information about the target table data, please see [link / details]. Figure 3 And its related descriptions.

[0191] The second similarity refers to a numerical value that characterizes the degree of similarity between the target table data and the user input in terms of table semantics or field matching.

[0192] In some embodiments, the second similarity can be the cosine similarity between the feature vector corresponding to the target table data and the feature vector corresponding to the user input.

[0193] Step 630: Based on the first similarity, the second similarity, the first weight, and the second weight, generate the target confidence score of the target output.

[0194] Target confidence refers to a numerical value that characterizes the credibility of the target output.

[0195] In some embodiments, when the consistency status of the target text data and the target table data is consistent, the processor can use a first weight as the weight of the first similarity and a second weight as the weight of the second similarity, and perform a weighted calculation on the first similarity and the second similarity to generate the target confidence score. When the consistency status is missing (e.g., missing target text data), the processor multiplies the similarity of the data that is not missing with the weight (e.g., the second similarity and the second weight) to generate the target confidence score. When the consistency status is conflicting, the processor calculates the product of the first similarity and the first weight, and the product of the second similarity and the second weight, respectively, and sets the lower of the two products as the target confidence score.

[0196] In some implementations of this specification, by setting different weights for the target text data and the target table data respectively, and combining the similarity with the user input to generate the target confidence score of the target output, the reliability of the final output can be identified, making it easier for users to judge the search results.

[0197] In some embodiments, the processor adds a risk label to the target output based on the consistency status of the target text data and the target table data, as well as the target confidence level.

[0198] Risk identification refers to identification information that indicates that the target output has a risk.

[0199] In some embodiments, the processor can classify the risk level of the target output based on the consistency status of the target text data and the target table data, as well as the target confidence level, and add a risk label accordingly. When the consistency status is consistent and the target confidence level is not lower than a first threshold, the processor determines the risk level of the target output to be low risk and adds a "low risk" risk label. When the consistency status is missing or the target confidence level is lower than the first threshold but not lower than the second threshold, the processor determines the risk level of the target output to be medium risk, and can highlight the target output (e.g., use a yellow color or bold it) and add a "medium risk" risk label, and add a risk information field to the target output (e.g., "Relevant description found only in technical documents, not explicitly listed in business documents, please manually verify"). When the consistency status is conflicting or the target confidence level is lower than the second threshold, the processor determines the risk level of the target output to be high risk. It can then highlight the target output (e.g., using red or other colors, or bolding and flashing it) and add a "High Risk" risk label. A risk information field can also be added to the target output (e.g., "Section 3.2 of the technical document describes it as 'Stone Paint,' while item 15 of the quotation details in the business document records 'Elastic Latex Paint.' Please verify manually."). The first and second thresholds can be set by those skilled in the art based on experience, with the first threshold being greater than the second threshold.

[0200] In some embodiments, the processor can record the location information (such as page number and line number) of the target text data and target table data with a consistency conflict in the original file, so that users can quickly locate the original text for manual review.

[0201] In some embodiments of this specification, by dynamically adding corresponding risk identifiers and risk information fields to the target output, the reliability and potential risk level of the target output can be intuitively reflected, making it easier for users to quickly identify high-risk information and conduct manual verification, thereby effectively reducing decision-making and compliance risks caused by information conflicts or omissions.

[0202] It should be noted that the above descriptions of processes 300, 400, and 600 are for illustrative purposes only and do not limit the scope of this specification. Those skilled in the art can make various modifications and changes to processes 300, 400, and 600 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0203] Some embodiments of this specification also provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the method described in any of the above embodiments.

[0204] Furthermore, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.

[0205] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are sometimes modified by the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. Although the numerical ranges and parameters used to confirm their breadth in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0206] If there is any inconsistency or conflict between the descriptions, definitions, and / or terms used in the materials referenced in this specification and the content described in this specification, the descriptions, definitions, and / or terms used in this specification shall prevail.

Claims

1. A method for extracting information from architectural documents, characterized in that, The method includes: Obtain the original files of the architectural documents, which include text content and table content; Scan the original file and divide it into text areas and table areas; Extract the text data from the text region and store the text data in a text database; Extract the table data from the table area and store the table data in the table database; Based on user input, retrieve target text data and target table data corresponding to the user input from the text database and the table database; Determine the consistency status of the target text data and the target table data, and generate the target output.

2. The method as described in claim 1, characterized in that, The step of scanning the original file and dividing it into text and table regions includes: The original file is scanned using a first model to obtain a first feature corresponding to the text content and a second feature corresponding to the table content. The text region is determined based on the first feature; The table region is determined based on the second feature.

3. The method as described in claim 2, characterized in that, The step of extracting the table data from the table area and storing the table data in the table database includes: Based on one or more header rows in the table area, determine one or more data names; Based on the data rows corresponding to the one or more header rows, determine the data values ​​corresponding to the one or more data names; Based on the one or more data names and their corresponding data values, determine the structured or semi-structured tabular data; The structured or semi-structured tabular data is stored in the tabular database.

4. The method as described in claim 2, characterized in that, The table area includes a first table and a second table spanning multiple pages. Extracting the table data from the table area and storing the table data in the table database further includes: Identify the header row of the first table and the header row of the second table; Determine the similarity between the header row of the first table and the header row of the second table; When the similarity meets the similarity condition, the first table and the second table are determined to be the same table; When the similarity does not meet the similarity condition, the first table and the second table are determined to be different tables.

5. The method as described in claim 1, characterized in that, The step of retrieving target text data and target table data corresponding to the user input from the text database and the table database based on the user input includes: Keywords are determined based on the user input; Based on the keywords, determine text-related terms and table-related terms; Based on the text-related keywords, a text query statement is determined, and based on the text query statement, the target text data is retrieved from the text database; Based on the table related terms, a table query statement is determined, and based on the table query statement, the target table data is retrieved from the table database.

6. The method as described in claim 5, characterized in that, The step of retrieving the target text data from the text database based on the text query statement includes: Based on the text query statement, determine the target feature vector; Based on multiple text feature vectors in the text database, the similarity between the multiple text feature vectors and the target feature vector is determined; Based on the similarity, a preset number of text data are selected as the target text data.

7. The method as described in claim 5, characterized in that, The process of determining the consistency status of the target text data and the target table data, and generating the target output, includes: Based on the original text corresponding to the target text data, determine the core text parameters related to the keywords; Based on the aforementioned core text parameters, the text object is determined; Based on the original text of the target table data, determine the core table parameters related to the keywords; Based on the core parameters of the table, determine the table object; Based on the knowledge graph, the consistency status between the text object and the table object is determined; The target output is generated based on the text object, the table object, and the consistency status between the text object and the table object.

8. The method as described in claim 7, characterized in that, The process of determining the text object based on the core text parameters includes: Based on the core text parameters, determine the text parameter content corresponding to the core text parameters from the original text. The text object is determined based on the core text parameters and the content of the text parameters; The process of determining the table object based on the core parameters of the table includes: Based on the keywords and the core parameters of the table, determine the table parameter content corresponding to the core parameters of the table from the original text of the table. The table object is determined based on the core parameters and the content of the table parameters.

9. The method as described in claim 7, characterized in that, The process of determining the consistency status between the text object and the table object based on the knowledge graph includes: Based on the knowledge graph, the core parameters of the text, and the core parameters of the table, a mapping table is determined; Based on the mapping table, the standard fields corresponding to the core parameters of the text and the core parameters of the table are determined respectively, and the standard fields are classified. Based on the categorized standard fields and the corresponding text parameter content and / or table parameter content, determine one or more entity pairs; Based on the knowledge graph, the consistency state of the one or more entity pairs is determined, and based on the consistency state of the one or more entity pairs, the consistency state of the text object and the table object is determined.

10. The method as described in claim 9, characterized in that, The categorized standard fields include a first standard field, which simultaneously maps to a set of text core parameters and table core parameters, and also corresponds to a set of text parameter content and table parameter content. Determining the consistency state of the one or more entity pairs based on the knowledge graph includes: Based on the knowledge graph, a consistency comparison is performed on the text parameter content and table parameter content of the entity pairs corresponding to the first standard field. If the text parameter content and table parameter content in the entity pair corresponding to the first standard field are consistent or have an inclusive relationship, the consistency status of the entity pair corresponding to the first standard field is determined to be consistent. If the text parameter content and table parameter content in the entity pair corresponding to the first standard field are mutually exclusive, the consistency state of the entity pair corresponding to the first standard field is determined to be conflicted.

11. The method as described in claim 9, characterized in that, The categorized standard fields include a second standard field, which maps to a text core parameter or a table core parameter. The second standard field corresponds to text parameter content or table parameter content. Determining the consistency state of the one or more entity pairs based on the knowledge graph includes: The consistency status of the entity pair corresponding to the second standard field is determined to be missing.

12. The method as described in claim 8, characterized in that, The process of determining the consistency status between the text object and the table object based on the knowledge graph includes: Based on the text object, the table object, and the target-related content in the knowledge graph, prompt words are determined; The prompt word is input into the second model to obtain the consistency status between the text object and the table object.

13. The method as described in claim 10 or 12, characterized in that, The step of generating the target output based on the text object, the table object, and the consistency status between the text object and the table object includes: When the consistency status of the text object and the table object is consistent, the table object is used as the first output in the target output, and the text object is used as the second output in the target output.

14. The method as described in claim 10 or 12, characterized in that, The step of generating the target output based on the text object, the table object, and the consistency status between the text object and the table object includes: When the consistency status of the text object and the table object is in conflict, the text object or the table object is determined as the target output based on the user's selection.

15. The method as described in claim 14, characterized in that, The method further includes: Record the user's selection; When there is a conflict between the consistency status of the future text object and the future table object, the future text object or the future table object is determined as the target output based on the user selection.

16. The method as described in claim 1, characterized in that, The target text data has a first weight, and the target table data has a second weight, where the second weight is greater than the first weight. Determining the consistency status of the target text data and the target table data, and generating the target output, includes: Determine the first similarity between the target text data and the user input; Determine the second similarity between the target table data and the user input; Based on the first similarity, the second similarity, the first weight, and the second weight, the target confidence score of the target output is generated.

17. The method as described in claim 16, characterized in that, The method further includes: Based on the consistency status of the target text data and the target table data, as well as the target confidence level, a risk label is added to the target output.

18. The method as described in claim 1, characterized in that, The step of scanning the original file and dividing it into text and table regions includes: Scan the original file to determine the coordinate information of the text area and the coordinate information of the table area; The step of extracting the text data from the text region and storing the text data in a text database includes: The text data and the coordinate information of the text region are stored in the text database; The step of extracting the table data from the table area and storing the table data in the table database includes: Store the table data and the coordinate information of the table area into the table database; The process of determining the consistency status of the target text data and the target table data, and generating the target output, includes: The target output is bound to the coordinate information of the text area and the coordinate information of the table area.

19. An information extraction system for architectural documents, characterized in that, The system includes: The acquisition module is configured to acquire the original files of the building documents, which include text content and table content; The processing module is configured as follows: Scan the original file and divide it into text areas and table areas; Extract the text data from the text region and store the text data in a text database; Extract the table data from the table area and store the table data in the table database; Based on user input, retrieve target text data and target table data corresponding to the user input from the text database and the table database; Determine the consistency status of the target text data and the target table data, and generate the target output.

20. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions from the storage medium, the computer executes the method as described in any one of claims 1 to 18.

Citation Information

Patent Citations

  • Intelligent extraction system and method for financial document information

    CN110889310A

  • Log inspection intelligent question-answering system based on multi-modal domain knowledge base and construction method of log inspection intelligent question-answering system

    CN121189466A

  • Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

    CN121278157A

  • Geological reservoir feature automatic identification and integration system, method and equipment based on multi-modal data and medium

    CN121365350A

  • Intelligent document compliance auditing system and method based on multi-modal deep learning

    CN121365360A