Langchain-based building specification graphical representation multi-modal question and answer method and equipment
By building a multimodal database and deep learning model using the Langchain framework, the problem of multimodal information processing of building code illustrations and text was solved, achieving efficient and accurate question answering of building code illustrations, improving query efficiency and reducing labor costs.
Patent Information
- Application Number
- CN202510882341.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies struggle to effectively process the multimodal information in building code illustrations and text, resulting in one-sided query results that fail to meet the actual needs of the construction industry. Furthermore, multimodal models perform poorly in feature extraction and semantic parsing when processing vector graphics in the construction industry.
A multimodal database of building codes is constructed using the Langchain framework. Through a deep learning-based text vectorization model and a distributed storage strategy, consistency verification of text and images is achieved. Combined with multi-dimensional retrieval and reordering, structured answers are generated.
It enables efficient and accurate Q&A on building code diagrams, improves query efficiency, reduces labor costs, and provides support for intelligent applications in the construction industry.
Smart Images

Figure CN120849547A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary field of artificial intelligence and the construction industry, and in particular to a multimodal question-answering method and device for building code illustrations based on Langchain. Background Technology
[0002] In the wave of digital transformation in the construction industry, the demand for accurate and efficient access to specification and graphic information in architectural design and construction is becoming increasingly urgent. Drawings, as the core information carrier, are deeply coupled with specification provisions. However, traditional single-modal processing technology has fundamental flaws, failing to achieve an effective correlation between drawings and specification provisions, resulting in one-sided query results that are difficult to match the needs of actual engineering scenarios. At the same time, the construction industry's specification system is vast and complex, covering multiple professional fields such as architecture, structure, water supply, heating, and electricity. The limitations of professional personnel's technical accumulation make specification and graphic query efficiency low, and the promotion of new specifications faces numerous obstacles, severely restricting the industry's intelligent development process.
[0003] With the rapid development of large-scale model technology, knowledge base question-answering systems built on retrieval-enhanced generation (RAG) have become a research hotspot. Applying this technology to the field of building code illustration question-answering is expected to significantly reduce the burden of memorizing regulations for professionals and avoid safety risks caused by delayed updates to regulations. However, building industry regulations primarily use vector graphics, which, compared to conventional bitmaps (such as images of people or natural landscapes), have lower information density and simpler visual features. This makes it difficult for multimodal models to capture image-text matching relationships and effectively distinguish similar illustrations, greatly increasing the technical difficulty of building code illustration question-answering.
[0004] Existing multimodal models, such as CILP, can train alignment vectors to describe graphic content using text-image pairs. However, they have stringent requirements for image quality. When processing vector graphics in the construction industry, they are limited by the special geometric structure and professional semantic expression, making effective feature extraction and semantic parsing difficult. Their performance in matching text and graphics to building codes is unsatisfactory. Even though large open-source models like Deepseek and Qwen have significantly improved natural language processing capabilities in recent years by expanding training corpora and optimizing architecture algorithms, effectively modeling human natural language expression and thinking abilities, and achieving efficient integration with knowledge bases through RAG technology, a targeted solution is still lacking in the vertical field of building code graphic question answering.
[0005] While the large-scale model development framework Langchain possesses powerful input / output control capabilities, enabling precise regulation of large-scale model output through prompt word templates, response format definitions, and structured output parsers, and constructing development pipelines through a "chain" structure, its application in multimodal question-and-answer scenarios for building code diagrams remains unexplored. Overcoming existing technological bottlenecks and deeply integrating the Langchain framework with the needs of the construction industry has become a critical technical issue that urgently needs to be addressed. Summary of the Invention
[0006] This application provides a Langchain-based multimodal question-answering method and device for architectural specification illustrations, which addresses the difficulties in processing multimodal information and low query efficiency in architectural specification illustration question-answering.
[0007] The technical solution adopted in this application is as follows:
[0008] On the one hand, this application provides a multimodal question-answering method for building code diagrams based on Langchain, the method comprising:
[0009] Construct a multimodal database of building codes and store and manage the multimodal data of building codes;
[0010] Based on user query commands, the multimodal database is retrieved from different dimensions;
[0011] Based on the search results, generate a structured answer corresponding to the user's query.
[0012] The structured answer is validated for consistency between text and image. If the validation passes, the corresponding answer to the user's query command is output.
[0013] In one possible implementation of this application, the building code multimodal data includes building code text and building code illustrations;
[0014] Storage and management of multimodal data related to building codes, including:
[0015] The data storage fields include NAME, CONTENT, VECTOR, URL, BASE64, VAILD, and UPDATE fields;
[0016] The NAME field stores the name and number of the building code illustration; the CONTENT field stores the building code text in UTF-8 encoding; the VECTOR field stores the semantic vector of the building code text with a preset dimension generated by a pre-trained model; the URL field stores the address of the building code illustration corresponding to the building code text; the BASE64 field stores the Base64 encoded string of the building code illustration; the VALID field is a boolean type, where true indicates that the building code multimodal data is valid, and false indicates that the building code multimodal data is obsolete; and the UPDATE field records the update time of the building code multimodal data.
[0017] In one possible implementation of this application, the storage and management of multimodal data of building codes includes:
[0018] For building code texts, a deep learning-based text vectorization model is used to extract and transform deep semantic features, including:
[0019] The building code text is segmented into sentences or paragraphs and then input into the BERT model so that the BERT model can generate a corresponding hidden state vector for each input token;
[0020] Through average pooling, the hidden state vectors corresponding to all tokens are merged into a fixed-length text feature vector v. text ;
[0021] Obtain the text feature vector v text Then, the building code text is stored in the CONTENT field, and the text feature vector v is also stored in the CONTENT field. text Store in the VECTOR field.
[0022] In one possible implementation of this application, the storage and management of multimodal data of building codes includes:
[0023] For building code drawings, a storage strategy combining distributed storage and Base64 encoding conversion is adopted, including:
[0024] A unique storage path is generated for each building code illustration, and the storage path is stored in the URL field in URL format to form the network access identifier of the building code illustration;
[0025] Read the binary data corresponding to the building code drawing, convert the binary data into a string according to the Base64 encoding rules, and store the string in the BASE64 field.
[0026] In one possible implementation of this application, constructing a multimodal database of building codes and storing and managing the multimodal data of building codes further includes:
[0027] Obtain the latest building code text;
[0028] By using a text comparison algorithm, the latest building code text is compared with the current building code text line by line, and the newly added, modified and deleted content is automatically identified to obtain updated data;
[0029] The process of updating the multimodal database of building codes is triggered by first updating the building code text to a vectorized form, and then regenerating the text feature vector v. text Then fill in the VECTOR field, update the text content of the CONTENT field, replace the building code illustration with the correct version, and update the corresponding URL address.
[0030] After the update is complete, a notification of the data change is sent via a message queue.
[0031] In one possible implementation of this application, the multimodal database is retrieved from different dimensions based on user query instructions, including:
[0032] The user's query command is input into the BERT model to generate the query feature vector v corresponding to the user's query command. query ;
[0033] Calculate the query feature vector v query Compared with all text feature vectors v in the building code multimodal database text Cosine similarity between them;
[0034] Architectural code texts with similarity values exceeding a preset threshold are selected as candidate set C1;
[0035] The BM25 algorithm is used to match the keywords of the user's query with the building code text, and the top N building code text fragments with the highest BM25 scores are selected to form a candidate set C2;
[0036] Using the hybrid similarity calculation formula, candidate set C3 is generated by combining cosine similarity and BM25 score;
[0037] Merge candidate sets C1, C2, and C3, remove duplicates, and obtain the multi-path recall candidate set C. text ;
[0038] The multi-path recall candidate set C is reordered using a re-ranking model. text Sort the results and select the top K candidate results by score as the search results.
[0039] In one possible implementation of this application, a structured answer corresponding to the user's query instruction is generated based on the search results, including:
[0040] The Prompt template is obtained by combining the user's query command, search results and output format requirements;
[0041] Using the Langchain framework, the return content of the Qwen large model is defined as JSON format, and two key values are set: Content and URL List;
[0042] The value corresponding to Content is the text answer generated by the Qwen big model based on the Prompt template, used to answer the user's query command about building codes. The value corresponding to URL List is a list containing image URLs related to the text answer.
[0043] In one possible implementation of this application, after obtaining the returned content in JSON format, the method further includes:
[0044] Sensitive word filtering and compliance verification are performed on the text answers in the Content field.
[0045] In one possible implementation of this application, the structured answer is subjected to a text-image consistency check, including:
[0046] Extract the list of image URLs contained in the URL List field;
[0047] Establish a mapping database between building code texts and building code illustrations;
[0048] For the building code clause number referenced in the structured answer, retrieve the set of building code illustrations associated with the building code clause in the multimodal building code database;
[0049] The set of building code illustrations is compared with the set of illustrations corresponding to the list of image URLs. If there are differences between the two, it is determined that the text and images are inconsistent.
[0050] On the other hand, this application also provides a Langchain-based multimodal question-answering device for architectural specification diagrams, the device comprising:
[0051] At least one processor; and,
[0052] A memory communicatively connected to the at least one processor; wherein,
[0053] The memory stores instructions that can be executed by the at least one processor, enabling the at least one processor to perform the following:
[0054] Construct a multimodal database of building codes and store and manage the multimodal data of building codes;
[0055] Based on user query commands, the multimodal database is retrieved from different dimensions;
[0056] Based on the search results, generate a structured answer corresponding to the user's query.
[0057] The structured answer is validated for consistency between text and image. If the validation passes, the corresponding answer to the user's query command is output.
[0058] This application provides a Langchain-based multimodal question-answering method and device for architectural specification diagrams, which has the following beneficial effects:
[0059] This application, through the aforementioned methods, achieves a closed-loop process encompassing multimodal data storage and updating, information retrieval and sorting, structured answer generation, and text-image consistency verification. It constructs an efficient and accurate multimodal question-and-answer system based on a large text-based question-and-answer model, effectively improving the accuracy and query efficiency of building code illustration questions and answers, significantly reducing the manual costs of code consultation, and providing strong support for the intelligent application of building industry code information. In other words, this application achieves efficient and accurate multimodal question-and-answer for building codes by constructing a complete question-and-answer system methodology. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0061] Figure 1 This application provides a flowchart of a multimodal question-answering method for building code illustration based on Langchain;
[0062] Figure 2 This application provides a schematic diagram illustrating the working principle of a multimodal question-answering method for building code illustrations based on Langchain.
[0063] Figure 3 This application provides a structural schematic diagram of a multimodal question-and-answer device based on Langchain for architectural specification illustrations. Detailed Implementation
[0064] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0065] This application focuses on technological innovation in multimodal question-answering for building code illustrations. Specifically, it proposes a large-model structured generation technology based on the Langchain framework, integrating key technologies such as multimodal data storage and updating, text vectorization processing, and multi-path recall retrieval and rearrangement. By constructing a multimodal question-answering system for building codes, it achieves efficient integration of multimodal data such as text and images, accurately matching building code provisions with illustrated information. This technology is applicable to scenarios such as intelligent review of architectural design drawings, real-time consultation on construction codes, and intelligent supervision of building engineering quality, providing an efficient and accurate intelligent solution for the retrieval and application of building code information, and promoting the digital and intelligent upgrading of the building industry.
[0066] The method in this application will be described in detail below with reference to the accompanying drawings.
[0067] Figure 1 A flowchart of a multimodal question-answering method for building code illustrations based on Langchain is provided for this application, such as... Figure 1 As shown, the multimodal question answering method in this application includes at least the following execution steps:
[0068] Step 101: Construct a multimodal database of building codes and store and manage the multimodal data of building codes.
[0069] This step employs an advanced data storage architecture to achieve efficient data storage and rapid retrieval. Simultaneously, a dynamic update mechanism is implemented. When building industry standards change, the updated content is automatically identified, and the standard clauses and corresponding illustrations in the database are updated promptly. This ensures that the Q&A system always provides services based on the latest industry standards, effectively addressing the impact of industry standard updates on Q&A effectiveness.
[0070] Specifically, this application chooses unstructured databases as the multimodal data storage medium, primarily based on the complexity and diversity of building code data. Building code texts contain numerous freely formatted clauses, while illustrations exist in the form of unstructured graphic files. Unstructured databases, such as MongoDB, possess flexible data models, enabling efficient storage and processing of this heterogeneous data. They do not require predefined fixed table structures and can dynamically expand as data grows and changes, greatly improving the adaptability of data storage. Simultaneously, their distributed architecture supports horizontal scaling, easily meeting the storage needs of massive amounts of building code data and ensuring stable performance even under high-concurrency access scenarios.
[0071] In terms of data storage format, a key-value structure is used. Each building code data unit corresponds to a unique key, which allows for quick location and access to its associated multimodal building code data values. This storage method significantly improves data retrieval efficiency and is particularly suitable for application scenarios where question-and-answer systems can quickly respond to query requests.
[0072] In one possible implementation of this application, the data storage field includes seven core fields: NAME, CONTENT, VECTOR, URL, BASE64, VAILD, and UPDATE. These fields work together to fully record the multimodal information of the building code.
[0073] The NAME field stores the name and number of the illustration, coded according to industry standard naming rules, such as "14X505-1-Design Specification for Automatic Fire Alarm Systems," where "14X505-1" represents the specification number and "Design Specification for Automatic Fire Alarm Systems" is the specification name. This naming method not only facilitates manual identification and management but also allows for quick association of specification text with corresponding illustrations within the system through numbering, achieving accurate matching of multimodal data.
[0074] The CONTENT field stores the text content of the specification. It uses UTF-8 encoding to fully preserve special characters, punctuation marks, and technical terms in the specification, ensuring efficient access to the specification text.
[0075] The VECTOR field stores semantic vectors generated from the canonical text using pre-trained models such as BERT and BGE-M3. Each vector has a fixed dimension of 768, using the BERT-base model as an example. Simultaneously, a vector index, such as the FAISS index, is established to support efficient vector similarity retrieval, enabling the system to quickly match canonical text semantically similar to the user's query.
[0076] The URL field stores the address of the illustration corresponding to the specified text, using the Uniform Resource Locator (URL) standard format. For illustrations stored in the file system, the URL includes information such as the network protocol for file storage (e.g., HTTP / HTTPS), server address, file path, and filename. The system can directly access and download the illustration file through this URL, enabling rapid association and display of text and image information.
[0077] The BASE64 field stores the Base64 encoded string of the image file. Converting architectural code diagrams to Base64 format facilitates the direct embedding of text information during processing, avoiding performance degradation caused by frequent file read / write operations. Furthermore, Base64 encoded diagram data can be directly rendered and displayed on the front-end interface, improving user convenience in viewing the diagrams.
[0078] The VALID field is a boolean type. `true` indicates that the specification fragment is valid, while `false` indicates that it is obsolete. When performing question-and-answer searches, the system prioritizes specification data with a `VALID` field of `true` to ensure that the information provided to users is currently valid, avoiding engineering risks caused by using obsolete specifications.
[0079] The UPDATE field records the update time of the specification, accurate to the second. This field enables version management of specification data. When a specification is updated, historical version data is automatically retained, allowing users to easily trace the history of specification changes. It also provides a timestamp for the dynamic update mechanism, ensuring that the data in the database is always synchronized with the latest industry standards.
[0080] In one possible implementation of this application, the text and illustrations in the construction industry standard system have a close and complex correspondence, significantly different from conventional multimodal question-and-answer scenarios. Each standard provision often requires multiple illustrations to explain it from different perspectives, i.e., a one-to-many relationship, while the same illustration may correspond to multiple different standard provisions, i.e., a many-to-many relationship. This complex mapping relationship places extremely high demands on data storage and retrieval. If not handled properly, it can easily lead to problems such as incorrect text-image matching and incomplete information retrieval, seriously affecting the accuracy and reliability of the question-and-answer system. Therefore, building an efficient text-image association mechanism from the data storage stage becomes a key breakthrough for achieving accurate multimodal question answering.
[0081] This application employs a deep learning-based text vectorization model to extract and transform deep semantic features from building code texts. Taking the BERT model as an example, its pre-training process is based on a large-scale general text corpus, and its bidirectional Transformer architecture can fully capture textual context information and accurately understand semantic connotations. When processing building code texts, the text is segmented into sentences or paragraphs and then input into the BERT model. The model generates a corresponding hidden state vector for each input token. Through average pooling, the hidden state vectors of all tokens are fused into a fixed-length text feature vector v. text Its generation formula is:
[0082]
[0083] Where n is the number of text tokens, BERT(text) i This is the hidden state vector of the i-th token output by the BERT model.
[0084] This semantic vector not only contains surface-level lexical information of the text but also uncovers the deep semantic logic of the regulatory provisions, greatly improving the accuracy and comprehensiveness of the text's semantic representation. After obtaining the text semantic vector, the original regulatory text is stored in the CONTENT field, preserving the complete content of the provisions; simultaneously, the semantic vector is stored in the VECTOR field. To achieve efficient semantic retrieval, the Facebook AI Similarity Search (FAISS) library is used to construct a vector index. FAISS provides various efficient index structures, such as Hierarchical NSW (HNSW) clustering and Product Quantization (PQ). Considering the characteristics of building code semantic vector data, the HNSW index structure is selected, as it exhibits excellent retrieval performance and memory efficiency in high-dimensional vector retrieval scenarios. Through the FAISS index, the system can quickly calculate the cosine similarity or Euclidean distance between the user's query vector and the stored vector, achieving millisecond-level semantic matching and significantly reducing the error rate caused by semantic misunderstanding when answering user questions.
[0085] For building code illustrations, this application adopts a storage strategy combining distributed storage and Base64 encoding conversion. Considering the large data volume and highly specialized nature of the illustration files, a unique storage path is generated for each illustration during storage, and this path is stored in the URL field in URL format to form the network access identifier for the illustration file.
[0086] To further improve the efficiency of data transmission and processing within the system, the read image data is converted to Base64 format. Base64 encoding is a method of converting binary data into printable ASCII characters, which can directly embed text information for transmission and storage. Specifically, the binary data of the image file is first read, then converted into a string according to Base64 encoding rules, and the string is stored in the BASE64 field. This storage method effectively avoids frequent file read / write operations, reduces I / O overhead during data transmission, and facilitates the direct display of Base64 encoded images on the front-end interface using HTML's `img` tag or relevant rendering libraries, enabling rapid presentation of text and image information.
[0087] Through the aforementioned text and image storage strategies, a precise correspondence between text and images is established at the data level. The system can quickly associate the image number in the NAME field with the standard text of the CONTENT field, the semantic vector of the VECTOR field, the image address in the URL field, and the image data in the BASE64 field. Whether dealing with complex one-to-many or many-to-many text-image relationships, it can achieve efficient and accurate text-image matching and retrieval, providing a solid data foundation for subsequent modules such as multi-path retrieval, answer generation, and text-image consistency verification. This ensures that the question-and-answer system can accurately respond to the multimodal query needs of users in the construction industry.
[0088] In addition, this application includes a dynamic update mechanism when storing multimodal data of building codes. Specifically, the latest code documents are obtained periodically by crawling official websites and standard release platforms of the building industry through a code update monitoring system. Using text comparison algorithms, such as edit distance algorithms, the new code text is compared clause by clause with existing codes in the database to automatically identify added, modified, and deleted content.
[0089] After confirming the update, the system automatically triggers the database update process: First, the affected text data is vectorized and updated, semantic vectors are regenerated and filled into the VECTOR field, and the text content of the CONTENT field is updated simultaneously. Next, the relevant illustrations are version-replaced, and the corresponding image URLs are updated. For specifications involving illustration changes, the VAILD and UPDATE fields in the database are used to implement document obsolescence and updates. After the update is complete, other modules are notified via a message queue to synchronize the data changes, ensuring the consistency and timeliness of data throughout the question-and-answer system.
[0090] Step 102: Based on the user's query command, retrieve the multimodal database from different dimensions.
[0091] This step primarily involves accurately locating and filtering information relevant to the user's query from the multimodal database of building codes. This step employs a "broad net first, then fine-tuning" strategy, effectively balancing the comprehensiveness and accuracy of the search through multi-dimensional retrieval and refined sorting, providing a high-quality data foundation for subsequent answer generation. Specifically, it includes the following two stages:
[0092] 1) Multi-channel recall phase
[0093] This phase employs a parallelized multi-strategy retrieval mechanism to search the multimodal database from different dimensions, maximizing the coverage of potentially relevant information. The specific implementation includes the following key steps:
[0094] Query vector processing: Input the user's query command into the BERT model, which uses the same model as the text storage stage, and generate a semantic vector v of the query text. query That is, querying the feature vector v query To enhance query comprehension, a query expansion technique is employed. Based on a pre-trained domain knowledge graph, it automatically identifies and supplements synonyms, hyponyms, and other extended terms related to the query keywords, thereby expanding the semantic scope of the query.
[0095] Multi-pronged recall strategies include:
[0096] Semantic vector recall: Calculating the query feature vector v query Compared with all text feature vectors v in the database text The cosine similarity is calculated using the following formula:
[0097]
[0098] Filter out similarity Sim cosine Canonical texts exceeding the threshold are selected as candidate set C1.
[0099] Keyword matching recall: The BM25 algorithm is used to match the query keywords corresponding to the user's query command with the standard text, and the top N text fragments with the highest BM25 scores are selected to form a candidate set C2.
[0100] Hybrid Similarity Recall: Combining the advantages of semantic vectors and keyword matching, a hybrid similarity calculation formula is proposed:
[0101]
[0102] Where α is the adjustment coefficient (experimentally determined to be optimal at 0.65), BM25 max and BM25 min These represent the minimum and maximum scores of the BM25 score, respectively. The candidate set C3 is generated by balancing the weights of semantic understanding and keyword matching using this formula.
[0103] The semantic vector recall candidate set C1, the keyword matching recall candidate set C2, and the hybrid similarity recall candidate set C3 are merged and deduplicated to obtain the multi-path recall candidate set C. text .
[0104] 2) Reordering phase
[0105] This stage employs a re-ranking model to refine the ranking of the recall results. BAAI General Embedding (BGE) is a general text embedding model trained on a large-scale corpus. The BGE-Rerank model is optimized for ranking tasks, enabling it to more accurately capture the semantic relationships between queries and candidate documents.
[0106] Score i =Rerank(Text i ,Query)
[0107] Among them, Text i Let be the i-th canonical text in the multi-way recall candidate set, Query be the user question, and Score be... i To score the i-th standardized text, the top K candidate results are obtained as the search results, and then the subsequent structured answer generation stage begins.
[0108] Step 103: Based on the search results, generate a structured answer corresponding to the user's query command.
[0109] The structured answer generation step is crucial for achieving accurate multimodal question-and-answer output in building codes. Based on the Langchain framework, it deeply integrates the language processing capabilities of the Qwen large-scale model, transforming the high-quality information output from the multi-path recall and re-ranking steps—i.e., the search results—into structured answers that meet user needs. This step ensures that the answers are both professionally accurate and meet the requirements of structured and visual presentation through carefully designed Prompt templates, strict response format constraints, and comprehensive content inspection mechanisms.
[0110] Specifically, Prompt Template Construction: The Prompt template is a core tool for guiding the large model to generate answers that meet expectations. The Prompt template designed in this step adopts a hierarchical, multi-dimensional constraint structure, organically combining user questions, search results, and output format requirements. Its specific structure is as follows:
[0111] P final =(P template Text, Query, P schema )
[0112] Among them, P templateFor the overall framework, Text represents the K reordered candidate documents, integrated into the template in descending order of relevance; Query represents the user's query text; and P... schema To constrain the output format, this Prompt template provides explicit input and output requirements for the large model, ensuring that the generated answers are structured and standardized. Response Schema Constraints: The Response Schema is used to strictly regulate the output format of the large model, ensuring that the generated content can be efficiently parsed and applied by the system. This step utilizes the response format control capabilities provided by the Langchain framework to define the content returned by the Qwen large model as JSON format and set two core key values:
[0113] Content: The value corresponding to this key is the text answer generated by the large model based on the Prompt, used to answer user questions about building codes. To ensure answer quality, constraints are placed on the length of the Content field and the use of technical terms. For example, the answer length is required to be no more than 500 characters, and key code provisions must be clearly cited, such as "according to Article 5.5.18 of GB50016-2014". Simultaneously, a technical terminology dictionary is established to automatically validate the content generated by the large model. If terms outside the dictionary or incorrect terms are found, an error correction mechanism is triggered.
[0114] URL List: The value corresponding to this key is a list of image URLs related to the text answer. The system extracts the storage URLs of the corresponding images and text determined during the reordering phase and populates this list. To ensure URL validity, URL format validation and access testing are performed before population. If an invalid URL is found, a new available image link is retrieved from the database. For example, when the storage path of an image changes, the system automatically updates the corresponding link in the URL List.
[0115] With strict Response Schema constraints, the generated JSON format answers are clear, easy to parse, and can be seamlessly integrated into subsequent display and application stages, providing users with intuitive and accurate multimodal information.
[0116] In one possible implementation of this application, this step further includes content generation and security detection. Specifically, the filled-in final template is input into the Qwen large model to generate the final JSON format answer. After obtaining the JSON format answer, the system automatically performs security detection on the text content of the Content field, including the following two aspects:
[0117] Sensitive word filtering: A dedicated sensitive word database for the construction industry has been established, covering sensitive information such as prohibited expressions, outdated regulations, and erroneous data. A combination of regular expressions and semantic analysis is used to scan the answer text word by word, and sensitive words are immediately corrected upon detection.
[0118] Compliance Verification: The answer content is compared with a database of currently valid building codes to verify the accuracy and validity of the code provisions cited in the answer. If an incorrect citation or use of obsolete codes is found in the answer, the system automatically retrieves the correct content from the database and replaces it.
[0119] Step 104: Perform a text-image consistency check on the structured answer. If the check passes, output the corresponding answer to the user's query command.
[0120] The image-text consistency verification step is a core step in ensuring the accuracy of multimodal question-and-answer results in building codes. It bears the important responsibility of deeply verifying the image and text content output by the structured answer generation step. If inconsistencies are found, the system will automatically trigger a re-search and answer generation process. This step mainly includes the following two stages:
[0121] 1) Image matching based on a list of URLs
[0122] From the JSON data output by the structured answer generation module, the system extracts the image URL list contained in the "URL List" field. This list details all image resource addresses associated with the answer content. Using this list as an index, the system performs precise image matching. Before retrieving images, the system performs secondary validity verification on the URLs. In addition to format validation, it also checks the HTTP response status code by initiating a lightweight HEAD request. If the status code is 200, the URL is valid; if it is 404 or another error code, the system immediately retrieves the latest URL for the image from the multimodal database and updates the "URL List" to prevent matching failures due to changes in image storage paths.
[0123] 2) Image-text consistency verification and answer regeneration
[0124] Establish a mapping database between building code clauses and graphic content. For the code clause number cited in the answer, retrieve the standard graphic set associated with that clause in the multimodal database, and compare the graphic in the answer with the standard graphic set to check whether the graphic meets the standard requirements of the code clause. If there is a difference, the graphic is determined to be inconsistent with the text.
[0125] The maximum number of retries for answer regeneration is set to 3. When the consistency verification of images and text fails, the system immediately clears the current search cache to avoid incorrect results due to repeated searches. Based on the user's original query and the specific information of the inconsistency between images and text, such as contradictory keywords and conflicting regulatory clauses, the system adjusts the search strategy, increases the weight of search keywords, expands the search scope, and re-executes multi-way recall and re-ranking operations.
[0126] Through the collaborative work of the above four steps, this application realizes a closed-loop process from multimodal data storage and updating, information retrieval and sorting, structured answer generation to text-image consistency verification, and builds an efficient and accurate multimodal question-and-answer system based on a large text-based question-and-answer model. This effectively improves the accuracy and query efficiency of building code illustration questions and answers, greatly reduces the manual cost of code consultation, and provides strong support for the intelligent application of building industry code information.
[0127] Based on the same inventive concept, this application also provides a Langchain-based multimodal question-answering device for architectural specification diagrams, the structure of which is as follows: Figure 3 As shown.
[0128] Figure 3 This application provides a structural schematic diagram of a multimodal question-and-answer device based on Langchain for illustrating building codes. Figure 3 As shown, the Langchain-based multimodal question-answering device 300 for architectural specifications diagrams in this application specifically includes: at least one processor 301; and a memory 303 communicatively connected to the at least one processor 301 (connected via a bus 302); wherein the memory 303 stores instructions executable by the at least one processor 301 to enable the at least one processor 301 to execute a Langchain-based multimodal question-answering method for architectural specifications diagrams as described in the above embodiments.
[0129] In one possible implementation of this application, the aforementioned processor 301 is used to perform the following actions: constructing a multimodal database of building codes and storing and managing the multimodal data of building codes; searching the multimodal database from different dimensions based on user query instructions; generating a structured answer corresponding to the user query instructions based on the search results; performing a text-image consistency check on the structured answer; and outputting the answer corresponding to the user query instructions after the check passes.
[0130] This application proposes a Langchain-based multimodal question-answering method and device for architectural code illustrations. It innovatively utilizes the Langchain large-model structured output framework to endow the text-based question-answering large-model system with image-text matching capabilities, enabling multimodal intelligent question-answering for architectural industry illustrations. The solution constructs a complete question-answering system technology architecture, covering four core steps: multimodal data storage and updating, multi-path recall and reordering, large-model structured generation based on the Langchain framework, and image-text matching. A specially designed dynamic update mechanism for the multimodal architectural code database effectively addresses the impact of industry standard iterations on question-answering performance. Compared to traditional text-based large-model systems, the multimodal question-answering system built using this method significantly improves the accuracy and query efficiency of architectural code illustration question-answering, greatly reduces the manual cost of standard consultation, and provides an efficient solution for intelligent information retrieval and application in the architectural industry.
[0131] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0132] The equipment and method provided in this application are one-to-one correspondences. Therefore, the equipment also has similar beneficial technical effects as its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the equipment will not be repeated here.
[0133] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0134] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0135] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A multimodal question-answering method for building code illustrations based on Langchain, characterized in that, The method includes: Construct a multimodal database of building codes and store and manage the multimodal data of building codes; Based on user query commands, the multimodal database is retrieved from different dimensions; Based on the search results, generate a structured answer corresponding to the user's query. The structured answer is validated for consistency between text and image. If the validation passes, the corresponding answer to the user's query command is output.
2. The Langchain-based multimodal question-answering method for architectural code illustrations according to claim 1, characterized in that, The multimodal data of the building code includes building code text and building code illustrations; Storage and management of multimodal data related to building codes, including: The data storage fields include NAME, CONTENT, VECTOR, URL, BASE64, VAILD, and UPDATE fields; The NAME field stores the name and number of the building code illustration; the CONTENT field stores the building code text in UTF-8 encoding; the VECTOR field stores the semantic vector of the building code text with a preset dimension generated by a pre-trained model; the URL field stores the address of the building code illustration corresponding to the building code text; the BASE64 field stores the Base64 encoded string of the building code illustration; the VALID field is a boolean type, where true indicates that the building code multimodal data is valid, and false indicates that the building code multimodal data is obsolete; and the UPDATE field records the update time of the building code multimodal data.
3. The Langchain-based multimodal question-answering method for architectural code illustrations according to claim 2, characterized in that, Storage and management of multimodal data related to building codes, including: For building code texts, a deep learning-based text vectorization model is used to extract and transform deep semantic features, including: The building code text is segmented into sentences or paragraphs and then input into the BERT model so that the BERT model can generate a corresponding hidden state vector for each input token; Through average pooling, the hidden state vectors corresponding to all tokens are merged into a fixed-length text feature vector v. text ; Obtain the text feature vector v text Then, the building code text is stored in the CONTENT field, and the text feature vector v is also stored in the CONTENT field. text Store in the VECTOR field.
4. The Langchain-based multimodal question-answering method for architectural code illustrations according to claim 2, characterized in that, Storage and management of multimodal data related to building codes, including: For building code drawings, a storage strategy combining distributed storage and Base64 encoding conversion is adopted, including: A unique storage path is generated for each building code illustration, and the storage path is stored in the URL field in URL format to form the network access identifier of the building code illustration; Read the binary data corresponding to the building code drawing, convert the binary data into a string according to the Base64 encoding rules, and store the string in the BASE64 field.
5. The Langchain-based multimodal question-answering method for architectural code illustrations according to claim 1, characterized in that, Constructing a multimodal database of building codes, and storing and managing the multimodal data of building codes, also includes: Obtain the latest building code text; By using a text comparison algorithm, the latest building code text is compared with the current building code text line by line, and the newly added, modified and deleted content is automatically identified to obtain updated data; The process of updating the multimodal database of building codes is triggered by first updating the building code text to a vectorized form, and then regenerating the text feature vector v. text Then fill in the VECTOR field, update the text content of the CONTENT field, replace the building code illustration with the correct version, and update the corresponding URL address. After the update is complete, a notification of the data change is sent via a message queue.
6. The Langchain-based multimodal question-answering method for architectural code illustrations according to claim 1, characterized in that, Based on user query commands, the multimodal database is retrieved from different dimensions, including: The user's query command is input into the BERT model to generate the query feature vector v corresponding to the user's query command. query ; Calculate the query feature vector v query Compared with all text feature vectors v in the building code multimodal database text Cosine similarity between them; Architectural code texts with similarity values exceeding a preset threshold are selected as candidate set C1; The BM25 algorithm is used to match the keywords of the user's query with the building code text, and the top N building code text fragments with the highest BM25 scores are selected to form a candidate set C2; Using the hybrid similarity calculation formula, candidate set C3 is generated by combining cosine similarity and BM25 score; Merge candidate sets C1, C2, and C3, remove duplicates, and obtain the multi-path recall candidate set C. text ; The multi-path recall candidate set C is reordered using a re-ranking model. text Sort the results and select the top K candidate results by score as the search results.
7. A multimodal question-answering method for building code illustrations based on Langchain according to claim 1, characterized in that, Based on the search results, a structured answer corresponding to the user's query is generated, including: The Prompt template is obtained by combining the user's query command, search results and output format requirements; Using the Langchain framework, the return content of the Qwen large model is defined as JSON format, and two key values are set: Content and URL List; The value corresponding to Content is the text answer generated by the Qwen big model based on the Prompt template, used to answer the user's query command about building codes. The value corresponding to URL List is a list containing image URLs related to the text answer.
8. A multimodal question-answering method for building code illustrations based on Langchain according to claim 7, characterized in that, After obtaining the returned content in JSON format, the method further includes: Sensitive word filtering and compliance verification are performed on the text answers in the Content field.
9. A multimodal question-answering method for building code illustrations based on Langchain according to claim 8, characterized in that, The structured answer is subjected to a consistency check of text and image, including: Extract the list of image URLs contained in the URL List field; Establish a mapping database between building code texts and building code illustrations; For the building code clause number referenced in the structured answer, retrieve the set of building code illustrations associated with the building code clause in the multimodal building code database; The set of building code illustrations is compared with the set of illustrations corresponding to the list of image URLs. If there are differences between the two, it is determined that the text and images are inconsistent.
10. A multimodal question-and-answer device for building code illustrations based on Langchain, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, enabling the at least one processor to perform the following: Construct a multimodal database of building codes and store and manage the multimodal data of building codes; Based on user query commands, the multimodal database is retrieved from different dimensions; Based on the search results, generate a structured answer corresponding to the user's query. The structured answer is validated for consistency between text and image. If the validation passes, the corresponding answer to the user's query command is output.
Citation Information
Patent Citations
Streaming question and answer illustration method and system
CN118035416A
Question and answer system construction method fusing large language model and domain knowledge
CN118113832A
Building standard knowledge question and answer model construction method, computer program product, storage medium and electronic equipment
CN118296118A
Solution retrieval method applied to multi-path question-answering system in field of electric power design
CN118568236A
Knowledge base question and answer method and device and computer readable storage medium
CN119128096A
Cited By
Building specification graphical representation retrieval method, equipment and medium
CN121524307A