Construction image-oriented multi-modal knowledge retrieval method and system

Through the multimodal knowledge retrieval method, the construction images are analyzed using intelligent large models and multimodal coding models, vectorized data are generated and dynamically adapted to the knowledge base topics, which solves the shortcomings of information mining and user intention understanding in traditional methods, and realizes accurate knowledge content recall and recommendation.

CN120448567AActive Publication Date: 2025-08-08ZERO-YI TONGZHI (QINGDAO) DIGITAL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510285637.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-08-08
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Traditional construction image analysis methods are difficult to dig deep into multi-level information, cannot effectively integrate multi-field knowledge, and lack a deep understanding of user query intentions, resulting in low matching of search results with requirements, and it is difficult to adapt to changes in complex construction scenarios.

Method used

The multimodal knowledge search method is adopted to analyze construction images through intelligent large models, generate related images and perform vectorization processing, combine multimodal coding models to match the knowledge base topics, and use intelligent sorting recommendation algorithm to output accurate answers, and dynamically adapt the search range.

Benefits of technology

It realizes the accuracy and context-related knowledge content recall of construction images, improves the level of intelligence, can output the most relevant knowledge content in cross-domain application scenarios, and improves retrieval efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448567A_ABST
    Figure CN120448567A_ABST
Patent Text Reader

Abstract

The invention provides a construction image-oriented multi-modal knowledge retrieval method and system, and the method comprises the steps: obtaining a construction image, and carrying out the data enhancement processing of the construction image, so as to generate a related construction image; analyzing the construction image by using the intelligent large model to extract a construction subject, a knowledge field and an application scene of the construction image, and determining a knowledge base theme related to the construction image; utilizing a multi-modal coding model to perform vectorization processing on related construction images to generate semantic vectors which serve as retrieval vectors to recall a matched first knowledge item set; taking a construction subject of the construction image as a query statement to recall a matched second knowledge item set; performing intelligent sorting recommendation on knowledge entries in the knowledge entry set; and selecting knowledge entries ranking the top K, constructing a prompt project according to the analysis content of the construction image, the construction image and the knowledge entries ranking the top K, performing task analysis by using the intelligent large model again, and finally outputting accurate and context-related answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal knowledge retrieval method and system for construction images. Background Art

[0002] As modern engineering construction expands in scale and complexity, the need for information-based management of construction sites is becoming increasingly important. Construction images, a crucial tool for documenting project progress, managing quality, and overseeing safety, contain a wealth of on-site information and have become critical data in construction management. However, with the rapid growth of construction site image data, traditional image analysis and retrieval technologies are increasingly exposed to deficiencies, exhibiting significant limitations in analyzing complex scenes, linking multi-domain knowledge, and meeting personalized user needs.

[0003] Traditional image analysis methods mostly rely on single-modality feature extraction techniques. Analysis of image content is often limited to surface features, making it difficult to deeply explore the multi-layered information contained within an image. For example, in a construction scene, images may simultaneously contain multiple elements such as machinery and equipment, building components, and human activity. However, existing technologies struggle to comprehensively identify these elements and their associated characteristics. Furthermore, identifying potential issues in construction images (such as safety hazards or quality defects) often requires manual intervention, resulting in a low degree of automation and insufficient real-world application. Existing technologies are particularly inadequate in understanding and analyzing multi-domain characteristics in construction scenarios involving multiple disciplines (such as civil engineering and electromechanical installation). Knowledge bases in the construction field typically cover multiple specialized areas, such as concrete construction, steel structure engineering, and earthwork operations, resulting in complex content and unclear categorization. Existing knowledge retrieval systems often rely on keyword or tag matching, lacking a deep understanding of user query intent, resulting in poor matching of search results with user needs. Current retrieval methods often struggle to dynamically adapt to changes in the construction scene and are unable to adjust the knowledge base's search scope based on the specific content in the image. In addition, for complex construction scenarios involving multiple fields, traditional methods find it difficult to effectively integrate knowledge items from different fields, resulting in limited retrieval efficiency and accuracy of the system. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed to provide a multimodal knowledge retrieval method and system for construction images that overcome the above problems.

[0005] One aspect of the present invention provides a multimodal knowledge retrieval method for construction images, the method comprising:

[0006] Obtain user question data, where the question data is a construction image, and perform data enhancement processing on the construction image to generate n related construction images, where n is greater than or equal to 2;

[0007] The construction image is parsed using a preset intelligent large model to extract the parsed content of the construction image. The target knowledge base topics involved in the construction image are determined based on the parsed content. The parsed content includes construction subjects, knowledge fields, and application scenarios.

[0008] Using a preset multimodal encoding model, each relevant construction image obtained after data augmentation is vectorized to generate a first semantic vector. Each first semantic vector is used as a first retrieval vector to perform similarity calculations with vector data stored in the knowledge base corresponding to the target knowledge base topic, so as to recall a first set of knowledge items that match the construction image.

[0009] Taking the construction subject of the construction image as a query statement, obtaining knowledge items matching the construction subject from a knowledge base corresponding to the target knowledge base subject according to the query statement, so as to recall a second set of knowledge items matching the construction image;

[0010] Intelligently sort and recommend knowledge items in the first knowledge item set and the second knowledge item set;

[0011] The top K knowledge items are selected from the intelligent sorting recommendation results. A prompt project is constructed based on the parsed content of the construction image, the construction image, and the top K knowledge items. The intelligent big model is then used again in combination with the prompt project to perform task analysis on the construction image to obtain the final retrieval results.

[0012] Furthermore, the target knowledge base topics involved in the construction image are determined based on the parsed content of the construction image, including:

[0013] Determining a first knowledge base topic involved in the construction image according to the knowledge domain and application scenario of the construction image;

[0014] Semantically encoding the construction image to generate a corresponding semantic vector, matching the semantic vector with vector data in a preset knowledge base, and determining the subject of the knowledge base to which the vector data matching the semantic vector belongs as the second knowledge base subject involved in the construction image;

[0015] Perform visual feature recognition on the construction image, search a predefined matching relationship table to obtain application scenarios and / or knowledge domains that match the visual features, and determine the third knowledge base topic related to the construction image based on the application scenarios and / or knowledge domains that match the visual features; the matching relationship table includes the correspondence between the visual features of the image and the application scenarios and / or knowledge domains;

[0016] The first knowledge base topic, the second knowledge base topic and the third knowledge base topic are used as target knowledge base topics involved in the construction image.

[0017] Furthermore, the construction subject of the construction image is used as a query statement, and knowledge items matching the construction subject are obtained from the knowledge base corresponding to the target knowledge base subject according to the query statement, including:

[0018] The construction subject of the construction image is used as a query statement, and the query statement is vectorized using a preset multimodal encoding model to generate a second semantic vector. The second semantic vector is used as a second retrieval vector and a similarity calculation is performed with the vector data stored in the knowledge base corresponding to the target knowledge base topic to recall a first sub-set of knowledge items that match the construction image;

[0019] The construction subject of the construction image is used as a query statement, keywords in the query statement are extracted, and similarity calculation is performed between the keywords and keywords stored in the knowledge base corresponding to the target knowledge base topic to recall the second sub-knowledge item set matching the construction image.

[0020] Furthermore, performing intelligent sorting and recommendation on the knowledge items in the first knowledge item set and the second knowledge item set includes:

[0021] Obtaining interest model data of the current user, and performing a first recommendation score on the knowledge items in the first knowledge item set and the second knowledge item set according to the interest model data;

[0022] performing content matching with the historical search output results of the current user based on the content features of the knowledge items in the first knowledge item set and the second knowledge item set, and performing a second recommendation score based on the matching results;

[0023] Calculate the final score of each knowledge item based on the preset scoring method weight ratio and the first recommended score and second recommended score corresponding to each knowledge item;

[0024] Intelligently sort and recommend each knowledge item based on the final score.

[0025] Furthermore, the current user's interest model data is obtained, including:

[0026] Obtain the current user's user role, the project information the user is responsible for, and the user's historical operation records. The historical operation records include the user's historical evaluation records of the content generated by the intelligent big model, as well as the user's historical query, click, and browsing records;

[0027] Generate a co-occurrence matrix based on the user role, project information, and historical operation records, and find K users most similar to the current user based on the co-occurrence matrix;

[0028] Based on the historical operation records of the K users on the content generated by the intelligent big model, the interest preferences of the K users in knowledge base topics in different knowledge fields and / or application scenarios are analyzed, and the current user's interest preferences in knowledge items of different knowledge base topics are predicted based on the interest preferences of the K users to obtain the current user's interest model data.

[0029] Furthermore, performing data enhancement processing on the construction image to generate n related construction images includes:

[0030] Positioning a foreground image region of the construction image, extracting the foreground image region from the construction image, and sequentially splicing the foreground image region into different preset background images to obtain a plurality of related construction images; and / or

[0031] Randomly cropping different sub-regions of the construction image to obtain multiple related construction images; and / or

[0032] Performing at least one geometric transformation of multi-angle rotation, horizontal flipping, vertical flipping, and affine transformation on the construction image to simulate construction images from different perspectives, thereby obtaining a plurality of related construction images; and / or

[0033] Utilizing the OpenCV library to optimize and adjust the color, brightness, contrast, and / or saturation of the construction images to obtain multiple related construction images; and / or

[0034] Performing different degrees of denoising and blurring on the construction images to obtain multiple related construction images; and / or

[0035] Different key features in the construction image are highlighted using edge detection and texture enhancement techniques to obtain multiple related construction images.

[0036] Furthermore, before obtaining the user question data, the method further includes:

[0037] Constructing a knowledge base for knowledge retrieval, wherein the knowledge base is divided into knowledge base themes according to knowledge domains and application scenarios in the construction field, and the knowledge base includes a metadata database and a vector database;

[0038] The constructing of a knowledge base for knowledge retrieval includes:

[0039] Each text data in the preset construction field document dataset is divided into several knowledge base topics according to the knowledge domain and application scenarios of the construction field, and the text data corresponding to each knowledge base topic is determined;

[0040] Each text data is segmented to form a knowledge item dataset with complete semantic information;

[0041] Perform word segmentation processing on each knowledge item in the knowledge item data set corresponding to each text data to extract the keywords included in each knowledge item, assign a unique ID to each knowledge item in the knowledge item data set, establish a correspondence between each keyword of the knowledge item and the unique ID, and store the keywords included in each knowledge item and the correspondence between each keyword and the unique ID in a metadata database corresponding to the text data;

[0042] Semantic encoding is performed on each knowledge item in the knowledge item data set corresponding to each text data to generate corresponding vector data, and the vector data is stored in the vector database corresponding to the text data.

[0043] Furthermore, after obtaining the user question data, the method further includes:

[0044] A picture-specific link in the form of a link is generated for the construction image, and an effective use time is set for the picture-specific link, so that the intelligent large model can call the construction image through the picture-specific link within the effective use time.

[0045] Another aspect of the present invention provides a multimodal knowledge retrieval system for construction images, which includes a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor. The processor executes the computer program / instructions to implement the steps of the above-mentioned multimodal knowledge retrieval method for construction images.

[0046] Another aspect of the present invention further provides a computer program product having a computer program stored thereon, which implements the steps of the above-mentioned multimodal knowledge retrieval method for construction images when executed by a processor.

[0047] The multimodal knowledge retrieval method and system for construction images provided by the embodiments of the present invention supports dynamic adaptation of multiple knowledge bases for knowledge content recall, and can flexibly select the search scope based on the identified knowledge base topics. For cross-domain application scenarios, the present invention can simultaneously search multiple knowledge bases. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content in complex scenarios. Furthermore, the present invention deeply integrates the image analysis content with the recall results of the first retrieval, and then inputs them again into the intelligent large model to perform task retrieval, thereby guiding the model's generation capabilities to the precise field of user needs, and ultimately outputting accurate and context-related answers.

[0048] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Various other advantages and benefits will become apparent to those skilled in the art by reading the detailed description of the preferred embodiment below. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention. In the accompanying drawings:

[0050] Figure 1 This is a flowchart of a multimodal knowledge retrieval method for construction images according to an embodiment of the present invention;

[0051] Figure 2 This is a specific implementation flowchart of a multimodal knowledge retrieval method for construction images according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0053] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.

[0054] This invention provides a multimodal knowledge retrieval method for construction images, overcoming existing limitations. Traditional image retrieval techniques struggle to comprehensively and accurately extract key information from construction images, especially when multiple specialized fields are involved. Furthermore, this invention effectively enhances the intelligence of knowledge retrieval for construction images, moving beyond keyword or tag matching to understanding the deeper meaning behind images in conjunction with user behavior, thereby providing more precise answers and services.

[0055] Example 1

[0056] The embodiment of the present invention provides a multimodal knowledge retrieval method for construction images, such as Figure 1 As shown, the multimodal knowledge retrieval method for construction images proposed in the present invention includes the following steps:

[0057] S11. Obtain user question data, where the question data is a construction image, and perform data enhancement processing on the construction image to generate n related construction images, where n≥2.

[0058] The present invention provides a multimodal knowledge retrieval method for construction images, in which the user's question data is the construction image uploaded by the user. The present invention provides a user-friendly image upload interface to obtain construction images. Specifically, the image upload interface provided by the present invention supports the upload of construction images in common image formats, and ensures that the format of the uploaded file is correct and the size is reasonable by configuring a preset file verification mechanism. Furthermore, after obtaining the user's question data, that is, after obtaining the uploaded construction image, a picture-specific link in the form of a link can be generated for the construction image, and an effective use time can be set for the picture-specific link so that the intelligent large model can call the construction image through the picture-specific link within the effective use time. In a specific embodiment, after the picture is uploaded, the system can generate a picture-specific link with a validity period of one hour for the large model to call and process. After the expiration, the link will become invalid, thereby ensuring that the picture will not be exposed to the public network for a long time, thereby improving the security of image retrieval.

[0059] Furthermore, in order to improve the recall accuracy of the multimodal knowledge base, the present invention also implements a series of preprocessing steps on the construction image before inputting the construction image into the multimodal encoding model to generate multiple related construction images that are related or similar to the construction image, so as to realize data enhancement processing of the construction image. Specifically, the data enhancement processing of the construction image in this step to generate n related construction images specifically includes: locating the foreground image area of the construction image, extracting the foreground image area from the construction image, and sequentially splicing the foreground image area into different preset background images to obtain multiple related construction images; and / or randomly cropping different sub-areas of the construction image to obtain multiple related construction images; and / or performing at least one geometric transformation of multi-angle rotation, horizontal flipping, vertical flipping and affine transformation on the construction image to simulate construction images from different perspectives to obtain multiple related construction images; and / or using the OpenCV library to optimize and adjust the color, brightness, contrast and / or saturation of the construction image to obtain multiple related construction images; and / or performing different degrees of denoising and blurring processing on the construction image to obtain multiple related construction images; and / or using edge detection and texture enhancement technology to highlight different key features in the construction image to obtain multiple related construction images.

[0060] The data enhancement processing in this embodiment includes color correction of the original image, adjusting brightness, contrast, and saturation to reduce the impact of ambient light on the image. In a specific embodiment, the present invention may also consider converting the original image into a grayscale image to eliminate the potential interference of color information on knowledge recall. In terms of geometric transformation, the present invention may also rotate the original image at multiple angles (for example, ±15°, ±30°, ±45°), as well as flip it horizontally and vertically to simulate different shooting angles and directions. At the same time, affine transformation is also applied to slightly adjust the shape of the image to enhance the model's adaptability to different viewing angles. To address the noise problem, the present invention may also use a denoising algorithm to reduce random noise in the image and use a slight blurring process to reduce the impact of detail differences. Edge detection and texture enhancement techniques may also be used to highlight key features in the image to improve the recall effect of the knowledge base. In a specific embodiment, in addition to the above data enhancement processing methods, the present invention may also randomly crop sub-regions of the image and / or mix the foreground image area of the image with different backgrounds to increase the diversity of samples and improve the generalization ability of the knowledge base.

[0061] S12. Analyze the construction image using a pre-set intelligent large model to extract its analytical content. Based on the analytical content, the target knowledge base topics related to the construction image are determined. The analytical content includes construction subjects, knowledge domains, and application scenarios. The intelligent large model is fine-tuned and trained based on the open-source model llama3.2-vision:11b.

[0062] In this embodiment, the intelligent big model is first used to extract knowledge domain and application scenario information from user-uploaded construction images through a prompt project. This information is then used to determine the relevant knowledge base topics, laying the foundation for subsequent knowledge retrieval and prompt enhancement, thereby achieving more refined semantic extraction and question answering. Specifically, the preset intelligent big model performs multi-level analysis of the image content, including extracting the construction subject, knowledge domain, application scenario, and potential problems in the image. Based on the analysis results, the system then identifies the potentially relevant knowledge base topics and corresponding knowledge bases. Based on the identified knowledge base topics, the system then automatically queries the relevant knowledge bases.

[0063] In an embodiment of the present invention, the knowledge base is divided into knowledge base topics according to the knowledge fields and application scenarios in the construction field, and the knowledge base includes a metadata database and a vector database. Specifically, when constructing the knowledge base, it is divided into several knowledge base topics according to the knowledge fields and application scenarios in the construction field. For example, the knowledge base topics of the knowledge fields include concrete engineering, earthwork engineering, steel structure engineering, etc. The knowledge base topics of the application scenarios include safety management, construction specifications, etc. Each knowledge base topic covers a specific knowledge field and application scenario to ensure the clarity and completeness of the knowledge classification. This classification method facilitates the rapid matching of relevant content in subsequent searches. In a specific embodiment, the knowledge fields can be divided into: foundation engineering, masonry structure engineering, waterproofing engineering, HVAC engineering, electrical engineering, water supply and drainage engineering, curtain wall engineering, fire protection engineering, concrete construction, steel structure engineering and earthwork engineering; the application scenarios can be divided into: safety management, construction specifications, quality management and progress management.

[0064] Furthermore, the target knowledge base topic involved in the construction image is determined based on the parsed content of the construction image, including: determining a first knowledge base topic involved in the construction image based on the knowledge domain and application scenario of the construction image in the parsed results; semantically encoding the construction image to generate a corresponding semantic vector, and matching the semantic vector with vector data in a preset knowledge base, determining the topic of the knowledge base to which the vector data matching the semantic vector belongs as a second knowledge base topic involved in the construction image; performing visual feature recognition on the construction image, searching a predefined matching relationship table to obtain the application scenario and / or knowledge domain matching the visual feature, and determining a third knowledge base topic involved in the construction image based on the application scenario and / or knowledge domain matching the visual feature; the matching relationship table includes the correspondence between the visual features of the image and the application scenario and / or knowledge domain; and determining the first knowledge base topic, the second knowledge base topic, and the third knowledge base topic as the target knowledge base topic involved in the construction image. Specifically, after the user uploads the image, the system combines the image with a pre-set specific prompt word template Prompt input and uses a multimodal large model to perform a preliminary analysis of the image content. The Prompt project plays a key role here. By designing task prompts for the construction field, it clearly states what information the model needs to extract from the image, such as the construction subject (such as mechanical equipment, building components), application scenarios (such as construction environment, spatial layout) and knowledge domain. At the same time, the model can also automatically associate it with the knowledge base topics that may be involved by parsing the semantic information in the image content. Furthermore, the model can also match it to pre-defined application scenarios and / or knowledge domains based on the elements and features in the image, such as concrete construction corresponding to the concrete knowledge base, and equipment operation corresponding to the mechanical equipment knowledge base.

[0065] S13. Use a preset multimodal encoding model to vectorize each relevant construction image obtained after data enhancement processing to generate a first semantic vector, and use each first semantic vector as a first retrieval vector to perform similarity calculation with the vector data stored in the knowledge base corresponding to the target knowledge base topic, so as to recall the first set of knowledge items that match the construction image.

[0066] S14. Using the construction subject of the construction image as a query statement, obtaining knowledge items matching the construction subject from a knowledge base corresponding to the target knowledge base subject according to the query statement, so as to recall a second set of knowledge items matching the construction image.

[0067] S15. Intelligently sort and recommend the knowledge items in the first knowledge item set and the second knowledge item set.

[0068] S16. Select the top K knowledge items from the intelligent sorting recommendation results, build a prompt project based on the analysis content of the construction image, the construction image, and the top K knowledge items, and again use the intelligent big model combined with the prompt project to perform task analysis on the construction image to obtain the final retrieval results.

[0069] In this embodiment, the image analysis content is deeply integrated with the refined knowledge items and then re-input into the multimodal large model to generate accurate and context-relevant answers. Specifically, in the fusion stage, the system will select the top K knowledge items from the image analysis content and the intelligent sorting recommendation results to construct a new prompt template Prompt, and clearly specify the problem or task that the model needs to solve. The new prompt template Prompt will integrate the construction subject, application scenario, knowledge field and refined knowledge items extracted from the image, guide the model's generation capabilities to the precise field of user needs, and ultimately output accurate answers.

[0070] The multimodal knowledge retrieval method for construction images provided by an embodiment of the present invention supports dynamic adaptation of multiple knowledge bases for knowledge content recall, and can flexibly select the search scope based on the identified knowledge base topics. For cross-domain application scenarios, the present invention can simultaneously search multiple knowledge bases. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content in complex scenarios. Furthermore, the present invention deeply integrates the image analysis content with the recall results of the first retrieval, and then inputs the results into the intelligent large model to perform task retrieval again, thereby guiding the model's generation capabilities to the precise field of user needs, and ultimately outputting accurate and context-related answers.

[0071] In an embodiment of the present invention, before obtaining user question data, it is necessary to construct a knowledge base for knowledge retrieval, wherein the knowledge base is divided into knowledge base themes according to the knowledge domain and application scenario of the construction field, and the knowledge base includes a metadata database and a vector database. Specifically, constructing a knowledge base for knowledge retrieval includes: dividing each text data in a preset construction field document data set into a number of knowledge base themes according to the knowledge domain and application scenario of the construction field, and determining the text data corresponding to each knowledge base theme; performing text segmentation on each text data to form a knowledge item data set with complete semantic information; performing word segmentation on each knowledge item in the knowledge item data set corresponding to each text data to extract the keywords included in each knowledge item, and assigning a unique ID to each knowledge item in the knowledge item data set, establishing a correspondence between each keyword of the knowledge item and the unique ID, and storing the keywords included in each knowledge item and the correspondence between each keyword and the unique ID in the metadata database of the corresponding text data; performing semantic encoding on each knowledge item in the knowledge item data set corresponding to each text data to generate corresponding vector data, and storing it in the vector database of the corresponding text data.

[0072] In this embodiment, the "knowledge base" has two storage forms, text and vector. Each knowledge base first obtains knowledge items by document splitting, and assigns a unique ID to each knowledge item. The keywords in the knowledge item are obtained through jieba word segmentation, and all keywords point to this unique ID, which is convenient for full-text indexing. The knowledge item semantics are then encoded using the open source MiniCPM-V model to generate vector data, and stored in the PostgreSQL database. At the same time, the text data corresponding to the knowledge item is stored in the metadata, which is convenient for retrieving more accurate knowledge items when using hybrid retrieval. Among them, documents refer to professional documents, specifications, guidelines, etc. related to the construction field. Segmentation is to parse the content of the document and divide it into independent knowledge units (i.e. knowledge items). In a specific example, the specific segmentation method can be: extract the text in the document, pre-segment it according to the length standard of 600, and search for any one of "period, semicolon, line break, tab, etc." within the range of 50 characters above and below, and split it here; if not found, split it at the 600th position.

[0073] In an embodiment of the present invention, recalling knowledge items that match construction images is to vectorize multiple related construction images after data enhancement processing through a multimodal coding model, and at the same time automatically query the relevant knowledge base in combination with the target knowledge base topics extracted in the preliminary recognition stage to recall the matching knowledge items.

[0074] Specifically, a batch of relevant construction images after data enhancement processing are first generated into a high-dimensional semantic vector representation through the MiniCPM-V multimodal coding model. The multimodal coding model can integrate the visual features of the image with the semantic information behind it, and convert the complex content of the image into a quantifiable vector representation. Subsequently, the generated vector is used as the first retrieval vector to perform similarity calculation with the predefined vector in the selected knowledge base (that is, the knowledge base corresponding to the target knowledge base topic). Each knowledge entry in the knowledge base has been vectorized during construction and stored in the vector database. Through the efficient KNN nearest neighbor algorithm, the system can quickly recall the knowledge entries with the highest relevance to the image content. For example, if the picture involves "formwork erection", the system may recall entries related to formwork fixing methods, material specifications, and construction steps.

[0075] Furthermore, in step S14, using the construction subject of the construction image as a query statement and obtaining knowledge items matching the construction subject from the knowledge base corresponding to the target knowledge base topic based on the query statement specifically includes: using the construction subject of the construction image as the query statement, vectorizing the query statement using a preset multimodal encoding model to generate a second semantic vector, using the second semantic vector as a second retrieval vector, and performing similarity calculation with vector data stored in the knowledge base corresponding to the target knowledge base topic to recall a first sub-set of knowledge items matching the construction image; using the construction subject of the construction image as the query statement, extracting keywords from the query statement, and performing similarity calculation with keywords stored in the knowledge base corresponding to the target knowledge base topic to recall a second sub-set of knowledge items matching the construction image. The first sub-set of knowledge items and the second sub-set of knowledge items are combined into a second set of knowledge items.

[0076] In this embodiment, in addition to direct retrieval based on image vectorization, the present invention also uses the construction subject (such as "concrete wall" or "reinforced steel material") identified by the intelligent large model as a query statement to perform a hybrid search in the selected knowledge base (i.e., the knowledge base corresponding to the target knowledge base subject). The hybrid search includes: first, inputting these query statements into the MiniCPM-V model for vectorization to obtain a query vector, using the KNN nearest neighbor algorithm to calculate the vectors of the K closest knowledge items, and obtaining several knowledge items with the highest correlation therewith. The K value can be set according to user needs and can be selected as 5, 10, 20, 30 or other values. The present invention does not make specific limitations on this; second, segmenting the query statement through jieba word segmentation to extract keywords in the query, and then calculating the correlation between the keywords and the knowledge item keywords through the BM25 algorithm and sorting them, and obtaining several knowledge items with the highest correlation therewith.

[0077] The knowledge content recall function of the present invention supports dynamic adaptation of multiple knowledge bases, and can flexibly select the search scope according to the identified knowledge fields and application scenarios. For example, for concrete application scenarios, the system prioritizes searching the concrete construction specification library; for earthwork operations, it queries the earthwork engineering safety management library. In this way, the system avoids interference from irrelevant content, improves retrieval efficiency and the accuracy of recall results. In addition, for cross-domain application scenarios (such as scenarios involving both mechanical equipment and building components), the system can simultaneously search multiple knowledge bases and assign weights according to the importance of the recall results. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content in complex scenarios. Among them, the assigned weights refer to assigning different weights to the knowledge items queried based on the construction subject information, the construction subject information vector, and the pre-processed image vector to obtain the final knowledge items.

[0078] In an embodiment of the present invention, the intelligent sorting and recommendation of knowledge items in the first knowledge item set and the second knowledge item set in step S15 includes: obtaining the interest model data of the current user, and performing a first recommendation score on the knowledge items in the first knowledge item set and the second knowledge item set based on the interest model data; performing content matching with the historical retrieval output results of the current user based on the content features of the knowledge items in the first knowledge item set and the second knowledge item set, and performing a second recommendation score based on the matching results; calculating the final score of each knowledge item based on the preset scoring method weight ratio and the first recommendation score and the second recommendation score corresponding to each knowledge item, wherein the weight ratio of the first recommendation score and the second recommendation score can be optionally 4:6; and intelligently sorting and recommending each knowledge item based on the final score. Among them, obtaining the interest model data of the current user specifically includes: obtaining the user role of the current user, the project information for which the user is responsible, and the user's historical operation records, the historical operation records including the user's historical evaluation records of the content generated by the intelligent big model and the user's historical query, click, and browsing records; generating a co-occurrence matrix based on the user role, project information, and historical operation records, and searching for the K users most similar to the current user based on the co-occurrence matrix; analyzing the interest bias of the K users in knowledge base topics in different knowledge fields and / or application scenarios based on the historical operation records of the K users on the content generated by the intelligent big model, and predicting the current user's interest bias in knowledge items of different knowledge base topics based on the interest bias of the K users, to obtain the interest model data of the current user.

[0079] In actual construction management, different user roles have significantly different requirements for knowledge content. For example, project managers are more concerned with progress and cost, while construction supervisors prioritize safety. Existing systems generally lack the ability to identify and personalize user roles, making it difficult to provide accurate recommendations based on a user's historical operation history, current scenario requirements, or project type. Furthermore, existing ranking algorithms are often relatively simple and unable to dynamically adjust the priority of recommended results based on user preferences, resulting in a poor user experience. Furthermore, as construction management evolves towards intelligent and automated processes, the shortcomings of existing traditional systems in intelligent retrieval, recommendation, and dynamic adaptation are becoming increasingly prominent. Existing technologies lack intelligent ranking mechanisms that consider user scenario requirements when recommending knowledge items, resulting in recommendations that often deviate from actual user needs. For complex application scenarios, the system cannot dynamically adapt the search scope of cross-domain knowledge bases, making it difficult to provide comprehensive and accurate answers to users. These limitations significantly reduce the efficiency of construction management systems and user satisfaction. To address these issues, the present invention utilizes an intelligent recommendation ranking algorithm to recommend recalled knowledge items. This intelligent recommendation ranking algorithm combines collaborative filtering and content-based recommendation algorithms, generating a personalized ranking of knowledge items by weighted fusion of the recommendations from these two algorithms. Specifically, the main goal of this algorithm is to intelligently analyze the user's interests and needs based on the user's role information, project information, operation records, etc., so as to sort the most relevant knowledge items and output the Top-K results that best meet the user's needs.

[0080] When sorting, the user's behavior and history must first be analyzed. This includes user role analysis, which identifies user roles such as construction supervisor, project manager, and administrator. Users with different roles have different interests in knowledge items. Project information analysis, which identifies the knowledge areas a user may be interested in based on the type of project they are working on, such as architecture, civil engineering, and infrastructure, and regional information, such as the main construction categories in the area. Operation record analysis, which analyzes user interest preferences based on their historical evaluations of content generated by the intelligent large model, as well as their historical queries, clicks, and browsing history. This analysis then models the user's interest data and identifies their long-term interests.

[0081] Based on multi-dimensional data such as user behavior analysis, role information, and project information, the collaborative filtering algorithm recommends knowledge items that the user may be interested in based on the user's behavior history (such as the similarity of behavior with other similar users). The content-based recommendation algorithm recommends knowledge items that match the user's known interests based on the content characteristics of the knowledge items (such as keywords, categories, topics, etc.) and the user's interest vector. Finally, the two are weighted and fused, and the user's query intent and actual behavior are comprehensively considered. Based on the weighted scores, the system will sort all recalled knowledge items to optimize the final recommendation ranking. Usually, the system will select the top K items and return them to the user. The value of K is set according to actual needs. Among them, the user's interest vector refers to the vectorized representation of the historical retrieval output results.

[0082] Figure 2 The following is a flowchart showing the specific implementation of the multimodal knowledge retrieval method for construction images provided by the present invention. Figure 2 , by providing a user-friendly image upload interface to obtain construction images; using the large model combined with the prompt project, the image content is initially analyzed at multiple levels, including extracting the construction subjects, scenes and potential problems in the image, and identifying the knowledge base that may be involved; on the other hand, the construction image is preprocessed to generate n related construction images, and then according to the identified knowledge domain, a mixed query is automatically performed on the relevant knowledge base, and the construction image and the construction subjects identified by the large model are vectorized using the MiniCPM-V multimodal encoding model, and the matching knowledge items are recalled through vector retrieval, and the construction subjects identified by the large model are used for keyword retrieval to recall the matching knowledge items; the intelligent recommendation sorting algorithm is used to deduplicate and sort the knowledge items recalled from the image and content, and the topK knowledge items that are more consistent are extracted to form a complete knowledge set; the extracted knowledge items are merged with the new prompts generated by the image content, and then input into the large model for processing again to output accurate answers.

[0083] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0084] Example 2

[0085] An embodiment of the present invention provides a multimodal knowledge retrieval system for construction images, the system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, the processor executing the computer program / instruction to implement the steps of the multimodal knowledge retrieval method for construction images. Figure 1 Steps S11-S16 are shown.

[0086] In the specific implementation process of the second embodiment, reference may be made to the first embodiment, and the corresponding technical effects are achieved.

[0087] Example 3

[0088] An embodiment of the present invention provides a computer program product having a computer program stored thereon. When the computer program is executed by a processor, the steps in the embodiment of the multimodal knowledge retrieval method for construction images are implemented, for example: Figure 1 Steps S11-S16 are shown.

[0089] In the specific implementation process of Example 3, reference may be made to Example 1, and the corresponding technical effects are achieved.

[0090] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of the present invention and to form different embodiments. For example, any of the claimed embodiments may be used in any combination.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multimodal knowledge retrieval method for construction images, characterized by: The method comprises: Obtain user question data, where the question data is a construction image, and perform data enhancement processing on the construction image to generate n related construction images, where n is greater than or equal to 2; The construction image is parsed using a preset intelligent large model to extract the parsed content of the construction image. The target knowledge base topics involved in the construction image are determined based on the parsed content. The parsed content includes construction subjects, knowledge fields, and application scenarios. Using a preset multimodal encoding model, each relevant construction image obtained after data augmentation is vectorized to generate a first semantic vector. Each first semantic vector is used as a first retrieval vector to perform similarity calculations with vector data stored in the knowledge base corresponding to the target knowledge base topic, so as to recall a first set of knowledge items that match the construction image. Taking the construction subject of the construction image as a query statement, obtaining knowledge items matching the construction subject from a knowledge base corresponding to the target knowledge base subject according to the query statement, so as to recall a second set of knowledge items matching the construction image; Intelligently sort and recommend knowledge items in the first knowledge item set and the second knowledge item set; The top K knowledge items are selected from the intelligent sorting recommendation results. A prompt project is constructed based on the parsed content of the construction image, the construction image, and the top K knowledge items. The intelligent big model is then used again in combination with the prompt project to perform task analysis on the construction image to obtain the final retrieval results.

2. The method according to claim 1, characterized in that The target knowledge base topics involved in the construction images are determined based on the parsed content of the construction images, including: Determining a first knowledge base topic involved in the construction image according to the knowledge domain and application scenario of the construction image; Semantically encoding the construction image to generate a corresponding semantic vector, matching the semantic vector with vector data in a preset knowledge base, and determining the subject of the knowledge base to which the vector data matching the semantic vector belongs as the second knowledge base subject involved in the construction image; Perform visual feature recognition on the construction image, search a predefined matching relationship table to obtain application scenarios and / or knowledge domains that match the visual features, and determine the third knowledge base topic related to the construction image based on the application scenarios and / or knowledge domains that match the visual features; the matching relationship table includes the correspondence between the visual features of the image and the application scenarios and / or knowledge domains; The first knowledge base topic, the second knowledge base topic and the third knowledge base topic are used as target knowledge base topics involved in the construction image.

3. The method according to claim 1, characterized in that The construction subject of the construction image is used as a query statement, and knowledge items matching the construction subject are obtained from the knowledge base corresponding to the target knowledge base subject according to the query statement, including: The construction subject of the construction image is used as a query statement, and the query statement is vectorized using a preset multimodal encoding model to generate a second semantic vector. The second semantic vector is used as a second retrieval vector and a similarity calculation is performed with the vector data stored in the knowledge base corresponding to the target knowledge base topic to recall a first sub-set of knowledge items that match the construction image; The construction subject of the construction image is used as a query statement, keywords in the query statement are extracted, and similarity calculation is performed between the keywords and keywords stored in the knowledge base corresponding to the target knowledge base topic to recall the second sub-knowledge item set matching the construction image.

4. The method according to claim 1, wherein Intelligently ranking and recommending knowledge items in the first knowledge item set and the second knowledge item set includes: Obtaining interest model data of the current user, and performing a first recommendation score on the knowledge items in the first knowledge item set and the second knowledge item set according to the interest model data; performing content matching with the historical search output results of the current user based on the content features of the knowledge items in the first knowledge item set and the second knowledge item set, and performing a second recommendation score based on the matching results; Calculate the final score of each knowledge item based on the preset scoring method weight ratio and the first recommended score and second recommended score corresponding to each knowledge item; Intelligently sort and recommend each knowledge item based on the final score.

5. The method according to claim 4, characterized in that Get the current user's interest model data, including: Obtain the current user's user role, the project information the user is responsible for, and the user's historical operation records. The historical operation records include the user's historical evaluation records of the content generated by the intelligent big model, as well as the user's historical query, click, and browsing records; Generate a co-occurrence matrix based on the user role, project information, and historical operation records, and find K users most similar to the current user based on the co-occurrence matrix; Based on the historical operation records of the K users on the content generated by the intelligent big model, the interest preferences of the K users in knowledge base topics in different knowledge fields and / or application scenarios are analyzed, and the current user's interest preferences in knowledge items of different knowledge base topics are predicted based on the interest preferences of the K users to obtain the current user's interest model data.

6. The method according to claim 1, characterized in that Performing data enhancement processing on the construction image to generate n related construction images includes: Positioning a foreground image region of the construction image, extracting the foreground image region from the construction image, and sequentially splicing the foreground image region into different preset background images to obtain a plurality of related construction images; and / or Randomly cropping different sub-regions of the construction image to obtain multiple related construction images; and / or Performing at least one geometric transformation of multi-angle rotation, horizontal flipping, vertical flipping, and affine transformation on the construction image to simulate construction images from different perspectives, thereby obtaining a plurality of related construction images; and / or Utilizing the OpenCV library to optimize and adjust the color, brightness, contrast, and / or saturation of the construction images to obtain multiple related construction images; and / or Performing different degrees of denoising and blurring on the construction images to obtain multiple related construction images; and / or Different key features in the construction image are highlighted using edge detection and texture enhancement techniques to obtain multiple related construction images.

7. The method according to claim 1, characterized in that Before obtaining the user question data, the method further includes: Constructing a knowledge base for knowledge retrieval, wherein the knowledge base is divided into knowledge base themes according to knowledge domains and application scenarios in the construction field, and the knowledge base includes a metadata database and a vector database; The constructing of a knowledge base for knowledge retrieval includes: Each text data in the preset construction field document dataset is divided into several knowledge base topics according to the knowledge domain and application scenarios of the construction field, and the text data corresponding to each knowledge base topic is determined; Each text data is segmented to form a knowledge item dataset with complete semantic information; Perform word segmentation processing on each knowledge item in the knowledge item data set corresponding to each text data to extract the keywords included in each knowledge item, assign a unique ID to each knowledge item in the knowledge item data set, establish a correspondence between each keyword of the knowledge item and the unique ID, and store the keywords included in each knowledge item and the correspondence between each keyword and the unique ID in a metadata database corresponding to the text data; Semantic encoding is performed on each knowledge item in the knowledge item data set corresponding to each text data to generate corresponding vector data, and the vector data is stored in the vector database corresponding to the text data.

8. The method according to claim 1, characterized in that After obtaining the user question data, the method further includes: A picture-specific link in the form of a link is generated for the construction image, and an effective use time is set for the picture-specific link, so that the intelligent large model can call the construction image through the picture-specific link within the effective use time.

9. A multimodal knowledge retrieval system for construction images, comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, characterized in that: The processor executes the computer program / instructions to implement the steps of the method according to any one of claims 1 to 8.

10. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Agricultural multi-mode intelligent retrieval technology and system based on multi-source heterogeneous data

    CN117573882A

  • Construction method and device of knowledge base question-answering system, equipment and storage medium

    CN119293164A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1