Construction image-oriented multi-modal knowledge retrieval method and system
By parsing and vectorizing construction images using a multimodal knowledge retrieval method and combining them with a user interest model for intelligent sorting, the problem of low matching degree of information mining and retrieval results in traditional methods is solved, and accurate knowledge content retrieval is achieved in complex scenarios.
Patent Information
- Application Number
- CN202510285637.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Traditional construction image analysis methods struggle to delve into multi-layered information, effectively integrate knowledge from multiple fields, and yield low matching results with user needs, making it difficult to adapt to changes in complex construction scenarios.
A multimodal knowledge retrieval method is adopted. Construction images are analyzed through an intelligent large model to generate multiple related images. Multimodal coding model is used for vectorization processing, and user interest model is combined for intelligent ranking and recommendation to dynamically adapt to the knowledge base retrieval scope.
It enables accurate retrieval of knowledge content from construction images, and can output the most relevant knowledge content in cross-domain application scenarios, thereby improving the intelligence level of retrieval and user experience.
Smart Images

Figure CN120448567B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a multi-modal knowledge retrieval method and system for construction images. BACKGROUND
[0002] With the expansion of the scale and gradual increase in complexity of modern engineering construction, the informatization management demand of the construction site is becoming increasingly important. As an important tool for recording engineering progress, conducting quality management and implementing safety supervision, construction images contain a large amount of site information and have become key data in construction management. However, with the rapid growth of the amount of picture data on the construction site, traditional image analysis and retrieval techniques have gradually exposed many shortcomings, and there are obvious limitations in analyzing complex scenes, correlating multi-domain knowledge and meeting user individual needs.
[0003] Traditional image analysis methods mostly rely on single-modal feature extraction techniques, and the analysis of picture content is usually limited to surface features, making it difficult to deeply mine the multi-level information contained in the pictures. For example, in a construction scene, a picture may contain multiple elements such as mechanical equipment, building components, personnel activities, etc., but existing techniques are difficult to comprehensively identify these elements and their associated characteristics. In addition, for the identification of potential problems in construction images (such as safety hazards or quality defects), traditional methods usually need to rely on human intervention, have low automation, and are difficult to meet actual needs. Especially in construction scenes involving multiple fields (such as the combination of civil construction and mechanical and electrical installation), existing techniques are even less capable in understanding and analyzing multi-domain characteristics. The knowledge base in the construction field usually covers multiple professional fields such as concrete construction, steel structure engineering, earthwork operation, etc., and the content is complex and not clearly classified. Existing knowledge retrieval systems mostly rely on keyword or tag matching, lack a deep understanding of user query intent, and result in a low matching degree between the retrieval results and user needs. Current retrieval methods are often difficult to dynamically adapt to changes in the construction scene, and cannot adjust the retrieval range of the knowledge base according to the specific content in the picture. In addition, for complex construction scenes involving multiple fields, traditional methods are difficult to effectively integrate knowledge items from different fields, resulting in limitations in the retrieval efficiency and accuracy of the system. SUMMARY
[0004] In view of the above problems, the present application is proposed to provide a multi-modal knowledge retrieval method and system for construction images to overcome the above problems.
[0005] In one aspect of the present application, a multi-modal knowledge retrieval method for construction images is provided, the method comprising:
[0006] obtaining user query data, the query data being a construction image, performing data enhancement processing on the construction image to generate n relevant construction images, n≥2;
[0007] The construction images are analyzed using a pre-set intelligent big model to extract the analysis content. Based on the analysis content, the target knowledge base topics involved in the construction images are determined. The analysis content includes the construction subject, knowledge domain, and application scenario.
[0008] The first semantic vector is generated by vectorizing each relevant construction image obtained after data augmentation using a preset multimodal coding model. Each first semantic vector is then used as a first retrieval vector and its similarity is calculated with the vector data stored in the target knowledge base topic-corresponding knowledge base to recall the first set of knowledge entries that match the construction image.
[0009] Using the main body of the construction image as the query statement, knowledge entries matching the main body are retrieved from the knowledge base corresponding to the target knowledge base topic, in order to recall a second set of knowledge entries matching the construction image.
[0010] Intelligent sorting and recommendation of knowledge items in the first and second knowledge item sets;
[0011] The top K knowledge items in the intelligent ranking recommendation results are selected. Based on the parsed content of the construction images, the construction images, and the top K knowledge items, a prompt project is constructed. The intelligent big model is then used again in conjunction with the prompt project to perform task parsing on the construction images to obtain the final search results.
[0012] Furthermore, based on the parsed content of the construction images, the target knowledge base topics related to the construction images are determined, including:
[0013] The primary knowledge base topics involved in the construction images are determined based on the knowledge domain and application scenario of the construction images;
[0014] The construction image is semantically encoded to generate a corresponding semantic vector, and the semantic vector is matched with the vector data in the preset knowledge base. The topic of the knowledge base to which the vector data that matches the semantic vector belongs is determined as the second knowledge base topic involved in the construction image.
[0015] Visual feature recognition is performed on construction images, and a predefined matching relationship table is searched to obtain the application scenarios and / or knowledge domains that match the visual features. The third knowledge base topic involved in the construction images is determined based on the application scenarios and / or knowledge domains that match the visual features. The matching relationship table includes the correspondence between the visual features of the images and the application scenarios and / or knowledge domains.
[0016] The first knowledge base topic, the second knowledge base topic, and the third knowledge base topic are used as the target knowledge base topics involved in the construction images.
[0017] Furthermore, the main body of the construction image is used as the query statement. Based on the query statement, knowledge entries matching the main body are retrieved from the knowledge base corresponding to the target knowledge base topic, including:
[0018] The construction subject of the construction image is used as the query statement. The query statement is vectorized using a preset multimodal coding model to generate a second semantic vector. The second semantic vector is used as the second retrieval vector and similarity is calculated with the vector data stored in the target knowledge base topic corresponding to the knowledge base to recall the first sub-knowledge item set that matches the construction image.
[0019] The main body of the construction image is used as the query statement. Keywords in the query statement are extracted, and the similarity between the keywords and the keywords stored in the target knowledge base topic is calculated to recall the second sub-knowledge item set that matches the construction image.
[0020] Furthermore, intelligent ranking and recommendation are performed on the knowledge items in the first and second knowledge item sets, including:
[0021] Obtain the current user's interest model data, and perform a first recommendation score on the knowledge items in the first knowledge item set and the second knowledge item set based on the interest model data;
[0022] Content matching is performed based on the content characteristics of knowledge items in the first and second knowledge item sets and the current user's historical search results. A second recommendation score is then given based on the matching results.
[0023] The final score for each knowledge item is calculated based on the preset scoring method weighting ratio and the first and second recommended scores corresponding to each knowledge item.
[0024] The knowledge items are intelligently sorted and recommended based on the final score.
[0025] Furthermore, obtain the current user's interest model data, including:
[0026] Obtain the current user's user role, the project information the user is responsible for, and the user's historical operation records. The historical operation records include the user's historical evaluation records of the content generated by the intelligent big model, as well as the user's historical query, click, and browsing records.
[0027] A co-occurrence matrix is generated based on the user role, project information, and historical operation records, and the K users most similar to the current user are found based on the co-occurrence matrix;
[0028] Based on the historical operation records of the K users on the content generated by the intelligent big model, analyze the interest bias of the K users on knowledge base topics in different knowledge domains and / or application scenarios, and predict the current user's interest bias on knowledge items of different knowledge base topics based on the interest bias of the K users, so as to obtain the current user's interest model data.
[0029] Further, performing data augmentation processing on the construction images to generate n related construction images includes:
[0030] The foreground image region of the construction image is located, extracted from the construction image, and then sequentially stitched together with different preset background images to obtain multiple related construction images; and / or
[0031] Randomly cropping different sub-regions of the construction image yields multiple related construction images; and / or
[0032] Perform at least one of the following geometric transformations on the construction images: multi-angle rotation, horizontal flipping, vertical flipping, and affine transformation, to simulate construction images from different perspectives, thereby obtaining multiple related construction images; and / or
[0033] The OpenCV library was used to optimize and adjust the color, brightness, contrast, and / or saturation of construction images, resulting in multiple related construction images; and / or
[0034] Different degrees of denoising and blurring were applied to the construction images to obtain multiple related construction images; and / or
[0035] By using edge detection and texture enhancement techniques, different key features in the construction images are highlighted to obtain multiple related construction images.
[0036] Furthermore, before acquiring user question data, the method further includes:
[0037] A knowledge base for knowledge retrieval is constructed. The knowledge base is divided into knowledge base topics according to the knowledge domains and application scenarios in the construction field. The knowledge base includes a metadata database and a vector database.
[0038] The knowledge base constructed for knowledge retrieval includes:
[0039] Each text data in the pre-set construction domain document dataset is divided into several knowledge base topics according to the knowledge domain and application scenario of the construction domain, and the text data corresponding to each knowledge base topic is determined.
[0040] Each text data point is segmented to form a knowledge entry dataset with complete semantic information.
[0041] Each knowledge entry in the knowledge entry dataset corresponding to each text data is segmented to extract the keywords included in each knowledge entry, and a unique ID is assigned to each knowledge entry in the knowledge entry dataset. The correspondence between each keyword of the knowledge entry and the unique ID is established, and the keywords included in each knowledge entry and the correspondence between each keyword and the unique ID are stored in the metadata database of the corresponding text data.
[0042] Semantic encoding is performed on each knowledge item in the knowledge item dataset corresponding to each text data to generate corresponding vector data, which is then stored in the vector database of the corresponding text data.
[0043] Furthermore, after obtaining user question data, the method further includes:
[0044] Generate dedicated image links in the form of links for the construction images, and set a valid usage time for the dedicated image links so that the intelligent large model can call the construction images through the dedicated image links within the valid usage time.
[0045] In another aspect, the present invention provides a multimodal knowledge retrieval system for construction images, the system comprising a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor, the processor executing the computer program / instructions to implement the steps of the above-described multimodal knowledge retrieval method for construction images.
[0046] In another aspect, the present invention also provides a computer program product storing a computer program that, when executed by a processor, implements the steps of the above-described multimodal knowledge retrieval method for construction images.
[0047] The multimodal knowledge retrieval method and system for construction images provided in this invention support dynamic adaptation to multiple knowledge bases for knowledge content retrieval. It can flexibly select the retrieval scope based on the identified knowledge base topics. For cross-domain application scenarios, this invention can simultaneously retrieve multiple knowledge bases. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content even in complex scenarios. Furthermore, this invention deeply integrates the image parsing content with the retrieval results of the initial search, and then inputs it again into an intelligent large-scale model to perform task retrieval. This guides the model's generation capabilities to the precise domain of the user's needs, ultimately outputting accurate and context-relevant answers.
[0048] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0049] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings:
[0050] Figure 1 This is a flowchart of a multimodal knowledge retrieval method for construction images according to an embodiment of the present invention;
[0051] Figure 2 This is a flowchart illustrating the specific implementation of a multimodal knowledge retrieval method for construction images according to an embodiment of the present invention. Detailed Implementation
[0052] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0053] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.
[0054] This invention provides a multimodal knowledge retrieval method for construction images, overcoming the limitations of existing technologies. Traditional image retrieval techniques struggle to comprehensively and accurately extract key information from construction images, especially when multiple professional fields are involved. Furthermore, this invention effectively improves the intelligence level of knowledge retrieval for construction images, moving beyond keyword or tag-based matching to understand the deeper meaning behind the images by incorporating user behavior, thereby providing more accurate answers and services.
[0055] Example 1
[0056] This invention provides a multimodal knowledge retrieval method for construction images, such as... Figure 1 As shown, the multimodal knowledge retrieval method for construction images proposed in this invention includes the following steps:
[0057] S11. Obtain user question data, wherein the question data is construction images, and perform data augmentation processing on the construction images to generate n related construction images, where n≥2.
[0058] This invention provides a multimodal knowledge retrieval method for construction images, where the user's query data is the construction image uploaded by the user. This invention provides a user-friendly image upload interface to obtain construction images. Specifically, the image upload interface supports the upload of construction images in common image formats and ensures the correct format and reasonable size of uploaded files through a pre-configured file verification mechanism. Furthermore, after obtaining the user's query data, i.e., after obtaining the uploaded construction image, a dedicated image link can be generated for the construction image, and a valid usage time can be set for the dedicated image link, allowing the intelligent large-scale model to access the construction image through the dedicated image link within the valid usage time. In one specific embodiment, after the image is uploaded, the system can generate a dedicated image link with a validity period of one hour for the large-scale model to access and process. After the link expires, it will become invalid, thereby ensuring that the image is not exposed to the public network for a long time and improving the security of image retrieval.
[0059] Furthermore, in order to improve the recall accuracy of the multimodal knowledge base, this invention performs a series of preprocessing steps on the construction images before inputting them into the multimodal coding model to generate multiple related or similar construction images, thereby achieving data augmentation processing of the construction images. Specifically, the data augmentation processing of the construction image in this step to generate n related construction images includes: locating the foreground image region of the construction image, extracting the foreground image region from the construction image, and sequentially stitching the foreground image region into different preset background images to obtain multiple related construction images; and / or randomly cropping different sub-regions of the construction image to obtain multiple related construction images; and / or performing at least one geometric transformation on the construction image, including multi-angle rotation, horizontal flip, vertical flip, and affine transformation, to simulate construction images from different perspectives to obtain multiple related construction images; and / or using the OpenCV library to optimize and adjust the color, brightness, contrast, and / or saturation of the construction image to obtain multiple related construction images; and / or performing different degrees of denoising and blurring processing on the construction image to obtain multiple related construction images; and / or using edge detection and texture enhancement techniques to highlight different key features in the construction image to obtain multiple related construction images.
[0060] The data augmentation process in this embodiment includes color correction of the original image, adjusting brightness, contrast, and saturation to reduce the impact of ambient light on the image. In one specific embodiment, the invention may also consider converting the original image to a grayscale image to eliminate potential interference from color information on knowledge recall. Regarding geometric transformations, the invention may also perform multi-angle rotations (e.g., ±15°, ±30°, ±45°) and horizontal and vertical flips on the original image to simulate different shooting angles and directions. Affine transformations are also applied to slightly adjust the image shape, enhancing the model's adaptability to different viewpoints. To address noise issues, the invention may employ denoising algorithms to reduce random noise in the image and use slight blurring to reduce the impact of detail differences. Edge detection and texture enhancement techniques can also be used to highlight key features in the image, improving the recall effect of the knowledge base. In one specific embodiment, in addition to the above data augmentation methods, the invention may also increase sample diversity and improve the generalization ability of the knowledge base by randomly cropping sub-regions of the image and / or mixing the foreground image region with different backgrounds.
[0061] S12. The construction images are analyzed using a pre-set intelligent large-scale model to extract their analytical content. Based on this content, the target knowledge base topics related to the construction images are determined. The analytical content includes the construction subject, knowledge domain, and application scenario. The intelligent large-scale model is fine-tuned and trained based on the open-source model llama3.2-vision:11b.
[0062] In this embodiment, the intelligent big data model is first used to extract knowledge domain and application scenario information from the images by combining the prompt project with user-uploaded construction images. This information is then used to determine the relevant knowledge base topics, laying the foundation for subsequent knowledge retrieval and prompt reinforcement, thereby achieving more refined semantic extraction and question answering. Specifically, the pre-set intelligent big data model is used to perform multi-level analysis of the image content, including extracting the construction subject, knowledge domain, application scenario, and potential problems from the image. Simultaneously, based on the analysis results, the system identifies potentially relevant knowledge base topics and corresponding knowledge bases, allowing the system to automatically query the relevant knowledge bases based on the identified topics.
[0063] In this embodiment of the invention, the knowledge base is divided into knowledge base topics according to the knowledge domains and application scenarios of the construction field. The knowledge base includes a metadata database and a vector database. Specifically, when constructing the knowledge base, it is divided into several knowledge base topics based on the knowledge domains and application scenarios of the construction field. For example, knowledge base topics for knowledge domains include concrete engineering, earthwork engineering, and steel structure engineering. Knowledge base topics for application scenarios include safety management and construction specifications. Each knowledge base topic covers a specific knowledge domain and application scenario, ensuring the clarity and completeness of knowledge classification. This classification method facilitates rapid matching of relevant content in subsequent searches. In a specific embodiment, knowledge domains can be divided into: foundation engineering, masonry structure engineering, waterproofing engineering, HVAC engineering, electrical engineering, water supply and drainage engineering, curtain wall engineering, fire protection engineering, concrete construction, steel structure engineering, and earthwork engineering; application scenarios can be divided into: safety management, construction specifications, quality management, and schedule management.
[0064] Further, the target knowledge base topics involved in the construction images are determined based on the parsed content of the construction images, including: determining the first knowledge base topic involved in the construction images based on the knowledge domain and application scenario of the construction images in the parsing results; semantically encoding the construction images to generate corresponding semantic vectors, and matching the semantic vectors with vector data in a preset knowledge base, determining the topic of the knowledge base to which the vector data matching the semantic vectors belongs as the second knowledge base topic involved in the construction images; performing visual feature recognition on the construction images, searching a predefined matching relationship table to obtain the application scenario and / or knowledge domain matching the visual features, and determining the third knowledge base topic involved in the construction images based on the application scenario and / or knowledge domain matching the visual features; the matching relationship table includes the correspondence between the visual features of the images and the application scenario and / or knowledge domain; and using the first knowledge base topic, the second knowledge base topic, and the third knowledge base topic as the target knowledge base topics involved in the construction images. Specifically, after the user uploads an image, the system combines the image with a pre-set specific prompt word template (Prompt) input, and uses a multimodal large model to perform preliminary parsing of the image content. The Prompt project plays a crucial role here, explicitly defining the information the model needs to extract from images by designing task prompts tailored to the construction domain. This includes information such as the main construction elements (e.g., machinery and building components), application scenarios (e.g., construction environment and spatial layout), and knowledge domains. Simultaneously, the model can automatically associate relevant knowledge base topics by parsing the semantic information within the image content. Furthermore, the model can match elements and features in the image to predefined application scenarios and / or knowledge domains; for example, concrete construction corresponds to a concrete knowledge base, and equipment operation corresponds to a machinery and equipment knowledge base.
[0065] S13. Using a preset multimodal coding model, each relevant construction image obtained after data augmentation is vectorized to generate a first semantic vector. Each first semantic vector is used as a first retrieval vector and its similarity is calculated with the vector data stored in the target knowledge base topic-corresponding knowledge base to recall the first set of knowledge entries that match the construction image.
[0066] S14. Using the main body of the construction image as a query statement, retrieve knowledge entries matching the main body from the knowledge base corresponding to the target knowledge base topic, in order to recall the second set of knowledge entries matching the construction image.
[0067] S15. Perform intelligent sorting and recommendation of knowledge items in the first knowledge item set and the second knowledge item set.
[0068] S16. Select the top K knowledge items from the intelligent ranking recommendation results. Construct a prompt project based on the parsed content of the construction image, the construction image, and the top K knowledge items. Then, use the intelligent big model in conjunction with the prompt project to perform task parsing on the construction image to obtain the final search results.
[0069] In this embodiment, the image parsing content is deeply integrated with refined knowledge entries, and then input again into a multimodal large model to generate accurate and context-relevant answers. Specifically, during the fusion phase, the system selects the top K knowledge entries from the image parsing content and intelligent ranking recommendation results to construct a new prompt template, explicitly specifying the problem or task that the model needs to solve. The new prompt template integrates the construction subject, application scenario, knowledge domain extracted from the image, and refined knowledge entries, guiding the model's generation capabilities to the precise domain of the user's needs, ultimately outputting an accurate answer.
[0070] The multimodal knowledge retrieval method for construction images provided in this invention supports dynamic adaptation to multiple knowledge bases for knowledge content retrieval. It can flexibly select the retrieval scope based on the identified knowledge base topics. For cross-domain application scenarios, this invention can simultaneously retrieve multiple knowledge bases. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content even in complex scenarios. Furthermore, this invention deeply integrates the image parsing content with the retrieval results of the initial search, and then inputs it again into an intelligent large-scale model to perform task retrieval. This guides the model's generation capabilities to the precise domain of the user's needs, ultimately outputting accurate and context-relevant answers.
[0071] In this embodiment of the invention, before acquiring user query data, a knowledge base for knowledge retrieval needs to be constructed. The knowledge base is divided into knowledge base topics according to the knowledge domains and application scenarios of the construction field. The knowledge base includes a metadata database and a vector database. Specifically, constructing the knowledge base for knowledge retrieval includes: dividing each text data in a preset construction field document dataset into several knowledge base topics according to the knowledge domains and application scenarios of the construction field, and determining the text data corresponding to each knowledge base topic; performing text segmentation on each text data to form a semantically complete knowledge entry dataset; performing word segmentation on each knowledge entry in the knowledge entry dataset corresponding to each text data to extract the keywords included in each knowledge entry, assigning a unique ID to each knowledge entry in the knowledge entry dataset, establishing a correspondence between each keyword of the knowledge entry and the unique ID, and storing the keywords included in each knowledge entry and the correspondence between each keyword and the unique ID in the metadata database of the corresponding text data; performing semantic encoding on each knowledge entry in the knowledge entry dataset corresponding to each text data to generate corresponding vector data, and storing it in the vector database of the corresponding text data.
[0072] In this embodiment, the "knowledge base" has two storage formats: text and vector. Each knowledge base is first split into knowledge entries by document segmentation, and each knowledge entry is assigned a unique ID. Keywords in the knowledge entries are obtained through jieba word segmentation, and all keywords point to this unique ID, facilitating full-text indexing. Then, the open-source MiniCPM-V model is used to semantically encode the knowledge entries to generate vector data, which is stored in a PostgreSQL database. Simultaneously, the text data corresponding to the knowledge entries is stored in metadata, facilitating the retrieval of more accurate knowledge entries when using hybrid retrieval. Here, "document" refers to professional literature, standards, guidelines, etc., related to the construction field. Segmentation involves parsing the document content and dividing it into independent knowledge units (i.e., knowledge entries). In a specific example, the segmentation method can be: extracting text from the document, pre-segmenting it according to a length of 600 characters, and searching for any one of the following characters within a 50-character range (period, semicolon, newline, tab, etc.) and segmenting at that point; if not found, segmenting at the 600th character.
[0073] In this embodiment of the invention, the recall of knowledge entries matching the construction images is achieved by vectorizing multiple related construction images after data augmentation using a multimodal coding model, and by automatically querying the relevant knowledge base based on the target knowledge base topics extracted in the preliminary identification stage to recall matching knowledge entries.
[0074] Specifically, after data augmentation, a batch of relevant construction images are first used to generate high-dimensional semantic vector representations using the MiniCPM-V multimodal coding model. The multimodal coding model integrates the visual features of the images with their underlying semantic information, transforming the complex content of the images into quantifiable vector representations. Subsequently, the generated vectors are used as the first retrieval vectors and their similarity is calculated with predefined vectors in a selected knowledge base (i.e., a knowledge base corresponding to the target topic). Each knowledge entry in the knowledge base has been vectorized during construction and stored in a vector database. Through the efficient KNN nearest neighbor algorithm, the system can quickly recall the knowledge entries most relevant to the image content. For example, if the image involves "template erection," the system may recall entries related to template fixing methods, material specifications, and construction steps.
[0075] Further, step S14, using the construction subject of the construction image as a query statement and retrieving knowledge entries matching the construction subject from the knowledge base corresponding to the target knowledge base topic, specifically includes: using the construction subject of the construction image as a query statement, vectorizing the query statement using a preset multimodal coding model to generate a second semantic vector, using the second semantic vector as a second retrieval vector and calculating the similarity between it and the vector data stored in the target knowledge base topic-corresponding knowledge base to recall a first sub-knowledge entry set matching the construction image; using the construction subject of the construction image as a query statement, extracting keywords from the query statement, and calculating the similarity between the keywords and the keywords stored in the target knowledge base topic-corresponding knowledge base to recall a second sub-knowledge entry set matching the construction image. The first sub-knowledge entry set and the second sub-knowledge entry set are combined to form a second knowledge entry set.
[0076] In this embodiment, in addition to direct retrieval based on image vectorization, the present invention also uses the construction subject (such as "concrete wall" or "steel reinforcement material") identified by the intelligent large model as a query statement for mixed retrieval in a selected knowledge base (i.e., a knowledge base corresponding to the topic of the target knowledge base). Mixed retrieval includes: First, inputting these query statements into the MiniCPM-V model for vectorization to obtain query vectors; using the KNN nearest neighbor algorithm to calculate the vectors of the K closest knowledge items; and obtaining the knowledge items with the highest relevance. The value of K can be set according to user needs, and can be selected as 5, 10, 20, 30, or other values; the present invention does not specifically limit this. Second, segmenting the query statements using jieba word segmentation to extract keywords from the query; then using the BM25 algorithm to calculate and sort the relevance between the keywords and the knowledge item keywords; and obtaining the knowledge items with the highest relevance.
[0077] The knowledge content retrieval function of this invention supports dynamic adaptation to multiple knowledge bases, enabling flexible selection of the search scope based on the identified knowledge domain and application scenario. For example, for concrete application scenarios, the system prioritizes searching the concrete construction specifications database; for earthwork operations, it queries the earthwork engineering safety management database. In this way, the system avoids interference from irrelevant content, improving search efficiency and the accuracy of retrieval results. Furthermore, for cross-domain application scenarios (such as those involving both machinery and building components), the system can simultaneously search multiple knowledge bases and assign weights based on the importance of the retrieval results. This dynamic adaptation mechanism ensures that the system can still output the most relevant knowledge content even in complex scenarios. The weight assignment refers to assigning different weights to knowledge items retrieved based on construction entity information, construction entity information vectors, and preprocessed image vectors to obtain the final knowledge items.
[0078] In this embodiment of the invention, step S15, which involves intelligently ranking and recommending knowledge items in the first and second knowledge item sets, includes: obtaining the current user's interest model data; assigning a first recommendation score to the knowledge items in the first and second knowledge item sets based on the interest model data; performing content matching between the content features of the knowledge items in the first and second knowledge item sets and the current user's historical search output results; assigning a second recommendation score based on the matching results; calculating the final score for each knowledge item based on a preset weight ratio of the scoring method and the first and second recommendation scores corresponding to each knowledge item, wherein the weight ratio of the first and second recommendation scores can be selected as 4:6; and performing intelligent ranking and recommendation on each knowledge item based on the final score. The process of obtaining the current user's interest model data specifically includes: obtaining the current user's user role, the project information the user is responsible for, and the user's historical operation records, including the user's historical evaluation records of the content generated by the intelligent big data model, as well as the user's historical query, click, and browsing records; generating a co-occurrence matrix based on the user role, project information, and historical operation records, and finding the K users most similar to the current user based on the co-occurrence matrix; analyzing the K users' historical operation records of the content generated by the intelligent big data model, analyzing the K users' interest biases towards knowledge base topics in different knowledge domains and / or application scenarios, and predicting the current user's interest bias towards knowledge items in different knowledge base topics based on the K users' interest biases, thus obtaining the current user's interest model data.
[0079] In actual construction management, different user roles have significantly different needs for knowledge content. For example, project managers are more concerned with schedule and cost, while construction supervisors focus more on safety. Existing systems generally lack the ability to identify user roles and provide personalized support, making it difficult to provide accurate recommendations based on users' historical operation records, current scenario requirements, or project type. Furthermore, most existing ranking algorithms are relatively simple and cannot dynamically adjust the priority of recommendation results based on user preferences, resulting in a poor user experience. Moreover, as construction management moves towards intelligence and automation, the shortcomings of existing traditional systems in intelligent retrieval, recommendation, and dynamic adaptation are becoming increasingly prominent. Existing technologies lack intelligent ranking mechanisms that consider user scenario needs when recommending knowledge items, leading to deviations between recommended results and actual user requirements. For complex application scenarios, the system cannot dynamically adapt to the retrieval scope of cross-domain knowledge bases, making it difficult to provide users with comprehensive and accurate answers. This limitation significantly reduces the efficiency and user satisfaction of construction management systems. To address these issues, this invention employs an intelligent recommendation ranking algorithm to recommend recalled knowledge items. This intelligent recommendation ranking algorithm combines collaborative filtering and content-based recommendation algorithms, generating a personalized ranking of knowledge items by weighted fusion of the recommendation results from these two algorithms. Specifically, the main goal of this algorithm is to intelligently analyze the user's interests and needs based on the user's role information, project information, operation records, etc., and then sort the most relevant knowledge items to output the Top-K results that best meet the user's needs.
[0080] When sorting, the first step is to analyze user behavior and historical records. This includes user role analysis, identifying user roles such as construction supervisors, project managers, and administrators, as different roles have different areas of interest in knowledge items; project information analysis, determining the knowledge areas that users might be interested in based on the type of project they are involved in, such as architecture, civil engineering, and infrastructure, and regional information, such as the main construction categories in the region; and operation record analysis, where the system analyzes users' historical evaluation records of content generated by the intelligent big data model, as well as their historical query, click, and browsing records, to analyze user interest biases, build user interest model data, and obtain users' long-term interests.
[0081] Based on multi-dimensional data such as user behavior analysis, role information, and project information, collaborative filtering algorithms recommend knowledge items that users may be interested in based on their behavioral history (e.g., similarity to behavior with other similar users). Content-based recommendation algorithms recommend knowledge items that match the user's known interests based on the content features of the knowledge items (e.g., keywords, categories, topics) and the user's interest vector. Finally, the two are weighted and fused, comprehensively considering the user's query intent and actual behavior. Based on the weighted score, the system ranks all recalled knowledge items to optimize the final recommendation ranking. Typically, the system selects the top K items to return to the user, where the value of K is set according to actual needs. Here, the user's interest vector refers to the vectorized representation of the historical search results.
[0082] Figure 2 A flowchart illustrating the specific implementation of the multimodal knowledge retrieval method for construction images provided by this invention is shown. See also... Figure 2 The system provides a user-friendly image upload interface to obtain construction images. Using a large model combined with a prompt project, it performs multi-level analysis of the image content, including extracting the construction subject, scene, and potential problems, while also identifying relevant knowledge bases. On another front, the construction images are preprocessed to generate n related construction images. Then, based on the identified knowledge domains, a mixed query is automatically performed on the relevant knowledge bases. The MiniCPM-V multimodal coding model is used to vectorize the construction images and the construction subject identified by the large model. Vector retrieval recalls matching knowledge entries, and keyword retrieval using the construction subject identified by the large model recalls matching knowledge entries. An intelligent recommendation and ranking algorithm is used to deduplicate and rank the knowledge entries retrieved from the images and content, extracting the top K most relevant knowledge entries to form a complete knowledge set. The extracted knowledge entries are then fused with new prompts generated from the image content and input back into the large model for further processing, outputting accurate solutions.
[0083] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0084] Example 2
[0085] This invention provides a multimodal knowledge retrieval system for construction images. The system includes a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor. The processor executes the computer program / instructions to implement the steps of the aforementioned multimodal knowledge retrieval method for construction images. For example... Figure 1 Steps S11-S16 are shown.
[0086] In the specific implementation process of Embodiment 2, you can refer to Embodiment 1, and it has the corresponding technical effects.
[0087] Example 3
[0088] This invention provides a computer program product storing a computer program. When executed by a processor, the computer program implements the steps described in the above-described embodiment of the multimodal knowledge retrieval method for construction images, for example... Figure 1 Steps S11-S16 are shown.
[0089] In the specific implementation process of Example 3, reference can be made to Example 1, and it has the corresponding technical effects.
[0090] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, any of the claimed embodiments can be used in any combination.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal knowledge retrieval method for construction images, characterized in that, The method includes: Obtain user question data, wherein the question data is construction images, and perform data augmentation processing on the construction images to generate n related construction images, where n≥2; The construction images are analyzed using a pre-set intelligent big model to extract the analysis content. Based on the analysis content, the target knowledge base topics involved in the construction images are determined. The analysis content includes the construction subject, knowledge domain, and application scenario. The first semantic vector is generated by vectorizing each relevant construction image obtained after data augmentation using a preset multimodal coding model. Each first semantic vector is then used as a first retrieval vector and its similarity is calculated with the vector data stored in the target knowledge base topic-corresponding knowledge base to recall the first set of knowledge entries that match the construction image. Using the main body of the construction image as the query statement, knowledge entries matching the main body are retrieved from the knowledge base corresponding to the target knowledge base topic, in order to recall a second set of knowledge entries matching the construction image. Intelligent sorting and recommendation of knowledge items in the first and second knowledge item sets; The top K knowledge items in the intelligent ranking recommendation results are selected. Based on the parsed content of the construction images, the construction images, and the top K knowledge items, a prompt project is constructed. The intelligent big model is then used again in conjunction with the prompt project to perform task parsing on the construction images to obtain the final search results.
2. The method according to claim 1, characterized in that, Based on the parsed content of the construction images, the target knowledge base topics related to the construction images are determined, including: The primary knowledge base topics involved in the construction images are determined based on the knowledge domain and application scenario of the construction images; The construction image is semantically encoded to generate a corresponding semantic vector, and the semantic vector is matched with the vector data in the preset knowledge base. The topic of the knowledge base to which the vector data that matches the semantic vector belongs is determined as the second knowledge base topic involved in the construction image. Visual feature recognition is performed on construction images, and a predefined matching relationship table is searched to obtain the application scenarios and / or knowledge domains that match the visual features. The third knowledge base topic involved in the construction images is determined based on the application scenarios and / or knowledge domains that match the visual features. The matching relationship table includes the correspondence between the visual features of the images and the application scenarios and / or knowledge domains. The first knowledge base topic, the second knowledge base topic, and the third knowledge base topic are used as the target knowledge base topics involved in the construction images.
3. The method according to claim 1, characterized in that, Using the main construction element of the construction image as the query statement, knowledge entries matching the main construction element are retrieved from the knowledge base corresponding to the target knowledge base topic, including: The construction subject of the construction image is used as the query statement. The query statement is vectorized using a preset multimodal coding model to generate a second semantic vector. The second semantic vector is used as the second retrieval vector and similarity is calculated with the vector data stored in the target knowledge base topic corresponding to the knowledge base to recall the first sub-knowledge item set that matches the construction image. The main body of the construction image is used as the query statement. Keywords in the query statement are extracted, and the similarity between the keywords and the keywords stored in the target knowledge base topic is calculated to recall the second sub-knowledge item set that matches the construction image.
4. The method according to claim 1, characterized in that, Intelligent ranking and recommendation of knowledge items in the first and second knowledge item sets, including: Obtain the current user's interest model data, and perform a first recommendation score on the knowledge items in the first knowledge item set and the second knowledge item set based on the interest model data; Content matching is performed based on the content characteristics of knowledge items in the first and second knowledge item sets and the current user's historical search results. A second recommendation score is then given based on the matching results. The final score for each knowledge item is calculated based on the preset scoring method weighting ratio and the first and second recommended scores corresponding to each knowledge item. The knowledge items are intelligently sorted and recommended based on the final score.
5. The method according to claim 4, characterized in that, Obtain the current user's interest model data, including: Obtain the current user's user role, the project information the user is responsible for, and the user's historical operation records. The historical operation records include the user's historical evaluation records of the content generated by the intelligent big model, as well as the user's historical query, click, and browsing records. A co-occurrence matrix is generated based on the user role, project information, and historical operation records, and the K users most similar to the current user are found based on the co-occurrence matrix; Based on the historical operation records of the K users on the content generated by the intelligent big model, analyze the interest bias of the K users on knowledge base topics in different knowledge domains and / or application scenarios, and predict the current user's interest bias on knowledge items of different knowledge base topics based on the interest bias of the K users, so as to obtain the current user's interest model data.
6. The method according to claim 1, characterized in that, Data augmentation processing of the construction images to generate n related construction images includes: The foreground image region of the construction image is located, extracted from the construction image, and then sequentially stitched together with different preset background images to obtain multiple related construction images; and / or Randomly cropping different sub-regions of the construction image yields multiple related construction images; and / or Perform at least one of the following geometric transformations on the construction images: multi-angle rotation, horizontal flipping, vertical flipping, and affine transformation, to simulate construction images from different perspectives, thereby obtaining multiple related construction images; and / or The OpenCV library was used to optimize and adjust the color, brightness, contrast, and / or saturation of construction images, resulting in multiple related construction images; and / or Different degrees of denoising and blurring were applied to the construction images to obtain multiple related construction images; and / or By using edge detection and texture enhancement techniques, different key features in the construction images are highlighted to obtain multiple related construction images.
7. The method according to claim 1, characterized in that, Before acquiring user question data, the method further includes: A knowledge base for knowledge retrieval is constructed. The knowledge base is divided into knowledge base topics according to the knowledge domains and application scenarios in the construction field. The knowledge base includes a metadata database and a vector database. The knowledge base constructed for knowledge retrieval includes: Each text data in the pre-set construction domain document dataset is divided into several knowledge base topics according to the knowledge domain and application scenario of the construction domain, and the text data corresponding to each knowledge base topic is determined. Each text data point is segmented to form a knowledge entry dataset with complete semantic information. Each knowledge entry in the knowledge entry dataset corresponding to each text data is segmented to extract the keywords included in each knowledge entry, and a unique ID is assigned to each knowledge entry in the knowledge entry dataset. The correspondence between each keyword of the knowledge entry and the unique ID is established, and the keywords included in each knowledge entry and the correspondence between each keyword and the unique ID are stored in the metadata database of the corresponding text data. Semantic encoding is performed on each knowledge item in the knowledge item dataset corresponding to each text data to generate corresponding vector data, which is then stored in the vector database of the corresponding text data.
8. The method according to claim 1, characterized in that, After obtaining user question data, the method further includes: Generate dedicated image links in the form of links for the construction images, and set a valid usage time for the dedicated image links so that the intelligent large model can call the construction images through the dedicated image links within the valid usage time.
9. A multimodal knowledge retrieval system for construction images, the system comprising a memory, a processor, and computer programs / instructions stored in the memory and executable on the processor, characterized in that, The processor executes the computer program / instructions to implement the steps of the method according to any one of claims 1-8.
10. A computer program product, characterized in that, The computer program product stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Agricultural multi-mode intelligent retrieval technology and system based on multi-source heterogeneous data
CN117573882A
Construction method and device of knowledge base question-answering system, equipment and storage medium
CN119293164A