Construction safety risk early warning method and system based on cross-modal visual language retrieval
By adopting a multi-grained semantic decoupling combined image retrieval method based on CLIP in construction safety risk warning technology, the problems of insufficient environmental complexity, information islandization and intelligence in the existing technology are solved, and high-precision cross-modal risk warning is achieved.
Patent Information
- Application Number
- CN202510028717.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-13
AI Technical Summary
The existing construction safety risk warning technology has limitations in environmental complexity, information islandization and insufficient intelligence, and it is difficult to achieve cross-modal semantic understanding and high-precision risk warning.
A multi-grained semantic decoupling combined image retrieval method based on the pre-trained visual language model CLIP is adopted to achieve cross-modal semantic alignment and combination of images and text through multi-grained image-text semantic feature decoupling, feature combination and multi-grained target alignment.
It improves the accuracy and efficiency of construction safety risk warning, can effectively match risk images in complex environments, and enhances the model's understanding of the construction safety field.
Smart Images

Figure CN120146549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a construction safety risk early warning method and system based on cross-modal visual language retrieval, and in particular to a combined image retrieval method based on fine-grained alignment of graph visual language regions, belonging to the technical fields of combined image retrieval and multi-modal retrieval. Background Art
[0002] The power industry is an important part of the country's infrastructure construction. In construction activities such as power line erection, substation construction, equipment installation and maintenance, the construction scene is often complex and changeable, involving various high-risk operation types such as high-altitude operation, confined space operation, and live operation. Once a safety accident occurs, it may cause serious casualties, equipment damage and economic losses. Therefore, the early warning and control of construction safety risks are the core issues in the safety management of the power industry and an important measure to ensure the life safety of construction workers and the smooth implementation of projects. However, the current construction site safety risk early warning technology still faces many limitations:
[0003] (1) The complex distribution of the construction environment limits the retrieval accuracy of the early warning technology. The construction environment usually covers various environments such as indoor, outdoor, and urban areas, and the internal safety risks of each environment are different;
[0004] (2) Information islanding limits the cross-modal representation ability of the early warning technology. A large amount of multi-modal data such as pictures and texts generated during the construction process lacks a unified representation;
[0005] (3) Insufficient intelligence limits the multi-modal semantic understanding of the early warning technology. Existing technologies mostly use risk identification methods based on a single modality (such as images or texts), lacking the fusion and in-depth semantic understanding of multi-modal information, and it is difficult to meet the risk early warning requirements in complex scenarios.
[0006] To address the limitations of the above early warning technology, our goal is to use cross-modal combined image retrieval technology to map the images at the construction site and the relevant risk early warning text data to the same semantic space, so that the two can be effectively combined and matched in the same framework. However, due to the following three challenges, this is not an easy task.
[0007] (1) Difficulty in decoupling multi-granularity features: The complexity of the construction environment makes the image contain various semantic information. For example, there may be various elements such as electrical equipment, staff, and environmental factors at the construction site at the same time. The relationships between these elements are complex, making it difficult for us to effectively decouple the specific semantic content corresponding to the text from the image. This diversity increases the difficulty of feature extraction and matching.
[0008] (2) Difficulty in cross-modal feature alignment: Images and texts belong to different modalities, and there are significant differences in data distribution, information expression, and semantic understanding. Images are usually high-dimensional pixel data, while texts are linearly structured symbolic information. Therefore, how to achieve efficient cross-modal feature alignment and ensure the consistency of images and texts at the semantic level is an extremely challenging task. This requires not only advanced algorithm support, but also a deep understanding of the intrinsic connection between the two modalities.
[0009] (3) Difficulty in cross-modal semantic fusion: In practical applications, effectively combining the semantics of construction site images and safety risk texts to make them close to the target risk images to be retrieved is a difficult problem that needs to be solved urgently. We need to design a mechanism that can achieve semantic complementarity and fusion while retaining their respective characteristics. This process requires not only efficient algorithms, but also a deep understanding of the model in the field of construction safety to ensure that the fused information can accurately reflect the actual safety risks. Summary of the invention
[0010] In view of the shortcomings of the prior art, the present invention provides a multi-granularity semantically decoupled combined image retrieval method based on a pre-trained visual language model CLIP (CLIP-based multi-gRanularitysEmanticdeCoupledneTwork, abbreviated as ERECT), which aims to perform cross-modal semantic alignment and combination of image text features extracted by CLIP, so as to achieve a high-precision combined image retrieval method.
[0011] The present invention aims to solve a complex combined image retrieval problem: how to find a target image that meets the multimodal combined query requirements from a set of images. is a set of N triples, where x r is the input image modality information, t m To input text information, x t is the target image to be retrieved. The core goal is to learn a metric space so that the multimodal query (x r ,t m ) and the corresponding target image x t The embedding features of in Represents the embedding functions of the multimodal query and target images that need to be learned.
[0012] To this end, the present invention first designs a multi-granularity image-text semantic feature decoupling module. The main purpose of this module is to effectively decouple the data distributions, information expression methods, and semantic understandings of images and texts, so as to extract aligned image and text features. This process provides a solid alignment foundation for subsequent cross-modal combinations, ensuring that features between different modalities can correspond to each other at the semantic level. Secondly, the present invention designs a multi-granularity feature combination module, aiming to combine features at different granularity levels to obtain combined semantic information at different granularity details, enrich the feature expression, enhance the model's understanding ability of complex semantic relationships, and enable the model to obtain superior retrieval accuracy when processing diverse data. Finally, the present invention designs a multi-granularity combination-target alignment module, which can drive the combined features to approach the target image features at the multi-granularity level, realize cross-modal alignment at the multi-granularity detail level, and more precisely align the semantic information between the multi-modal query and the target image, thereby improving the overall cross-modal understanding and application effects.
[0013] Term Explanation:
[0014] 1. Triplets: Refers to the metadata structure used in the training process of the combined image retrieval model. Specifically, it consists of three parts: input image modality information, input text information, and the corresponding target image. The input image modality information provides a benchmark for representing the visual content we hope to search for; the input text information is used to describe specific modifications, changes, categories, or descriptions of the reference image, helping the model understand how to generate the target image from the reference image; and the target image is the final result we expect the model to recognize or generate.
[0015] 2. Multi-modal query refers to the query input form in the combined image retrieval model, mainly including input image modality information and input text information. By combining these two different modality information together, the model can more comprehensively understand the user's query intention. The input image modality information provides specific visual information, while the input text information provides context and semantic guidance for the image. Such a query method enables the model to more accurately capture the user's needs when processing complex image retrieval tasks, thereby improving the relevance and accuracy of the retrieval.
[0016] 3. CLIP (Contrastive Language–Image Pre-training) is an advanced pre-trained vision-language model designed to extract unified embedding features for text and images. Trained on a large-scale dataset, CLIP can learn the deep relationships between text and images, thus performing excellently in multi-modal tasks. Its core idea is to optimize the model through contrastive learning, making similar text and images closer in the embedding space while pulling unrelated text and images farther apart. This feature enables CLIP to have powerful capabilities in tasks such as image classification, retrieval, and generation, making it an important tool in the field of multi-modal learning.
[0017] 4. Transformer architecture: A deep learning model architecture widely used in sequence-to-sequence processing. This architecture is completely based on the attention mechanism and can effectively capture semantic context information.
[0018] 5. Multilayer perceptron: A feedforward neural network composed of multiple neuron layers, usually used to handle non-linear problems, and realizes complex mappings between input and output through full connections between layers.
[0019] The technical solution of the present invention is as follows:
[0020] A construction safety risk early warning method based on cross-modal visual language retrieval, including the following steps:
[0021] (1). Understand each element of the given triple and extract the aligned image and text features respectively, providing a solid alignment basis for subsequent cross-modal combination, ensuring that different modal features correspond to each other at the semantic level. And based on the image-text alignment features, effectively decouple the data distribution, information expression mode, and semantics of images and texts, providing effective multi-granularity decoupled feature information for subsequent cross-modal combination;
[0022] (2). For the extracted multi-granularity decoupled feature information, complete the extraction of combined features for multi-modal queries from the multi-granularity level, obtain combined semantic information at different granularity details, and enrich the expression of combined features;
[0023] (3). Align the multi-granularity combined features with the multi-granularity target features, promote the combined features to approach the target image features from the multi-granularity level, realize cross-modal alignment at the multi-granularity detail level, and accurately align the semantic information between the multi-modal combination and the target image, thereby improving the cross-modal understanding and application of the overall model.
[0024] (4) After precisely aligning the semantic information between the multi-modal combination and the target image, a text-image multi-modal combination retrieval model is obtained. Using this model to process multi-modal query information, retrieve the corresponding risk target images for both, match the risk content, and issue a warning.
[0025] Preferably, in step (1), the specific steps of understanding each element of the given triple and separately extracting the aligned image and text features and effectively decoupling the semantic features of the image and text include:
[0026] 1-1. Based on the image-text encoder of the advanced pre-trained vision-language model CLIP, respectively perform aligned feature extraction on the input image modality and text modality to obtain aligned semantic features;
[0027] 1-2. Based on the semantic features of the image-text alignment in 1-1, use a multi-layer perceptron and a Transformer architecture to decouple multi-granularity channel features, providing effective multi-granularity decoupled information for subsequent cross-modal combinations.
[0028] Further preferably, in step 1-1, based on the image-text encoder of the advanced pre-trained vision-language model CLIP, respectively perform aligned feature extraction on the input image modality x r and text modality t m , and the corresponding target image to obtain aligned semantic features. Specifically,
[0029] First, extract the global and local granularity features of the input image modality x r , which are formulated as follows,
[0030]
[0031] where respectively represent the second-to-last layer and the last layer of the CLIP image encoder, represents the local feature of the input image modality, represents the global feature of the input image modality, D represents the feature embedding dimension of CLIP, C represents the number of image channels, and the global feature of the target image can be extracted in the same way and the local feature of the target image Similarly, use the second-to-last layer and the last layer of the CLIP text encoder to extract the global feature and local feature S represents the text sequence length.
[0032] Further preferably, in step 1-2, in order to decouple the specific semantic content corresponding to the text from the image, this step is based on the image-text alignment features, and uses a multi-layer perceptron and a Transformer architecture to decouple the multi-granularity channel features, providing effective multi-granularity decoupled feature information for subsequent cross-modal combinations. This step uses the local features of the image as an example, and the decoupling process is specifically described as follows:
[0033] Since the image channels may include multiple color channels (such as RGB), while the text sequence is composed of words or characters, and the quantity may be much less than the image channels. Therefore, in order to solve the asymmetry between the number of image channels and the number of text sequences, this step needs to perform channel decoupling, decoupling the channels of the image and the sequences of the text into the same channels, so that in the subsequent combination process, the model can effectively integrate the information of the image and the text, thereby improving the mutual understanding and matching ability of cross-modal features. Specifically, this step feeds the local features of the image into a multi-layer perceptron to learn the local features after decoupling the channels Formulated as follows,
[0034]
[0035] where T is the number of channels to be decoupled; after obtaining the local features after decoupling the channels, in order to enhance the learnability of the decoupled features and the semantic representation ability of each decoupled channel, this step uses the Transformer architecture to perform attention interaction between the decoupled features and the original local features to learn more detailed semantic representations, formulated as follows,
[0036]
[0037] where represents the local decoupled features of the image, [:T] represents taking the first T channels output by the Transformer architecture to ensure the alignment of the number of channels of subsequent image-text features;
[0038] Similarly, this step extracts the local decoupled features corresponding to the local features of the text in the same way, formulated as follows,
[0039]
[0040] where is the local feature after decoupling the text channels; represents the local decoupled features of the text;
[0041] In addition, for multi-granularity semantic understanding, corresponding decoupling operations are also performed on the global features of images and texts respectively in this step, which are formulated as follows.
[0042]
[0043] Among them are the global features after decoupling the channels of images and texts respectively. are the global decoupled features of images and texts respectively. Finally, in this step, the global and local decoupled features are concatenated to generate the multi-granularity decoupled features of images and texts respectively, which are formulated as follows.
[0044]
[0045] Among them are the multi-granularity decoupled features of images and texts respectively. For the target image, the same method is used in this step for decoupling to obtain the multi-granularity decoupled features of the target image.
[0046] Preferably, in step (2), in order to obtain the combined semantic information at different granularity details and enrich the expression of the combined features, for the extracted multi-granularity decoupled information, the combined feature extraction of multi-modal queries is completed from different granularity levels, so that the combined features not only focus on the unique information at each granularity level, but also enhance the mutual relationship between them to achieve more comprehensive and detailed feature integration. The specific steps include:
[0047] 2-1. In order to learn the semantic information at different granularity details, a multi-layer perceptron is used to identify the weights of the semantic corresponding parts in the image and text; adaptively learn the correlation weights w of the semantic corresponding parts in the multi-granularity image and text features, and effectively capture the correlation between the image and the text, which is formulated as follows.
[0048] w = MLP([F r , F m )(8)
[0049] Among them represents the corresponding part weights between each channel.
[0050] 2-2. Weight the weights of the semantic corresponding parts to the multi-granularity decoupled features of the image and text, so that the semantic corresponding parts of the image and text are semantically enhanced.
[0051] The calculated correlation weights w are respectively weighted into the multi-granularity decoupled features of the image and the text to highlight the corresponding parts between the image and the text, thereby enhancing their correlation in the multi-granularity feature representation. In this way, the matching degree of the image and the text at the semantic level is effectively improved, so that the features at each granularity level are more closely combined, and then the understanding ability and performance of the overall model are improved. It is formulated as follows,
[0052]
[0053] 2-3. Using the enhanced image-text features, perform multi-granularity combination to obtain combined semantic information at different granularity details and enrich the expression of the combined features. It is formulated as follows,
[0054] F c =F r +F m (10)
[0055] where represents the multi-granularity combined feature.
[0056] Preferably, in step (3), the specific steps for aligning the multi-granularity combined feature - multi-granularity target feature include:
[0057] Using the cosine similarity as the distance metric to calculate the similarity matrix between the multi-granularity combined feature of the multi-modal query and the multi-granularity feature of the target image; at the same time, using the batch-based classification loss to understand the similarity between the multi-granularity combined feature of the multi-modal query and the multi-granularity feature of the target image, and promoting the proximity of their semantic representations;
[0058] To make the multi-granularity combined feature F c closer to the multi-granularity decoupled feature F t of the multi-granularity of the target image, this step adopts the batch-based classification loss that has been widely used in the combined image retrieval task. Its formula is as follows,
[0059]
[0060] where B is the batch size, respectively represent the i-th average-pooled multi-granularity combined feature F c and the multi-granularity decoupled feature F t of the multi-granularity of the target image in the same batch; s is the cosine similarity calculation function, τ is the temperature coefficient; the subscript j is the traversal value in the denominator;
[0061] Finally, the optimization function of the following multi-granularity semantic decoupled combined image retrieval model ERECT is obtained,
[0062]
[0063] where Θ is the parameter to be learned for ERECT, and Θ * is the value of the parameter to be learned that minimizes the loss function.
[0064] Preferably, in step (4), the optimized multi-granularity semantic decoupling combined image retrieval model ERECT is used for target image risk warning:
[0065] For the input multi-modal query (x r , t m ), where x r is the input image modal information and t m is the input text information, the multi-modal query is input into the ERECT model for multi-granularity semantic decoupling of the image and text, and the two (image and text) are multi-modally combined to obtain the multi-granularity combined feature F c ; and there is a predefined candidate risk image library where represents the s-th candidate risk image, p represents the image, and S is the number of images; the present invention inputs all candidate risk images into the ERECT model for multi-granularity decoupling of the images to obtain the multi-granularity feature set of the candidate risk images
[0066] Subsequently, the present invention uses the cosine similarity as the measure of similarity, calculates the similarity between the multi-granularity combined feature F c and each multi-granularity feature in the multi-granularity feature set of the candidate risk images, and sorts them to obtain the similarity set If the first similarity is greater than the set threshold σ (usually set to 0.95), it is determined that the multi-modal combined query probably contains the content corresponding to the risk image, and a risk warning is issued.
[0067] A construction safety risk warning system based on cross-modal visual language retrieval includes:
[0068] A multi-granularity image-text semantic feature decoupling module, configured to: effectively decouple the data distribution, information expression mode, and semantic understanding of the image and text, so as to extract the aligned image and text features;
[0069] A multi-granularity feature combination module, configured to: combine the features from different granularity levels to obtain combined semantic information at different granularity details;
[0070] The multi-granularity combination - target alignment module is configured to: promote the combined features to approach the target image features from multiple granularity levels, achieve cross-modal alignment at the multi-granularity detail level, and more accurately align the semantic information between images and texts.
[0071] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the steps in the above-mentioned construction safety risk warning method based on cross-modal visual language retrieval.
[0072] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned construction safety risk warning method based on cross-modal visual language retrieval.
[0073] Through the multi-granularity semantic decoupling technology, the present invention decouples the local and global features of images and texts respectively, and at the same time considers the alignment between the multi-modal query and the target image, rather than the alignment between the image and the text. The introduction of the multi-modal query expands the input of the present invention to the scenario of simultaneous input of images and texts. Through the multi-granularity feature combination module designed by the present invention, the decoupled features of the image and the text in the multi-modal query are combined at multiple granularity levels to obtain the multi-granularity combined features of the multi-modal query. Subsequently, the present invention designs a multi-granularity combination - target alignment module to promote the combined features (simultaneously containing the semantic information of the image and the text in the multi-modal query) to approach the target image features from multiple granularity levels, achieve cross-modal alignment at the multi-granularity detail level, and more accurately align the semantic information between the multi-modal query and the target image, thereby improving the overall cross-modal understanding and application effect.
[0074] The beneficial effects of the present invention are as follows:
[0075] 1. The present invention designs a multi-granularity semantic decoupling and alignment mechanism. By using the pre-trained vision-language model CLIP, the image and text features of a given triple are effectively decoupled to ensure the consistency of different modal features at the semantic level. This decoupling mechanism provides a solid foundation for subsequent cross-modal combinations and can effectively extract and align the semantic features of images and texts.
[0076] 2. The present invention performs multi-granularity feature combination and enhancement, and extracts combined features of multi-modal queries from different granularity levels. By identifying and weighting the weights of the semantically corresponding parts in the image and the text, the expression ability of the combined features is enhanced, enabling the model to better capture complex semantic relationships and improving the richness of the features.
[0077] 3. The present invention designs a new combined image retrieval model, namely a multi-granularity semantic decoupling combined image retrieval method based on the pre-trained vision-language model CLIP, which achieves excellent retrieval performance on multiple international standard datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 FIG. is the overall flowchart of a construction safety risk early warning method based on cross-modal vision-language retrieval according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The present invention will be further described below by way of examples in conjunction with the accompanying drawings, but is not limited thereto.
[0080] Example 1:
[0081] A construction safety risk early warning method based on cross-modal vision-language retrieval, as Figure 1 shown, includes the following steps:
[0082] (1). Understand each element of the given triple and separately extract the aligned image and text features, providing a solid alignment basis for subsequent cross-modal combination to ensure that different modal features correspond to each other at the semantic level. And based on the image-text alignment features, effectively decouple the data distribution, information expression mode, and semantics of images and texts, providing effective multi-granularity decoupled feature information for subsequent cross-modal combination; specifically including:
[0083] 1-1. Based on the image-text encoder of the advanced pre-trained vision-language model CLIP, respectively extract the alignment features of the input image modality and text modality to obtain the aligned semantic features.
[0084] Based on the image-text encoder of the advanced pre-trained vision-language model CLIP, respectively extract the alignment features of the input image modality x r and the text modality t m , and the corresponding target image to obtain the aligned semantic features. Specifically, first extract the global and local granularity features of the input image modality x r , which are formulated as follows,
[0085]
[0086] where respectively represent the penultimate layer and the last layer of the CLIP image encoder, represents the local feature of the input image modality, represents the global feature of the input image modality, D represents the feature embedding dimension of CLIP, C represents the number of image channels, and the global feature of the target image can be extracted in the same way Local features of the target image Similarly, the penultimate layer and the last layer of the text encoder of CLIP are used to extract the global features of the text and local features S represents the length of the text sequence.
[0087] 1-2. Based on the semantic features of the image-text alignment in 1-1, a multi-layer perceptron and a Transformer architecture are used to decouple the multi-granularity channel features, providing effective multi-granularity decoupled information for subsequent cross-modal combinations.
[0088] In step 1-2, in order to decouple the specific semantic content corresponding to the text from the image, this step is based on the image-text alignment features, and a multi-layer perceptron and a Transformer architecture are used to decouple the multi-granularity channel features, providing effective multi-granularity decoupled feature information for subsequent cross-modal combinations. In this embodiment, the local features of the image are taken as an example, and the decoupling process is specifically described as follows:
[0089] Since the image channels may include multiple color channels (such as RGB), while the text sequence is composed of words or characters, and the quantity may be much less than the image channels, in order to solve the asymmetry between the number of image channels and the number of text sequences, this step needs to perform channel decoupling, decoupling the channels of the image and the sequences of the text into the same channels, so that in the subsequent combination process, the model can effectively integrate the information of the image and the text, thereby improving the mutual understanding and matching ability of cross-modal features. Specifically, this step takes the local features of the image and feeds them into a multi-layer perceptron to learn the local features after decoupling the channels Formulated as follows,
[0090]
[0091] where T is the number of channels to be decoupled; after obtaining the local features after decoupling the channels, in order to enhance the learnability of the decoupled features and the semantic representation ability of each decoupled channel, this step uses the Transformer architecture to perform attention interaction between the decoupled features and the original local features to learn more detailed semantic representations, formulated as follows,
[0092]
[0093] where represents the local decoupled features of the image, and [:T] represents taking the first T channels output by the Transformer architecture to ensure the alignment of the number of channels of subsequent image-text features;
[0094] Similarly, in the same way, this step extracts the local features of the text. The corresponding local decoupled features are formulated as follows:
[0095]
[0096] where is the local feature after text decoupling channel; represents the local decoupled feature of the text;
[0097] In addition, for multi-granularity semantic understanding, this step also performs corresponding decoupling operations on the global features of images and texts respectively, which are formulated as follows:
[0098]
[0099] where are the global features after image and text decoupling channels respectively, are the global decoupled features of images and texts respectively; Finally, this step splices the global and local decoupled features to generate the multi-granularity decoupled features of images and texts respectively, which are formulated as follows:
[0100]
[0101] where are the multi-granularity decoupled features of images and texts respectively. For the target image, this step uses the same method for decoupling to obtain the multi-granularity decoupled features of the target image.
[0102] (2) For the extracted multi-granularity decoupled feature information, complete the extraction of combined features for multi-modal queries from the multi-granularity level, obtain combined semantic information at different granularity details, and enrich the expression of combined features.
[0103] To obtain combined semantic information at different granularity details and enrich the expression of combined features, for the extracted multi-granularity decoupled information, complete the extraction of combined features for multi-modal queries from different granularity levels, so that the combined features not only focus on the unique information at each granularity level, but also enhance the relationship between them to achieve more comprehensive and detailed feature integration. The specific steps include:
[0104] 2-1. To learn semantic information at different granularity details, use a multi-layer perceptron to identify the weights of the semantically corresponding parts in images and texts; adaptively learn the correlation weights w of the semantically corresponding parts in multi-granularity image and text features, and effectively capture the correlation between images and texts, which is formulated as follows:
[0105] w = MLP([F r , F m )(8)
[0106] Among them represents the corresponding part weights between each channel.
[0107] 2-2. Weight the weights of the semantic corresponding parts to the multi-granularity decoupled features of the image and text, so that the semantic corresponding parts of the image and text are semantically enhanced respectively.
[0108] The calculated correlation weights w are weighted into the multi-granularity decoupled features of the image and text respectively to highlight the corresponding parts between the image and the text, thereby enhancing their correlation in the multi-granularity feature representation. In this way, the matching degree of the image and the text at the semantic level is effectively improved, so that the features at each granularity level are more closely combined, and then the understanding ability and performance of the overall model are improved. The formula is as follows.
[0109]
[0110] 2-3. Use the enhanced image-text features for multi-granularity combination to obtain combined semantic information at different granularity details and enrich the expression of the combined features. The formula is as follows.
[0111] F c = F r + F m (10)
[0112] Among them represents the multi-granularity combined feature.
[0113] (3). Align the multi-granularity combined feature - multi-granularity target feature, promote the combined feature to approach the target image feature from the multi-granularity level, achieve cross-modal alignment at the multi-granularity detail level, and accurately align the semantic information between the multi-modal combination and the target image, so as to improve the cross-modal understanding and application of the overall model.
[0114] Specifically, it includes: using the cosine similarity as the distance metric to calculate the similarity matrix between the multi-granularity combined feature of the multi-modal query and the multi-granularity feature of the target image; at the same time, using the batch-based classification loss to understand the similarity between the multi-granularity combined feature of the multi-modal query and the multi-granularity feature of the target image, and promoting the semantic representations of the two to be close.
[0115] In order to make the multi-granularity combined feature F c closer to the multi-granularity decoupled feature F t of the multi-granularity of the target image, this step adopts the batch-based classification loss that has been widely used in the combined image retrieval task. The formula is as follows.
[0116]
[0117] Among them, B is the batch size, respectively representing the i-th multi-granularity combined feature F after average pooling in the same batch c and the multi-granularity decoupled feature F of the target image's multi-granularity t ; s is the cosine similarity calculation function, τ is the temperature coefficient; the subscript j is the traversal value in the denominator;
[0118] Finally, the following optimization function of the multi-granularity semantic decoupled combined image retrieval model ERECT is obtained,
[0119]
[0120] where Θ is the parameter to be learned of ERECT, Θ * is the value of the parameter to be learned that minimizes the loss function.
[0121] (4) After accurately aligning the semantic information between the multi-modal combination and the target image, a text-image multi-modal combination retrieval model is obtained. Using this model to process multi-modal query information, retrieve the corresponding risk target image of the two, match the risk content and give an early warning.
[0122] Use the optimized multi-granularity semantic decoupled combined image retrieval model ERECT for target image risk early warning:
[0123] For the input multi-modal query (x r , t m ), where x r is the input image modal information, t m is the input text information. Input this multi-modal query into the ERECT model for multi-granularity semantic decoupling of the image and the text, and perform multi-modal combination on the two (image and text) to obtain the multi-granularity combined feature F of this multi-modal query c ; and there is a predefined candidate risk image library where represents the s-th candidate risk image, p represents the image, and S is the number of images; the present invention inputs all candidate risk images into the ERECT model for multi-granularity decoupling of the images to obtain the multi-granularity feature set of the candidate risk images
[0124] Subsequently, the present invention uses the cosine similarity as the similarity metric standard to calculate the similarity between the multi-granularity combined feature F c and each multi-granularity feature in the multi-granularity feature set of the candidate risk images, and performs sorting to obtain the similarity set If the first similarity is greater than the set threshold σ (usually set to 0.95), it is determined that the multimodal combination query probably contains the content corresponding to the risk image, and a risk warning is issued.
[0125] Embodiment 2:
[0126] A construction safety risk warning system based on cross-modal visual language retrieval, comprising:
[0127] A multi-granularity image-text semantic feature decoupling module, configured to: effectively decouple the data distribution, information expression mode, and semantic understanding of images and texts, so as to extract aligned image and text features;
[0128] A multi-granularity feature combination module, configured to: combine features from different granularity levels to obtain combined semantic information at different granularity details;
[0129] A multi-granularity combination-target alignment module, configured to: promote the combined features to approach the target image features from multiple granularity levels, realize cross-modal alignment at multiple granularity detail levels, so as to more accurately align the semantic information between images and texts.
[0130] Embodiment 3:
[0131] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in a construction safety risk warning method based on cross-modal visual language retrieval as described in Embodiment 1.
[0132] Embodiment 4:
[0133] An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in a construction safety risk warning method based on cross-modal visual language retrieval as described in Embodiment 1.
[0134] Experimental Example
[0135] The multi-granularity semantic decoupling and combination image retrieval method based on the pre-trained vision-language model CLIP proposed by the present invention, abbreviated as ERECT, was compared for retrieval recall rates on the globally recognized combined image retrieval datasets FashionIQ (including three subclasses: Dresses, Shirts, and Tops&Tees), Shoes, and CIRR datasets to verify the effectiveness of this embodiment, and the best retrieval recall rates were obtained. Table 1-3 respectively show the comparison results of the retrieval recall rates of this embodiment and each method on the four datasets of FashionIQ, Shoes, and CIRR.
[0136] Table 1
[0137]
[0138] Table 2
[0139]
[0140] Table 3
[0141]
[0142] As can be seen from the comparison in the above table, the solution of the present application has excellent retrieval performance.
Claims
1. A construction safety risk early warning method based on cross-modal visual language retrieval, characterized in that: The steps include: (1) Understand each element of a given triple and extract the aligned image and text features respectively, providing an alignment basis for subsequent cross-modal combination. Based on the image-text alignment features, decouple the data distribution, information expression, and semantics of the image and text, providing multi-granularity decoupled feature information for subsequent cross-modal combination. (2) Based on the extracted multi-granularity decoupled feature information, the combined feature extraction of multimodal query is completed from the multi-granularity level to obtain the combined semantic information at different granularity details; (3) Align the multi-granularity combination features with the multi-granularity target features, push the combination features to approach the target image features from the multi-granularity level, achieve cross-modal alignment at the multi-granularity detail level, and align the semantic information between the multi-modal combination and the target image; (4) After aligning the semantic information between the multimodal combination and the target image, a text-image multimodal combination retrieval model is obtained. The model is used to process the multimodal query information, retrieve the corresponding risk target image, match the risk content and issue an early warning.
2. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 1 is characterized in that: In step (1), the specific steps of understanding each element of a given triplet and extracting the aligned image and text features and decoupling the semantic features of the image and text include: 1-1. An image-text encoder based on the pre-trained visual language model CLIP performs alignment feature extraction on the input image modality and text modality to obtain aligned semantic features; 1-2. Based on the semantic features of image-text alignment in 1-1, multi-layer perceptron and Transformer architecture are used to decouple multi-granularity channel features to provide multi-granularity decoupling information for subsequent cross-modal combination.
3. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 2 is characterized in that: In step 1-1, the image-text encoder based on the advanced pre-trained visual language model CLIP is used to encode the input image modality x r With text modal m , and the corresponding target image to extract the aligned features and obtain the aligned semantic features. Specifically, First extract the input image modality x r The global and local granularity characteristics of are formulated as follows: in Represent the penultimate and last layers of the CLIP image encoder, respectively. Represents the local features of the input image modality, Represents the global features of the input image modality, D represents the feature embedding dimension of CLIP, and C represents the number of image channels. The global features of the target image are extracted in the same way. Local features of the target image Similarly, the penultimate and final layers of the CLIP text encoder are used to extract the global features of the text. With local features S represents the length of the text sequence.
4. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 2 is characterized in that: In step 1-2, based on the image-text alignment features, the multi-layer perceptron and Transformer architecture are used to decouple the multi-granularity channel features, providing multi-granularity decoupled feature information for subsequent cross-modal combination. The decoupling process is as follows: Channel decoupling is performed to decouple the image channel and text sequence into the same channel, so that the model can integrate the information of the image and text in the subsequent combination process. Specifically, the local features of the image are Send it to the multi-layer perceptron to learn the local features after the decoupling channel The formula is as follows, in T is the number of channels that need to be decoupled. After obtaining the local features after the decoupled channels, the Transformer architecture is used to perform attention interaction between the decoupled features and the original local features. The formula is as follows: in Represents the local decoupled features of the image, and [:T] represents the first T channels output by the Transformer architecture; Extract local features of text in the same way The corresponding local decoupling feature is formulated as follows: in The local features after decoupling the channels for text; Represents the local decoupling features of the text; Corresponding decoupling operations are also performed on the global features of images and texts, and the formula is as follows: in They are the global features after decoupling channels of image and text, They are the global decoupled features of image and text respectively; finally, this step concatenates the global and local decoupled features to generate multi-granularity decoupled features of image and text respectively, which are formulated as follows: where F m , They are the multi-granularity decoupled features of images and texts respectively. For the target image, the same method is used for decoupling to obtain the multi-granularity decoupled features of the target image.
5. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 1 is characterized in that: In step (2), for the extracted multi-granularity decoupled information, combined feature extraction of multimodal query is completed from different granularity levels. The specific steps include: 2-1. Use a multi-layer perceptron to identify the weights of the semantically corresponding parts of the image and text; adaptively learn the correlation weights w of the semantically corresponding parts of the multi-granular image and text features to capture the correlation between the image and text. The formula is as follows: w=MLP([F r ,F m ])(8) in Represents the corresponding weights between each channel; 2-2. The weights of the semantically corresponding parts are added to the multi-granularity decoupled features of the image and text, so that the semantically corresponding parts of the image and text are semantically enhanced; The calculated correlation weight w is weighted to the multi-granularity decoupled features of the image and text respectively, and the formula is as follows: 2-3. Using the enhanced image text features, we can perform multi-granularity combination to obtain the combined semantic information of different granularity details. The formula is as follows: F c =F r +F m (10) in Represents multi-granularity combined features.
6. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 1 is characterized in that: In step (3), the specific steps of aligning the multi-granularity combined features with the multi-granularity target features include: The cosine similarity is used as the distance metric to calculate the similarity matrix between the multi-granularity combined features of the multi-modal query and the multi-granularity features of the target image. At the same time, the batch-based classification loss is used to understand the similarity between the multi-granularity combined features of the multi-modal query and the multi-granularity features of the target image, and to promote the semantic representation of the two. In order to make the multi-granularity combined feature F c The multi-granularity decoupled feature F is closer to the multi-granularity of the target image. t ,This step adopts batch-based classification loss, and its formula is as follows, Where B is the batch size, Represents the multi-granularity combination feature F after pooling of the i-th in the same batch c and the multi-granularity decoupled features F of the target image t ; s is the cosine similarity calculation function, τ is the temperature coefficient; subscript j is the ergodic value on the denominator; Finally, the optimization function of the multi-granularity semantic decoupling combined image retrieval model ERECT is obtained as follows: Where Θ is the parameter to be learned of ERECT, Θ * is the value of the parameter to be learned that minimizes the loss function.
7. The construction safety risk early warning method based on cross-modal visual language retrieval according to claim 1 is characterized in that: In step (4), the optimized multi-granularity semantic decoupling combined image retrieval model ERECT is used to perform risk warning of target images: For the multimodal query (x r ,t m ), where x r is the input image modality information, t m To input text information, the multimodal query is input into the ERECT model to perform multi-granular semantic decoupling of image and text, and the two are multimodally combined to obtain the multi-granular combined feature F of the multimodal query. c ; and there is a pre-defined candidate risk image library in represents the sth candidate risk image, p represents the image, and S is the number of images; all candidate risk images are input into the ERECT model for multi-granular decoupling of the images to obtain a multi-granular feature set of the candidate risk images Then, cosine similarity is used as the similarity metric to calculate the multi-granularity combined feature F c Calculate the similarity with each multi-granularity feature in the multi-granularity feature set of the candidate risk image, and sort them to obtain a similarity set If the first similarity is greater than the set threshold σ, it is determined that the multimodal combined query contains the content corresponding to the risk image, and a risk warning is issued.
8. A construction safety risk early warning system based on cross-modal visual language retrieval, characterized in that: include: The multi-granular image-text semantic feature decoupling module is configured to effectively decouple the data distribution, information expression, and semantic understanding of images and texts, thereby extracting aligned image and text features; The multi-granularity feature combination module is configured to: combine features from different granularity levels to obtain combined semantic information at different granularity details; The multi-granularity combination-target alignment module is configured to: push the combination features to approach the target image features from a multi-granularity level, and achieve cross-modal alignment at a multi-granularity detail level to align the semantic information between the image and the text.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the steps in a construction safety risk early warning method based on cross-modal visual language retrieval as described in any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: It includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the construction safety risk early warning method based on cross-modal visual language retrieval as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Construction method of multi-modal large model for security risk identification
CN120493188A
Cross-modal representation learning and retrieval method and system for grain production
CN120705355A
A cross-modal representation learning and retrieval method and system for grain production
CN120705355B
Multi-modal data fusion modeling method and system based on multi-task learning
CN120744812A
Cross-modal semantic alignment driven power grid equipment fault diagnosis method and system
CN121071667A