Large-model-driven image local quality evaluation method and system in mine complex scene
By employing a vision-language model with guided prompts and cross-modal interaction, the method addresses the challenge of local quality evaluation in mining environments, enhancing the model's ability to assess both local and global image quality accurately.
Patent Information
- Application Number
- CN202510357242.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing blind image quality evaluation model is difficult to effectively extract image local quality features in mine scenarios, and lacks the ability to reasonably evaluate local areas of the image.
The large-model-driven method is adopted to obtain the local area of the image significance by using the open vocabulary detection model, combine the visual language big model and the large language model to generate description text, and cross-modal interaction is carried out through the visual language interaction model to build a local-global feature aggregation module to enhance the model's perceived evaluation ability for local areas.
The model's quality evaluation ability of local areas of the image is improved, while maintaining effective evaluation of the entire image, improving the safety of mine operations and the accuracy of image quality evaluation.
Smart Images

Figure CN120318162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image local quality assessment method and system driven by a large model in a complex mine scenario, belonging to the technical field of image quality evaluation. Background Art
[0002] In mine operation scenarios, restricted by factors such as complex underground lighting conditions, dust and water mist interference, and equipment vibration, the roadway structure images, equipment status images, etc. collected by industrial cameras, mine inspection robots, or drones often have low quality, which affects the decision-making accuracy of mine safety monitoring systems. Image quality evaluation provides an important guidance for optimizing mine imaging hardware facilities and image restoration algorithm design by accurately extracting key factors affecting mine image quality, helps improve the practicability of key technologies such as underground fog-penetrating imaging and remote diagnosis of equipment status, and significantly reduces the operation and maintenance costs and safety risks caused by image misjudgment.
[0003] According to whether a reference image is required, image quality evaluation can be divided into full-reference image quality evaluation (FR-IQA), reduced-reference image quality evaluation (RR-IQA), and no-reference image quality evaluation (NR-IQA), also known as blind image quality evaluation (BIQA). In mine scenarios, it is difficult to obtain an ideal reference image due to the dynamic changes in the underground environment, which makes no-reference image quality evaluation (NR-IQA) an indispensable core technology in mine intelligent construction.
[0004] In recent years, thanks to the "end-to-end" learning paradigm of deep learning, existing blind image quality evaluation models can effectively complete the inference from quality features to quality scores under the supervision of quality score labels. However, due to the widespread use of transfer learning methods and the lack of strong supervision means for extracting quality features, the existing models have insufficient ability to extract local image quality features. At the same time, limited by the lack of local image quality score labels, it is difficult for existing models to give reasonable evaluation scores for local images based on the overall image.
[0005] Currently, the interaction of cross-modal information has become an effective means to enhance the model's perception ability, and the emergence of vision-language large models provides detailed and comprehensive text-modal information for images. Vision-Language Models are deep learning models that combine vision and language processing, and they can simultaneously understand and generate the association between images and texts. Through multi-modal learning, these models combine visual information (such as images, videos) with language information (such as texts, descriptions) for common understanding and reasoning. In industrial scenarios such as mine intelligence, vision-language large models can quickly adapt to the dynamic underground environment and provide more robust semantic-level analysis capabilities for tasks such as image quality evaluation. Summary of the Invention
[0006] The object of the present invention is to provide a method and system for large model-driven image local quality assessment in complex mine scenarios. The method and system can enhance the evaluation ability of the blind image quality evaluation model for local quality features, enabling it to have the evaluation ability for both the entire image and local parts of the image, and forming effective evaluations for the overall image and any specified region.
[0007] To achieve the above object, the present invention provides a method for large model-driven image local quality assessment in complex mine scenarios, including the following steps:
[0008] S1. Obtain the significant local regions of the image based on an open vocabulary detection model, providing local visual semantic information for the model to enhance its local quality perception and evaluation ability;
[0009] S2. Set up prompts for a vision-language large model guided by the chain of thought, and use the prompts to obtain the descriptive text of the significant local regions of the image and the local-global quality inference text of the image, including information on both the content and quality of the image;
[0010] S3. Obtain a complementary text group of the descriptive text of the significant local regions of the image and the local-global quality inference text based on a large language model, providing more refined text modality supervision information for the model to enhance its local quality evaluation ability and form a good local-global quality inference logic;
[0011] S4. With the help of the visual feature encoder of the vision-language interaction model, use the method of spatial mapping to match the positions in the image pixel domain and feature domain one by one, thereby extracting the local and global features of the image;
[0012] S5. Construct a local-global feature aggregation module based on the transposed attention and multi-head attention mechanisms to aggregate the local and global features of the image obtained in S4;
[0013] S6. With the help of the cross-modal interaction ability of the vision-language interaction model, set the training paradigm of the model, and use the complementary text group related to the local description of the image obtained in S3 to guide the refinement of the local features of the image obtained in S4; use the complementary text group of the local-global quality inference obtained in S3 to guide the refinement of the local-global features of the image aggregated in S5, thereby enhancing the model's perception and evaluation ability for local regions and forming a local-global quality inference logic.
[0014] Further, the specific process of obtaining the significant local regions of the image based on the open vocabulary detection model in S1 is as follows:
[0015] S1.1. Select the open vocabulary detection model YOLO-World to detect the objects in the image, where N represents the number of detected objects, ti Let the detected target be represented by \(t\), the image to be detected be \(I\), and the YOLO - World model be \(OD\). The formula is as follows:
[0016] \(\{t_1,t_2,t_3,\cdots,t\}\) N \(= OD(I)\);
[0017] S1.2. Starting from two aspects: the size of the target area and the distinction between foreground and background, screen the detected target areas. By calculating the intersection - over - union (IOU) between each target area ij , for the two target areas \(t\) i and \(t\) j with a relatively large IOU, determine the target area with the smaller area as the foreground \(t\) f . In this way, screen out the three most significant targets with the largest area in the image and obtain their spatial coordinates \(\{b_1,b_2,b_3\}\); where \(t\) i and \(t\) j represent any two targets detected by YOLO - World, \(min\) represents the operation of judging the area sizes of two targets and taking the target with the smaller area, \(\tau\) represents the set IOU threshold, which is set to 0.85 here, \(t\) fi represents the obtained foreground target, \(M\) represents the number of foreground targets screened out, \(sort\) represents the operation of sorting by area and taking the coordinates of the top three targets with the largest area. The specific process is as follows:
[0018]
[0019] Furthermore, the specific process of S2 is as follows:
[0020] S2.1. Set the prompt words ["Please describe the content within the <**significant region**> box and the visual quality, ignoring other boxes' edges inside."]. Guide the visual - language large - model to generate descriptive text from two aspects: quality and content. Among them, \(I\) represents the input image, \(P\) i represents the prompt words for the three significant regions corresponding to this figure, \(M1\) represents the visual - language large - model MiniCPM - Llama3 - V2.5, \(T\) i represents the descriptive text of the significant local region of the generated image. The formula is as follows:
[0021] \(\{T_1,T_2,T_3\}=M1(I,P\) i ) \(i\in(1,2,3)\);
[0022] S2.2. Set the prompt ["The image quality is categorized as bad, poor, fair, good, and excellent. Ignore the box edges, provide a brief description of the entire image, and analyze why the image looks <**quality label**> in relation to the area of the detected boxes. Let's think step by step."], and explicitly infer the association between the saliency local and global quality relationships of the image, where P inf represents the prompt for the local-global quality inference relationship of the design, and T inf represents the text of the local-global quality inference relationship. The formula is as follows:
[0023] T inf = M1(I, P inf ).
[0024] Furthermore, the specific process of S3 is as follows:
[0025] Set the prompt ["Generate the completely opposite text descriptions based on the following text: <**source text**>"], and use this to guide the large language model to generate a complementary text group; where M2 represents the Llama3-8B-Instruct large language model, and P opp represents the prompt for obtaining complementary texts, and respectively represent the complementary texts of the three saliency local description texts of the image and the complementary text of the image local-global quality inference relationship text. The formula is as follows:
[0026]
[0027] Furthermore, the specific process of S4 is as follows:
[0028] Adopt the visual feature encoder E of the visual language interaction model Long-CLIP v , and use the method of spatial mapping to obtain the saliency local features f1, f2, f3 and the global feature f of the image in the feature domain using the pixel domain coordinates b1, b2, b3. The specific process is as follows: I as follows:
[0029] {f1, f2, f3, f I} = E v (I, b1, b2, b3).
[0030] Further, the specific process of S5 is as follows:
[0031] Construct a local-global aggregation module, that is, aggregate local-global features from semantic and spatial aspects based on the transposed attention module TA and the multi-head attention module MH to obtain the local-global aggregation feature f lg , and the specific process is as follows:
[0032] f lg = MH(TA(f1, f2, f3, f I ))
[0033] Further, the specific process of S6 is as follows:
[0034] S6.1. Extract features from the complementary text group by means of the text encoder of Long-CLIP: Set quality prompt words ["high quality photo", "low quality photo"], and use the text feature encoder E of Long-CLIP t to extract features from the text information, where {f t1 , f t2 , f t3 , f tinf , f p} represents the source text features of each local, local-global quality inference relationship text, and quality prompt word, represents the complementary text features of each local, local-global quality inference relationship text, and quality prompt word, and the specific process is as follows:
[0035]
[0036] S6.2. Utilize the cross-modal interaction ability of the visual language interaction model Long-CLIP and adopt the complementary text group contrast learning paradigm to enhance the model's perception and extraction ability for local regions: Use the contrast learning method, that is, calculate the similarity S i between the local visual feature f ti and the local description source text feature f i , and use the loss function l1 to guide the model to approach the correct feature extraction method and move away from the wrong feature extraction method, so as to enhance the model's local feature perception and extraction ability. The specific process is as follows:
[0037]
[0038] S6.3. Leverage the cross-modal interaction ability of Long-CLIP and adopt a complementary text group contrastive learning paradigm to enhance the model to form a reasonable local-global quality inference relationship: Using the contrastive learning method, calculate the similarity S between the local-global aggregated feature f lg and the source text feature f of the local-global quality inference relationship tinf , and use the loss function l2 to guide the model to form the correct local-global quality inference logic. The specific process is as follows: inf
[0039]
[0040] S6.4. Leverage the cross-modal interaction ability of Long-CLIP, use the image quality label m to increase the strong supervision signal, calculate the similarity s between the image feature f I and the source text feature f of the quality prompt word, that is, the predicted quality score, use the loss function l3 to guide and enhance the correlation between the image feature and the quality score, and combine the obtained loss functions to form the final loss function L to jointly optimize the model; among them, {λ1, λ2, λ3} represents the proportion of each component loss function, which is set to 0.1, 0.1, 1 here. The specific process is as follows: p
[0041]
[0042] The present invention also provides an image local quality assessment system driven by a large model in a complex mine scene, including a camera, a memory card, a saliency target detection module, a local description and local-global relationship mining module, and a complementary text group quality perception enhancement module;
[0043] The camera and the memory card are used to collect and store the images underground in the mine;
[0044] The saliency target detection module uses an open vocabulary detection model to detect targets and perform saliency target screening;
[0045] The local description and local-global relationship mining module includes a vision-language large model and a large language model; its design is based on the prompt words of the chain of thought to guide the vision-language large model to obtain the image saliency local description text and the image local-global quality relationship inference text, and obtain the complementary text group through the large language model;
[0046] The described complementary text group quality perception enhancement module includes a local-global feature aggregation module and a vision-language interaction model Long-CLIP. The vision-language interaction model Long-CLIP consists of a vision feature encoder and a text feature encoder. It uses a spatial mapping method to obtain local image features, and then obtains local-global aggregated vision features through the local-global feature aggregation module. Utilizing the cross-modal interaction ability of Long-CLIP, it guides and enhances the model's perception and evaluation ability of the local image in the form of a complementary text group, and finally obtains the quality scores of the global image and any specified local position.
[0047] The present invention uses a chain of thought to guide the generation of text descriptions and enhances the model's perception and evaluation ability of local areas using multi-modal information. First, it uses a vision-language large model to obtain text of the significant local areas of the image, and obtains image quality inference text under the guidance of the chain of thought. Then, it designs a local-global feature aggregation mechanism to effectively extract and aggregate local and global features of the image. Finally, it constructs a complementary text group contrastive learning framework, and optimizes through cross-modal interaction to strengthen the model's perception sensitivity to the quality features of image areas in complex scenes, enhancing the evaluation ability of the blind image quality evaluation model for any specified area of the image, enabling it to have the evaluation ability for both the whole image and local areas of the image, and can form an effective evaluation of the overall mine image and any specified area, improving the safety factor of mine operations. Brief Description of the Drawings
[0048] Figure 1 is a schematic diagram of the working process of the method of the present invention;
[0049] Figure 2 is the overall system framework diagram in an embodiment of the present invention;
[0050] Figure 3 is a schematic diagram of obtaining the description text of the significant local area of the image and the image local-global quality inference text using the vision-language large model in an embodiment of the present invention;
[0051] Figure 4 is a schematic diagram of obtaining a complementary text group of the image using the large language model in an embodiment of the present invention;
[0052] Figure 5 is the structure diagram of the local-global aggregation module in an embodiment of the present invention;
[0053] Figure 6 is a schematic diagram of the evaluation of any local area of the image in an embodiment of the present invention;
[0054] Figure 7 is a performance comparison diagram of the ablation test on the FLIVE library in an embodiment of the present invention. Detailed Embodiments
[0055] The present invention will be further described below in conjunction with the accompanying drawings.
[0056] As Figure 1 shown, an image local quality assessment method driven by a large model in a complex mine scenario includes the following steps:
[0057] S1. Obtain the significant local regions of the image based on the open vocabulary detection model, and provide local visual semantic information for enhancing the local quality perception and evaluation ability of the model;
[0058] S2. Set the prompt words of the vision-language large model guided by the chain of thought, and obtain the descriptive text of the significant local regions of the image and the image local-global quality inference text with the help of the prompt words, including information on both the content and quality of the image;
[0059] S3. Obtain the complementary text group of the descriptive text of the significant local regions of the image and the local-global quality inference text based on the large language model, and provide more refined text modality supervision information for enhancing the local quality evaluation ability of the model and forming a good local-global quality inference logic;
[0060] S4. With the help of the visual feature encoder of the vision-language interaction model, adopt the method of spatial mapping to match the positions of the image pixel domain and the feature domain one by one, so as to extract the local and global features of the image;
[0061] S5. Construct a local-global feature aggregation module based on the transposed attention and multi-head attention mechanisms to aggregate the local and global features of the image obtained in S4;
[0062] S6. With the help of the cross-modal interaction ability of the vision-language interaction model, set the training paradigm of the model, and use the complementary text group related to the local description of the image obtained in S3 to guide the refinement of the local features of the image obtained in S4; use the complementary text group of the local-global quality inference obtained in S3 to guide the refinement of the local-global features of the image aggregated in S5, so as to enhance the model's perception and evaluation ability of the local region and form a local-global quality inference logic.
[0063] Embodiment: The algorithm framework of the present invention is as Figure 2 shown, including a camera, a memory card, a significant target detection module, a local description and local-global relationship mining module, and a complementary text group quality perception enhancement module;
[0064] The camera and the memory card are used to collect and store the images underground in the mine;
[0065] The significant target detection module uses the open vocabulary detection model to detect targets and perform significant target screening;
[0066] The described local description and local-global relationship mining module includes a vision-language large model and a large language model; it is designed to guide the vision-language large model based on chain-of-thought prompts to obtain text for local descriptions of image saliency and text for reasoning about the local-global quality relationship of the image, and obtain a complementary text group through the large language model;
[0067] The described complementary text group quality perception enhancement module includes a local-global feature aggregation module and a vision-language interaction model Long-CLIP, where the vision-language interaction model Long-CLIP consists of a visual feature encoder and a text feature encoder; it uses a spatial mapping method to obtain local image features, and then obtains local-global aggregated visual features through the local-global feature aggregation module. Using the cross-modal interaction ability of Long-CLIP, it guides and enhances the model's perception and evaluation ability for local images in the form of a complementary text group, and finally obtains the quality scores of the global image and any specified local position.
[0068] The specific process of local image quality assessment is as follows:
[0069] (1) Use an open-vocabulary detection model to detect target regions in the image, and filter out salient target regions through a designed algorithm;
[0070] (2) As Figure 3 shown, design prompts to guide the vision-language large model MiniCPM-Llama3-V2.5 to obtain text for describing the salient local regions of the image, and obtain text for reasoning about the local-global quality relationship of the image through prompts based on the chain-of-thought technique;
[0071] (3) As Figure 4 shown, use the large language model Llama3-8B-Instruct to design prompts to obtain a complementary text group for local descriptions and global quality reasoning;
[0072] (4) As Figure 5 shown, construct a local-global feature aggregation module based on transposed attention and multi-head attention mechanisms to aggregate the local and global features of the image obtained in the fourth step;
[0073] (5) Rely on the cross-modal interaction ability of the vision-language interaction model to guide and enhance the model's perception and evaluation ability for local regions in the form of a complementary text group;
[0074] The technical effects of the present invention are described in detail in combination with performance tests and experimental analyses. Existing blind image quality assessment algorithms usually focus on improving the global quality assessment ability, but lack the ability to assess the quality of local regions of images. The present invention aims to enhance the model's perceptual assessment ability for local regions of images. To prove the model's perceptual assessment ability for local regions of images, the present invention uses local images and their quality assessment labels in the real distortion dataset FLIVE for testing. And it is compared with existing algorithms, and the results are shown in Table 1:
[0075] Table 1
[0076]
[0077]
[0078] As can be seen from Table 1, compared with existing SOTA algorithms, the present invention has achieved the best local region assessment ability. At the same time, to more intuitively demonstrate the assessment ability of our algorithm for local regions, the present invention conducts a qualitative test. As Figure 6 shown, the assessment scores of the present invention for the distorted and undistorted parts of the image are in line with human subjective opinions;
[0079] While enhancing the model's perceptual assessment ability for local regions of images, the guided training mechanism provided by the present invention also has a feedback effect on the quality assessment of the overall image, showing excellent results. To prove the model's assessment ability for the overall image, the present invention uses the real distortion dataset KONIQ for testing and compares it with existing algorithms, and the results are shown in Table 2:
[0080] Table 2
[0081]
[0082] As can be seen from Table 2, compared with existing SOTA algorithms, the present invention has achieved the best global assessment ability.
[0083] To further prove the effectiveness of the present invention, the present invention conducts an ablation experiment. The present invention directly trains the prompt and the Long-CLIP model as the baseline model and compares it with the present invention, and the results are as Figure 7 shown. As can be seen from Figure 7 it, if the strategy of local description, local-global quality inference relationship positive and negative text cross-modal guidance used in the present invention is not adopted, the model's perceptual assessment ability for local regions of images will be significantly reduced, that is, the present invention significantly improves the perceptual assessment ability for local regions of images.
Claims
1. An image local quality assessment method driven by a large model in a complex mine scene, characterized in that, It includes the following steps: S1. Obtain the locally significant regions of the image based on the open-vocabulary detection model, providing local visual semantic information for enhancing the local quality perception evaluation ability of the model; S2. Set the prompt words of the vision-language large model guided by the chain of thought, and obtain the descriptive text of the locally significant regions of the image and the local-global quality inference text of the image with the help of the prompt words, including information on both the content and quality of the image; S3. Obtain the complementary text group of the descriptive text of the locally significant regions of the image and the local-global quality inference text based on the large language model, providing more refined text modality supervision information for enhancing the local quality evaluation ability of the model and forming a good local-global quality inference logic; S4. With the help of the visual feature encoder of the vision-language interaction model, use the method of spatial mapping to match the positions in the image pixel domain and the feature domain one by one, so as to extract the local and global features of the image; S5. Construct a local-global feature aggregation module based on the transposed attention and multi-head attention mechanisms to aggregate the local and global features of the image obtained in S4; S6. With the help of the cross-modal interaction ability of the vision-language interaction model, set the training paradigm of the model, and use the complementary text group related to the local description of the image obtained in S3 to guide the model to refine the local features of the image obtained in S4; use the complementary text group of the local-global quality inference obtained in S3 to guide the refinement of the local-global features of the image aggregated in S5, so as to enhance the model's perception and evaluation ability of the local region and form a local-global quality inference logic.
2. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, wherein The specific process of obtaining the locally significant regions of the image based on the open-vocabulary detection model in S1 is as follows: S1.
1. Select the open-vocabulary detection model YOLO-World to detect the objects in the image, where N represents the number of detected objects, t i represents the detected object, I represents the image to be detected, OD represents the YOLO-World model, and the formula is as follows: {t1, t2, t3... t N} = OD(I); S1.
2. Starting from the two aspects of the size and foreground / background distinction of the target region, screen the detected target regions by calculating the intersection over union (IOU) between each target region ij , and determine the target region with the smaller area among the two target regions t i and t j with a larger intersection over union as the foreground t f . In this way, screen out the three most significant targets with the largest area in the image and obtain their spatial coordinates {b1, b2, b3}; among them, t i and t j represent any two targets detected by YOLO-World, min represents the operation of judging the area sizes of the two targets and taking the target with the smaller area, τ represents the set intersection over union threshold, t fi represents the obtained foreground target, M represents the number of foreground targets screened out, sort represents the operation of sorting by area and taking the coordinates of the top three targets in terms of area size. The specific process is as follows:
3. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, wherein The specific process of S2 is as follows: S2.
1. Set the prompt ["Please describe the content within the <**salient region**> box and the visual quality, ignoring the edges of other boxes inside.”], and guide the visual language model to generate descriptive text from two aspects: quality and content. Here, I represents the input image, and P i represents the prompts for the three salient regions corresponding to the figure, M1 represents the visual language model MiniCPM-Llama3-V2.5, and T i represents the descriptive text of the generated salient local region of the image. The formula is as follows: {T1, T2, T3} = M1(I, P i ) i ∈ (1, 2, 3); S2.
2. Set the prompt ["The image quality is categorized as bad, poor, fair, good, and excellent. Ignore the box edges, provide a brief description of the entire image, and analyze why the image looks <**quality label**> in relation to the area of the detected boxes. Let's think step by step.”], and explicitly infer the association between the saliency of the image and the relationship between local and global quality, where P inf represents the prompt for the local-global quality inference relationship of the design, and T inf represents the text of the local-global quality inference relationship. The formula is as follows: T inf = M1(I, P inf )。 4. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, characterized in that, The specific process of S3 is as follows: Set the prompt ["Generate the completely opposite text descriptions based on the following text: <**source text**>"] to guide the large language model to generate a complementary text group; where M2 represents the Llama3-8B-Instruct large language model, and P opp represents the prompt for obtaining complementary text, and respectively represent the complementary texts of the three significant local description texts of the image and the complementary text of the local-global quality inference relationship text of the image. The formula is as follows:
5. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, wherein The specific process of S4 is as follows: Visual feature encoder E of the visual language interaction model Long-CLIP v , using the method of spatial mapping, obtaining the significant local features f1, f2, f3 and global feature f of the image in the feature domain by using the pixel domain coordinates b1, b2, b3 I , the specific process is as follows: {f1,f2,f3,f I} = E v (I,b1,b2,b3).
6. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, wherein The specific process of S5 is as follows: Construct a local-global aggregation module, that is, aggregate local-global features from semantic and spatial aspects based on the transposed attention module TA and the multi-head attention module MH to obtain the local-global aggregation feature f lg , and the specific process is as follows: f lg = MH(TA(f1, f2, f3, f I ))。 7. The method for local image quality assessment driven by a large model in a complex mine scenario according to claim 1, wherein The specific process of S6 is as follows: S6.
1. Feature extraction of the complementary text group using the text encoder of Long-CLIP: Set the quality prompt words ["high quality photo", "low quality photo"], and use the text feature encoder E of Long-CLIP t to extract features from the text information, where {f t1 , f t2 , f t3 , f tinf , f p} represents the source text features of each local, local-global quality inference relationship text, and quality prompt word, and represents the complementary text features of each local, local-global quality inference relationship text, and quality prompt word. The specific process is as follows: S6.
2. Leverage the cross-modal interaction ability of the visual language interaction model Long-CLIP and adopt a complementary text group contrastive learning paradigm to enhance the model's perception and extraction ability for local regions: Use the contrastive learning method, that is, calculate the similarity S i between the local visual feature f ti and the local description source text feature f i , and use the loss function l1 to guide the model to approach the correct feature extraction method and move away from the wrong feature extraction method, so as to enhance the model's local feature perception and extraction ability. The specific process is as follows: S6.
3. Leverage the cross-modal interaction ability of Long-CLIP and adopt a complementary text group contrastive learning paradigm to enhance the model to form a reasonable local-global quality inference relationship: Using the method of contrastive learning, calculate the similarity S between the local-global aggregated feature f lg and the source text feature f of the local-global quality inference relationship tinf and use the loss function l2 to guide the model to form the correct local-global quality inference logic. The specific process is as follows: inf S6.
4. Utilize the cross-modal interaction ability of Long-CLIP, use the image quality label m to increase the strong supervision signal, and calculate the image feature f I and the source text feature f of the quality prompt p between the similarity s, that is, the predicted quality score, use the loss function l3 to guide and enhance the correlation between the image feature and the quality score, and combine the obtained loss functions to form the final loss function L to jointly optimize the model; among them, {λ1, λ2, λ3} represents the proportion of each component loss function, and the specific process is as follows:
8. An image local quality assessment system driven by a large model in a complex mine scenario, comprising a camera and a memory card, characterized in that, It also includes a significant object detection module, a local description and local-global relationship mining module, and a complementary text group quality perception enhancement module; The camera and memory card are used to store the images underground in the mine; The significant object detection module uses the open-vocabulary detection model to detect objects and perform significant object screening; The local description and local-global relationship mining module includes a vision-language large model and a large language model; it designs prompt words based on the chain of thought to guide the vision-language large model to obtain the descriptive text of the locally significant regions of the image and the local-global quality relationship inference text of the image, and obtains the complementary text group through the large language model; The complementary text group quality perception enhancement module includes a local-global feature aggregation module and the vision-language interaction model Long-CLIP, where the vision-language interaction model Long-CLIP consists of a visual feature encoder and a text feature encoder; it uses the method of spatial mapping to obtain the local features of the image, and then obtains the local-global aggregated visual features through the local-global feature aggregation module, and uses the cross-modal interaction ability of Long-CLIP to guide and enhance the model's perception and evaluation ability of the local part of the image in the form of a complementary text group, and finally obtains the quality scores of the global image and any specified local position.
Citation Information
Patent Citations
Knowledge-enhanced user multi-modal online comment quality evaluation method and system
CN116881689A
No-reference image quality evaluation method based on image features and semantic description
CN118608467A
Cited By
Intelligent document segmentation method based on Longform and Long-LIP models
CN121072531A
Document intelligent segmentation method based on longformer and long-clip model
CN121072531B