Method and system for evaluating the quality of image content inference enhancement in a mine restricted environment

By acquiring textual descriptions of mine image content through a large visual language model and utilizing zero-shot thought chain technology to deeply explore image quality relationships, text labels are generated to guide model training. This solves the problem of insufficient generalization of pre-trained models in the mine environment and improves the performance of the quality evaluation model.

CN120526429BActive Publication Date: 2026-05-01CHINA UNIV OF MINING & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2025-05-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing pre-trained models have poor generalization ability in image quality assessment in mining environments, making it difficult to effectively capture complex factors. Furthermore, they are prone to overfitting in scenarios with few labels, resulting in insufficient model prediction accuracy.

Method used

We employ a large visual language model to obtain unsupervised text descriptions of image content. We leverage the strong reasoning ability of the large model and zero-shot thought chain technology to deeply explore the relationship between image content and quality, generate text labels to guide model training, and enhance quality evaluation through the Long-CLIP model.

Benefits of technology

In the case of small samples, the model's understanding of the quality factors of mine environment images was improved, and the generalization ability and prediction accuracy of the quality assessment model were enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526429B_ABST
    Figure CN120526429B_ABST
Patent Text Reader

Abstract

The application discloses a kind of mine limited environment image content inference enhanced quality evaluation method and system, first with the aid of visual language big model, unsupervisedly obtain the mine limited environment image content text description including key object, scene, attribute etc. in image, then the powerful reasoning ability of language big model and zero sample thinking chain technology is used to in-depth mining the relationship between image content and quality, after the reasoning relationship between image content and quality is condensed by language big model, text label for guiding model training is generated, and it is used as guide information to guide model to further mine the factors affecting image quality, finally only a small amount of samples are used to complete the final training of quality evaluation model.The application can make the model more effectively learn and in-depth mine the complex factors affecting image quality on the basis of small sample, and then improve the generalization ability and practicability of quality evaluation model, especially suitable for image quality evaluation work in mine environment.
Need to check novelty before this filing date? Find Prior Art

Description

A Quality Assessment Method and System for Image Content Inference Enhancement in Confined Mining Environments Technical Field

[0001] This invention relates to an image quality assessment method and system, specifically a quality assessment method and system for image content inference enhancement in a confined environment of a mine, belonging to the field of image quality assessment technology. Background Technology

[0002] With the rapid progress of mine informatization, the "unmanned mining" model is steadily advancing. More and more image acquisition and video processing technologies are being integrated into various aspects of mining operations, aiming to achieve efficient visual monitoring and detection functions. However, the complex mine environment, characterized by significant electromagnetic interference and poor imaging lighting conditions, results in high noise and low illumination in acquired images. This low-light, noisy mine images severely restrict the widespread application of image and video technologies in mining. The importance of mine image quality assessment technology is increasingly prominent. This technology aims to comprehensively and accurately assess the quality of image data acquired in the mine environment through scientific algorithms and models, thereby guiding subsequent image processing techniques such as image enhancement, noise suppression, and target detection, improving the usability of image information and decision support capabilities. Therefore, in-depth research and development of image quality assessment technologies suitable for the mine environment is of great significance for promoting the intelligent and safe development of mining.

[0003] Image quality assessment, which mimics the human visual system, can be broadly categorized into subjective and objective assessments. Subjective assessment relies on the observer's visual perception and judgment; the observer scores the image based on their own visual experience, and the score is typically used as a reference score for image quality assessment. Objective assessment, on the other hand, uses mathematical models and algorithms to automatically evaluate image quality without human intervention. Based on the varying degrees of reference information provided by the source image, image quality assessment methods can be further subdivided into three main categories: Full-Reference (FR), Reduced-Reference (RR), and No-Reference (NR). FR assessment relies on a complete and high-quality reference image as a benchmark for comprehensive comparative analysis of the test image; RR assessment, while utilizing some form of reference information, does not provide as detailed and complete information as under full-reference conditions; NR assessment, however, does not rely on any pre-defined reference image or corresponding detailed reference information, relying solely on the image's inherent characteristics for independent quality assessment.

[0004] Currently, mainstream quality assessment algorithms typically use quality scores to fine-tune pre-trained models to adapt them to quality assessment tasks. Pre-training is a deep learning model training strategy that utilizes large-scale datasets to initially train the model, enabling it to learn general feature representations. This process is similar to the basic learning stage humans undergo before learning new knowledge, accumulating experience through extensive reading and observation. Pre-trained models can capture universal features in the data, which usually have good transferability to subsequent specific tasks. Therefore, pre-trained models become the starting point for many deep learning applications, significantly improving model performance on new tasks. Fine-tuning refers to further training the model on a task-specific dataset based on the pre-trained model to adjust the model parameters and better adapt it to the target task. While pre-trained models can significantly improve model performance on specific tasks, especially with limited training samples, achieving efficient fine-tuning through transfer learning, they also have some inherent limitations. Specifically, on the one hand, existing pre-trained models often rely on large-scale datasets for initial training. The distribution of these datasets may differ from the data distribution of the actual task, leading to poor generalization when the model is transferred to a new task. When fine-tuning the pre-trained model using quality scores, the model often struggles to fully learn and uncover the complex factors affecting image quality. Real distorted images often exhibit multi-dimensional distortion and a tight coupling between image content and distortion. When the model relies solely on the single metric of quality score for fine-tuning, it often fails to effectively capture and utilize these distortion characteristics. How to enable the model to fully understand the essential features of image quality during the learning process is a pressing issue that current model algorithms need to address. On the other hand, when using fewer labeled samples for fine-tuning, the model is prone to overfitting to the task without learning the characteristics of image distortion, making it difficult to guarantee the model's prediction accuracy. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a quality assessment method and system for image content reasoning enhancement in confined environments such as mines. This method enables the model to learn more effectively and delve deeper into the complex factors affecting image quality based on small samples, thereby promoting a deeper understanding of the essential characteristics of image quality during the training process. This enhances the generalization ability and practicality of the quality assessment model, making it particularly suitable for image quality assessment in mining environments.

[0006] To achieve the above objectives, the quality assessment method for image content inference enhancement in the confined environment of this mine specifically includes the following steps:

[0007] Step 1: Obtain detailed descriptions of image content based on a large visual language model: Using an unsupervised approach, textual descriptions of image content related to the confined environment of a mine are obtained with the help of a large visual language model.

[0008] Step 2, delve deeper into the relationship between image content and quality: leverage the powerful reasoning capabilities of large models and zero-shot thought chain technology to delve deeper into the relationship between image content and quality;

[0009] Step 3: Use the large language model to refine the inference relationship between image content and quality, and generate text labels to guide model training;

[0010] Step 4: Use the inference relationship text labels as guiding information to guide the model in mining the factors affecting image quality;

[0011] Step 5: First, calculate the final loss l for quality assessment. Then, while adapting to the quality assessment task, enhance the model's understanding of image quality factors in the confined environment of the mine by using the guidance of content and quality correlation. The final training of the quality assessment model is completed using only a small number of samples.

[0012] Furthermore, Step 1 is detailed as follows:

[0013] Step 1-1, Select the visual language large model: Select ShareGPT4V-Captioner as the visual language large model;

[0014] Step 1-2, Design prompts to generate image content descriptions: After inputting the image to be described, design a prompt: ["Describe the picture in detail"]. The visual language model provides a detailed text description, represented as:

[0015] t1 = M1(I1, p1)

[0016] In the formula: t1 represents the output content text description; I1 represents the input image; p1 represents the prompt text hint designed to guide the model to generate the descriptive text; M1 represents the ShareGPT4V visual language big model, which accepts image I1 and prompt p1 as input and outputs the content text description t1.

[0017] Furthermore, Step 2 is detailed below:

[0018] Step 2-1, Select the large language model: Select MiniCPM-Llama3-V as the large language model;

[0019] Step 2-2, Design prompts: ["Here is the content description of the image: <**content text description**>", "The image quality can be divided into five levels of 'Bad-Poor-Fair-Good-Excellent' from low to high. Why the quality of this image is <**quality label**>", "Please think step by step in combination with the image content."] The language model provides the textual relationship between image content and quality, represented as follows:

[0020] t2 = M2(t1, p2)

[0021] In the formula: t2 represents the text describing the relationship between the output image content and quality; p2 represents the prompt text designed to predict the relationship between image content and quality; M2 represents the MiniCPM-Llama3-V language big model, which takes the content text description t1 and prompt p2 obtained in Step1 as input and outputs the text describing the relationship between image content and quality t2.

[0022] Furthermore, Step 3 is detailed below:

[0023] Step 3-1: Select a large language model to refine the inference relationship between image content and quality: Select Llama3-8B-Instruct as the large language model to generate positive and negative sample pairs;

[0024] Step 3-2: Design prompts that predict the relationship between image content and quality in Step 2, thereby generating positive and negative text label pairs to guide model training. One set of prompts is designed: ["Condense three main descriptions from the following text, then generate three completely opposite descriptions: <**Text on the relationship between image content and quality**>"]. The language model outputs positive and negative text label pairs representing the relationship between image content and quality, as follows:

[0025] [t3,t4,t5,t6,t7,t8]=M3(t2,p3)

[0026] In the formula: [t3,t4,t5,t6,t7,t8] represents the positive and negative sample pairs of text labels relating to the relationship between image content and quality, where t3,t5,t7 indicate that the prompt belongs to positive samples, and t4,t6,t8 indicate that the prompt belongs to negative samples; p3 represents the designed prompt, which is the text prompt used to predict the positive and negative sample pairs of text labels; M3 represents the language big model Llama3-8B-Instruct, which takes the text t2 and prompt p3 relating to the relationship between image content and quality obtained in Step 2 as input, and outputs the positive and negative sample pairs of the relationship between image content and quality [t3,t4,t5,t6,t7,t8].

[0027] Furthermore, Step 4 is detailed below:

[0028] After obtaining the positive and negative sample pairs [t3,t4,t5,t6,t7,t8] relating image content and quality, the Long-CLIP model is used to compare the positive and negative sample pairs to obtain the content inference loss l1. This loss l1 is used as guiding information to instruct the model to discover factors affecting image quality, and is expressed as:

[0029] l1=∑M4([t3,t4,t5,t6,t7,t8])

[0030] In the formula: l1 represents the content inference loss; M4 represents the Long-CLIP model.

[0031] Furthermore, in Step 5, when calculating the final loss l for quality assessment, the original loss l0 is first obtained by directly performing quality assessment using the Long-CLIP model. Then, the original loss l0 is added to the content inference loss l1 obtained in Step 4 to obtain the final loss l, as detailed below:

[0032] The text encoder of the CLIP model directly predicts image quality and quality attributes by inputting a designed prompt into the text. The prompt is designed to contain both positive and negative descriptive text: positive prompt... 1 {high}quality photo and negative p 0 The low-quality photo, the design prompt is denoted as p. c Where c = 0, 1, representing that the prompt belongs to class c, and the CLIP text encoder and image encoder are denoted as E, respectively. t and E i Then CLIP's prediction score for image I is expressed as:

[0033]

[0034] In the formula: tc For text features; E t For CLIP's text encoder; p c The prompt is for the design; f represents the image features; E represents the design prompt. i For CLIP image encoder; m c Cosine similarity; I represents the image; s represents the prediction score of image I in the prompt for the design; m 1 For p 1 cosine similarity; m 0 For p 0 Cosine similarity; For m 1 The exponential function; e m0 For m 0 The exponential function;

[0035] The original loss l0 is obtained by directly performing quality assessment using the Long-CLIP model, and is expressed as:

[0036] l0 = M4(s,s0)

[0037] In the formula: l0 represents the original loss; M4 represents the Long-CLIP model; s represents the predicted score of image I obtained in the prompt for the design; s0 represents the true score of image I;

[0038] The original loss l0 is added to the content reasoning loss l1 obtained in Step 4 to obtain the final loss l, which is expressed as:

[0039] l=l0+λl1

[0040] In the formula: λ is the weight parameter of the content inference loss l1.

[0041] A quality assessment system for image content inference enhancement in confined mine environments, based on an image content description module, a content inference enhancement module, and an image quality assessment module, is provided.

[0042] Compared with existing technologies, considering that mainstream algorithms may overfit to quality assessment tasks when fine-tuning pre-trained models, the model's ability to understand content may decrease, especially in scenarios with few labels, this quality assessment method for image content reasoning enhancement in confined mine environments adopts a content reasoning enhancement strategy. First, it uses a large visual language model to unsupervisedly acquire text descriptions of image content in confined mine environments, covering key objects, scenes, and attributes in the images. Then, it leverages the powerful reasoning capabilities of the large language model and zero-shot thought chain technology to deeply explore the relationship between image content and quality. After further refining the reasoning relationship between image content and quality using the large language model, it generates text labels to guide model training and uses these labels as guiding information to further explore factors affecting image quality. Finally, it completes the final training of the quality assessment model using only a small number of samples. This quality assessment method, which enhances image content reasoning in confined mine environments, leverages the correlation between image content and quality to improve the model's understanding of image quality factors in confined mine environments while adapting to the quality assessment task. It enables the model to learn more effectively and delve deeper into the complex factors affecting image quality on a small sample basis, thereby promoting a deeper understanding of the essential characteristics of image quality during training. This enhances the generalization ability and practicality of the quality assessment model, making it particularly suitable for image quality assessment in mine environments. Attached Figure Description

[0043] Figure 1 is a flowchart of the present invention;

[0044] Figure 2 is a flowchart of the algorithm of the present invention;

[0045] Figure 3 is a diagram showing the ablation experiment results on different datasets provided in the embodiments of the present invention. Detailed Implementation

[0046] This quality assessment method, which enhances image content reasoning in confined mine environments, firstly uses a large visual language model to unsupervisedly acquire textual descriptions of the image content, covering key objects, scenes, and attributes within the images. Then, it leverages the powerful reasoning capabilities of the large model and zero-shot thought chain techniques to deeply explore the relationship between image content and quality. Next, it uses the large language model to condense the reasoning relationships between image content and quality, generating textual labels for these relationships to guide model training. These textual labels then serve as guiding information to further explore factors influencing image quality. Finally, the quality assessment model is trained using only a small number of samples. By leveraging the association between image content and quality, the method enhances the model's understanding of image quality factors in confined mine environments while adapting to the quality assessment task, effectively improving the model's predictive ability.

[0047] As shown in Figure 1, the quality assessment method for image content inference enhancement in the confined environment of this mine specifically includes the following steps:

[0048] Step 1: Obtain detailed descriptions of image content based on a large visual language model: Using an unsupervised approach, textual descriptions of the confined environment images in a mine are obtained using a large visual language model. These textual descriptions cover key objects, scenes, and attributes within the images. Details are as follows:

[0049] Step 1-1: Selecting a Large-Scale Visual Language Model. This invention selects ShareGPT4V-Captioner as the large-scale visual language model. ShareGPT4V has 100,000 high-quality image text labels collected from the advanced GPT4-Vision dataset. By labeling the model on this subset, it is further expanded to 1.2 million highly descriptive image description text labels, surpassing the diversity and information content of other existing datasets. It covers world knowledge, object attributes, spatial relationships, and aesthetic evaluation, and can generate more accurate and vivid image content text descriptions. It can focus on areas that are easily overlooked by the human eye. At the same time, ShareGPT4V also has cross-modal learning capabilities, and can process multiple data types such as images and text simultaneously, further improving the quality and diversity of generated text.

[0050] Step 1-2: Design prompts to generate image content descriptions. After inputting the image to be described, design a prompt: ["Describe the picture in detail"]. The model will then provide a detailed text description, which can be represented as:

[0051] t1 = M1(I1, p1)

[0052] In the formula: t1 represents the output content text description; I1 represents the input image; p1 represents the designed prompt, which is a text prompt used to guide the model to generate descriptive text; M1 represents the ShareGPT4V visual language big model, which accepts image I1 and prompt p1 as input and outputs content text description t1.

[0053] Step 2: Delve deeper into the relationship between image content and quality: Leveraging the powerful reasoning capabilities of large models and zero-shot thought chain techniques, delve into the relationship between image content and quality. Specifically:

[0054] Step 2-1: Select the large language model for text reasoning. This invention selects MiniCPM-Llama3-V 2.5 as the large language model. MiniCPM-Llama3-V 2.5 is the latest version of the MiniCPM-V series model, built based on SigLip-400M and Llama3-8B-Instruct, with a total of 8B parameters, achieving a significant performance improvement compared to existing algorithms. The features of MiniCPM-Llama3-V 2.5 include: ① Leading performance: MiniCPM-Llama3-V 2.5 achieved an average score of 65.1 on the OpenCompass leaderboard, which integrates 11 mainstream multimodal large model evaluation benchmarks. With a size of 8 bytes, it surpasses mainstream commercial closed-source multimodal large models such as GPT-4V-1106, Gemini Pro, Claude 3, and Qwen-VL-Max, and significantly outperforms other multimodal large models built on Llama 3; ② Excellent text recognition capabilities: MiniCPM-Llama3-V 2.5 can accept image input with any aspect ratio up to 1.8 megapixels, achieving an OCR Bench score of 725, surpassing commercial closed-source models such as GPT-4o, GPT-4V, Gemini Pro, and Qwen-VL-Max, reaching the best performance. Based on recent user feedback, MiniCPM-Llama3-V... Version 2.5 enhances high-frequency practical capabilities such as full-text text recognition information extraction and table image to Markdown conversion, and further strengthens instruction following and complex reasoning capabilities, which can bring a better multimodal interactive experience; ③ Trustworthy multimodal behavior: With the help of the latest RLAIF-V alignment technology, MiniCPM-Llama3-V 2.5 has more trustworthy multimodal behavior, and the illusion rate in Object HalBench has been reduced to 10.3%, which is significantly lower than GPT-4V-1106 (13.6%), reaching the best level in the open source community.

[0055] Step 2-2 utilizes the zero-shot Chain of Thought (CoT) technique in large models to explore the relationship between image content and quality. Chain of Thought is an improved Prompt technique that introduces intermediate reasoning steps into large language models, aiming to enhance the model's performance on complex reasoning tasks. This technique significantly enhances the model's arithmetic, common sense, and reasoning abilities by requiring the model to explicitly output the step-by-step reasoning process before outputting the final answer. CoT originated from a 2022 Google paper, its core idea being to decompose complex problems into a series of sub-problems and solve them step-by-step, forming a series of intermediate reasoning steps. When applying CoT, a complete Prompt typically includes three parts: instructions, logical basis, and examples. Instructions describe the problem and inform the model of the output format; the logical basis is the intermediate reasoning process of CoT, including the solution to the problem, intermediate steps, and external knowledge related to the problem; examples provide the basic format of input-output pairs in a few-shot manner, helping the model understand how to generate the reasoning process. Depending on whether examples are included, CoT can be divided into Zero-Shot-CoT and Few-Shot-CoT types. Zero-Shot-CoT only prompts the model to generate reasoning chains through specific cue text, while Few-Shot-CoT guides the model's reasoning by describing the solution steps in detail. The introduction of CoT technology can significantly improve the performance of large language models on complex reasoning tasks. Its principle lies in building bridges between the fundamental knowledge understood by the large model, allowing known information to form a coherent chain and preventing the model from deviating during the solution process. At the same time, CoT technology can also enhance the model's interpretability, credibility, controllability, flexibility, and creativity.

[0056] Design prompts that can predict the relationship between image content and quality (such as key objects, scenes, attributes, etc. in the image). Design prompts like: ["Here is the content description of the image: <**content text description**>", "The image quality can be divided into five levels of 'Bad-Poor-Fair-Good-Excellent' from low to high. Why the quality of this image is <**quality label**>", "Please think step by step in combination with the image content."]. The model can then output the text describing the relationship between image content and quality, which can be represented as:

[0057] t2 = M2(t1, p2)

[0058] In the formula: t2 represents the text describing the relationship between the image content and quality; p2 represents the designed prompt, which is a textual hint used to predict the relationship between the image content and quality (such as key objects, scenes, attributes, etc. in the image); M2 represents the MiniCPM-Llama3-V 2.5 language big model, which takes the content text description t1 and prompt p2 obtained in Step1 as input and outputs the text describing the relationship between the image content and quality t2.

[0059] Step 3: Utilize the large language model to refine the inference relationships between image content and quality, generating text labels to guide model training. Specifically:

[0060] Step 3-1: Select a large-scale language model to refine the inference relationship between image content and quality. This invention selects Llama3-8B-Instruct as the large-scale language model to generate positive and negative sample pairs. The Llama3-8B-Instruct model is a large-scale language model based on the Transformer architecture, developed by MetaWorks. It has 8 billion parameters, is pre-trained on a large-scale corpus, and possesses powerful natural language processing capabilities, enabling it to efficiently and accurately complete diverse tasks such as dialogue generation, text creation, and programming assistance.

[0061] Step 3-2: Design prompts that can predict the relationship between image content and quality in Step 2 and generate text label pairs to guide model training. One set of prompts is designed: ["Condense three main descriptions from the following text, then generate three completely opposite descriptions: <**Text on the relationship between image content and quality**>"]. The model can then output text label pairs representing the relationship between image content and quality, which can be represented as:

[0062] [t3,t4,t5,t6,t7,t8]=M3(t2,p3)

[0063] In the formula: [t3,t4,t5,t6,t7,t8] represents the positive and negative sample pairs of text labels relating to the relationship between image content and quality, where t3,t5,t7 indicate that the prompt belongs to positive samples, and t4,t6,t8 indicate that the prompt belongs to negative samples; p3 represents the designed prompt, which is the text prompt used to predict the positive and negative sample pairs of text labels; M3 represents the language big model Llama3-8B-Instruct, which takes the text t2 and prompt p3 relating to the relationship between image content and quality obtained in Step 2 as input, and outputs the positive and negative sample pairs of the relationship between image content and quality [t3,t4,t5,t6,t7,t8].

[0064] Step 4: Use the inference relationship text labels as guiding information to instruct the model to discover factors affecting image quality. Specifically:

[0065] After obtaining the positive and negative sample pairs [t3,t4,t5,t6,t7,t8] relating image content and quality, the Long-CLIP model is used to compare the positive and negative sample pairs to obtain the content inference loss l. 14 Using this as guiding information to instruct the model to discover factors affecting image quality can be expressed as:

[0066] l1=∑M4([t3,t4,t5,t6,t7,t8])

[0067] In the formula: l1 represents the content inference loss; M4 represents the Long-CLIP model.

[0068] Step 5: First, calculate the final loss l for quality assessment. Then, while adapting to the quality assessment task, enhance the model's understanding of image quality factors in the confined environment of the mine by using the guidance of content and quality correlation. The final training of the quality assessment model is completed using only a small number of samples.

[0069] To calculate the final loss *l* obtained from the quality assessment, the original loss *l0* is first obtained by directly performing the quality assessment using the Long-CLIP model. Then, the original loss *l0* is added to the content inference loss *l1* obtained in Step 4 to obtain the final loss *l*. Specifically:

[0070] Long-CLIP possesses strong zero-shot prediction capabilities, allowing direct prediction of image quality and quality attributes through a text encoder that inputs a designed prompt into the CLIP model. The prompt is designed with both positive and negative descriptive text: positive p 1 {high}quality photo and negative p 0 {low}quality photo. The design prompt is denoted as p. cWhere c = 0, 1, representing that prompt belongs to class c. The CLIP text encoder and image encoder are denoted as E. t and E i Then CLIP's prediction score for image I is expressed as:

[0071]

[0072] In the formula: t c For text features; E t For CLIP's text encoder; p c The prompt is for the design; f represents the image features; E represents the design prompt. i For CLIP image encoder; m c Cosine similarity; I represents the image; s represents the prediction score of image I in the prompt for the design; m 1 For p 1 cosine similarity; m 0 For p 0 Cosine similarity; For m 1 The exponential function; For m 0 The exponential function.

[0073] The original loss l0 obtained by directly performing quality assessment using the Long-CLIP model can be expressed as follows:

[0074] l0 = M4(s,s0)

[0075] In the formula: l0 represents the original loss; M4 represents the Long-CLIP model; s represents the predicted score of image I obtained in the prompt for the design; s0 represents the true score of image I.

[0076] Adding the original loss l0 to the content inference loss l1 obtained in Step 4 yields the final loss l, which can be expressed as follows:

[0077] l=l0+λl1

[0078] In the formula: λ is the weight parameter of the content inference loss l1.

[0079] Figure 2 shows the system framework of the image content inference enhancement quality assessment method for confined environments in mines. The system consists of three modules: image content description module, content inference enhancement module, and image quality assessment module. The ShareGPT 4v visual language model is used to unsupervisedly acquire textual descriptions of the confined environment images in the mine. These textual descriptions cover key objects, scenes, and attributes in the images. After obtaining the textual descriptions, the powerful inference capabilities and zero-shot reasoning techniques of the MiniCPM-Llama3-V 2.5 model are first utilized to deeply explore the relationship between image content and quality. Then, the Llama3-8B-Instruct language model is used to generate positive and negative sample pairs to further refine the inference relationship between image content and quality, generating text labels to guide model training. After obtaining the inference relationship labels, they are used as guiding information to guide the model to further explore factors affecting image quality. Finally, the Long-CLIP quality assessment model is trained using only a small number of samples. This invention enhances the model's understanding of image quality factors in the confined environment of mines by leveraging the correlation between image content and quality, thereby effectively improving the model's predictive ability while adapting to quality assessment tasks.

[0080] The technical effects of this invention will be described in detail below, based on performance testing and experimental analysis. Existing no-reference image quality assessment algorithms typically use 80% of the images to train the model. This invention aims to reduce the amount of training data required. To demonstrate the performance of this model with fewer labels, we selected 5% and 10% of the images from the real-world distortion dataset KonIQ-10k for training. The results are compared with existing algorithms such as the dubbed blind / referenceless imagespatial quality evaluator (BRISQUE), the unsupervised feature learning framework for no-reference image quality assessment (CORNIA), the blind image quality assessment in the wild guided by a self-adaptive hypernetwork (HyperNet), the blind image quality assessment using a deep bilinear convolutional neural network (DBCNN), and the active learning-based sample selection for label-efficient blind image quality assessment (AL-IQA). The results are shown in the table below.

[0081]

[0082]

[0083] As can be seen from the table above, compared with the existing best-performing model algorithms, the algorithm of this invention achieves the best few-label prediction capability.

[0084] To further verify the effectiveness of the quality assessment method for image content inference enhancement in the confined environment of this mine, an ablation experiment was conducted on KonIQ-10k using the algorithm of this invention. Following the general algorithm approach, images were randomly selected, and the same training method and network training model as the algorithm of this invention were used as baseline models for comparison. All models were trained using 5% of the KonIQ-10k images. The comparison results are shown in Figure 3. As can be seen from Figure 3, if the content inference enhancement strategy used in this invention is not employed, and instead a commonly used quality score is used for training, the model's predictive ability will be significantly reduced. Experimental verification shows that, in scenarios with a significantly reduced number of image training samples, the algorithm of this invention can significantly improve the model's predictive ability.

[0085] This quality assessment method for images in confined mine environments employs a content-based reasoning enhancement strategy. By leveraging the correlation between image content and quality, it enhances the model's understanding of image quality factors in confined mine environments while adapting to the quality assessment task. This approach enables the model to learn more effectively and delve deeper into the complex factors affecting image quality on a small sample basis. Consequently, it promotes a deeper understanding of the essential characteristics of image quality during training, thereby improving the generalization ability and practicality of the quality assessment model. This method is particularly suitable for image quality assessment in mine environments.

Claims

1. A quality assessment method for image content inference enhancement in a confined environment of a mine, characterized in that, Specifically, the following steps are included: Step 1: Obtain detailed image content description based on a visual language model: An unsupervised approach is used to obtain textual descriptions of images depicting a confined environment in a mine using a visual language model. Specifically: Step 1-1: Select a visual language model: ShareGPT4V-Captioner is selected as the visual language model. Step 1-2: Design prompts to generate image content descriptions: After inputting the image to be described, a prompt is designed: ["Describe the picture in detail"]. The visual language model provides a detailed textual description, expressed as: = ( In the formula: This represents a text description of the output content; This represents the input image; This indicates a prompt text tooltip designed to guide the model in generating descriptive text. This represents the ShareGPT4V visual language big model, which accepts images. and prompt As input, output a text description of the content. Step 2, In-depth exploration of the relationship between image content and quality: Utilizing the powerful reasoning capabilities of the large model and zero-shot thought chain technology, we delve into the relationship between image content and quality; specifically as follows: Step 2-1, Selecting the language large model: We select MiniCPM-Llama3-V as the language large model; Step 2-2, Designing prompts: ["Here is the content description of the image: <**content text description**>", "The image quality can be divided into five levels of 'Bad-Poor-Fair-Good-Excellent' from low to high. Why the quality of this image is <**quality label**>", "Please think step by step in combination with the image content."] The language large model provides the textual representation of the relationship between image content and quality, expressed as: = ( In the formula: Text indicating the relationship between the output image content and quality; The prompt text indicates a design feature for predicting the relationship between image content and quality. This refers to the MiniCPM-Llama3-V language model, which accepts the text description obtained in Step 1. and prompt As input, output text relating image content and quality. Step 3: Utilize a large language model to refine the inference relationship between image content and quality, generating text labels to guide model training; specifically as follows: Step 3-1: Select a large language model for refining the inference relationship between image content and quality: Select Llama3-8B-Instruct as the large language model to generate positive and negative sample pairs; Step 3-2: Design prompts that can predict the relationship between image content and quality in Step 2 and generate positive and negative text label sample pairs to guide model training. One set of prompts is designed: ["Condense three main descriptions from the following text, then generate three completely opposite descriptions: <**Text on the relationship between image content and quality**>"]. The large language model outputs positive and negative text label sample pairs representing the relationship between image content and quality, as follows: [ ]= ( In the formula: [ ] represents the positive and negative sample pairs of text labels indicating the relationship between image content and quality, where This indicates that prompt belongs to a positive sample. This indicates that prompt belongs to a negative sample; The prompt represents the design, which is a textual hint used to predict positive and negative sample pairs of text labels; The Llama3-8B-Instruct represents a large language model that accepts text representing the relationship between image content and quality obtained in Step 2. and prompt As input, output positive and negative sample pairs relating image content and quality. Step 4: Use the inference relationship text labels as guiding information to guide the model in mining the factors affecting image quality; Step 5: First, calculate the final loss for obtaining the quality assessment. Then, while adapting to the quality assessment task, the model's understanding of the quality factors of images in the confined environment of the mine is enhanced by the guidance of the correlation between content and quality, and the final training of the quality assessment model is completed using only a small number of samples.

2. The quality assessment method for image content inference enhancement in a confined environment of a mine according to claim 1, characterized in that, Step 4 is as follows: Obtain positive and negative sample pairs representing the relationship between image content and quality. Then, the content inference loss is obtained by comparing positive and negative sample pairs using the Long-CLIP model. This information is used as guiding information to instruct the model to identify factors affecting image quality, and is expressed as: = In the formula: This represents the content inference loss; This represents the Long-CLIP model.

3. The quality assessment method for image content inference enhancement in a confined environment of a mine according to claim 2, characterized in that, Step 5: Calculate the final loss for quality assessment. First, the Long-CLIP model is used to directly assess the quality and obtain the original loss. Then the original loss Content reasoning loss obtained in Step 4 The sum gives the final loss. Specifically, the text encoder of the CLIP model directly predicts the image quality and quality attributes by inputting a designed prompt into the prompt, which is a descriptive text in both positive and negative aspects: positive... {high} quality photo and negative The low-quality photo, the design prompt is denoted as ,in =0, 1, respectively, indicating that prompt belongs to the th... The classes are CLIP's text encoder and image encoder, respectively denoted as... and CLIP for images The predicted score is expressed as: In the formula: Text features; A text encoder for CLIP; A prompt for the design; Image features; For CLIP image encoder; Cosine similarity; Represents an image; Representing an image The predicted score obtained from the prompt for the design; for Cosine similarity; for Cosine similarity; for The exponential function; for The exponential function; the original loss is obtained by directly performing quality assessment using the Long-CLIP model. , is represented as: = ( In the formula: Indicates the original loss; Represents the Long-CLIP model; Representing an image The predicted score obtained from the prompt for the design; Representing an image The true score; the original loss Content reasoning loss obtained in Step 4 The sum gives the final loss. , is represented as: = In the formula: Loss for content inference The weight parameters.

4. A quality assessment system for image content inference enhancement in confined mine environments, based on the quality assessment method for image content inference enhancement as described in claim 1, characterized in that, It includes an image content description module, a content reasoning enhancement module, and an image quality evaluation module.

Citation Information

Patent Citations

  • Video quality assessment model training method, quality assessment method, device and medium

    CN119763018A