Mine complex environment content understanding-based less-label quality evaluation method and system

The detailed content description and language model are generated through the visual language model to obtain the "actual-antisation" text pair, combined with the global-local graphic similarity relationship distillation and the "actual-antisation" text description pair guidance, which enhances the model's understanding of the content of the complex environment images of the mine, solves the problem of high labeling cost of image data sets, and improves the quality evaluation accuracy and generalization ability.

CN120526428APending Publication Date: 2025-08-22CHINA UNIV OF MINING & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510619393.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The image dataset labeling work cost and low efficiency in complex mine environments. The existing pre-trained models are prone to overfitting during fine-tuning, resulting in a decrease in the model's ability to understand the image content, affecting the quality evaluation accuracy and generalization ability.

Method used

The visual language model is used to generate detailed content descriptions, and the "actual-antisation" text pairs are obtained through the language model, and the global-local graphic and text similarity relationship is used to distillate and "actual-antisation" text description pairs are guided to enhance the model's understanding of image content, and fine-tune the pre-trained model to adapt to quality evaluation tasks through multi-task learning.

Benefits of technology

It improves the image quality evaluation accuracy of the model in a complex mine environment, enhances the processing ability of content diversity, is suitable for labeling of image data sets with less labels, reduces the annotation cost and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526428A_ABST
    Figure CN120526428A_ABST
Patent Text Reader

Abstract

The invention discloses a low-label quality evaluation method and system based on mine complex environment content understanding, and the method comprises the steps: firstly generating detailed content description aligned with a quality score through a visual language large model, obtaining an'actual-antisense 'text pair of the content description through the language large model, and then carrying out the quality evaluation while adapting to the quality evaluation. By distilling the relationship between local and global text and graph similarities, the content understanding ability of the pre-training model is reserved, and the training model is close to actual description and far away from antisense description. According to the low-label quality evaluation method based on mine complex environment content understanding, when the pre-training model is finely adjusted to adapt to the quality evaluation task, the ability of the model to understand the image content and the ability of processing the content diversity can be enhanced, and the quality evaluation precision of the model can be improved on the basis of understanding the content; the method is especially suitable for image data set labeling work in a complex mine environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image quality assessment method and system, in particular to a few-label quality assessment method and system based on understanding the complex environment content of a mine, belonging to the technical field of image quality evaluation. Background Art

[0002] No-reference quality assessment algorithms for real-world distorted images in complex mine environments have broad applications in image acquisition, processing, and display. They provide strong support for coal mine safety, process monitoring, efficiency improvement, and regulatory decision-making, and have therefore garnered significant attention. The unique operating conditions of mine environments not only pose a threat to safe mining but also significantly impact the quality of image acquisition. Images acquired in complex mine environments often fail to meet the high-quality input data requirements of subsequent image processing. In critical applications such as mine safety monitoring and disaster warning, low-quality images can lead to false alarms and missed alarms in monitoring systems, directly threatening mine safety and personnel safety. With the promotion and implementation of the concept of intelligent mines, the demand for real-time monitoring and efficient management of mine environments is growing. Image quality has become a key factor restricting the advancement of intelligent mines. Therefore, comprehensive and systematic assessment and optimization of mine image quality, as well as the development of image enhancement and restoration technologies adapted to complex mine environments, have become important research directions for improving mine safety management effectiveness and ensuring safe production operations. Unlike simulated distortion, real-world distorted images lack direct reference information. Their complexity lies in the interweaving of multidimensional distortions and the close correlation between image content and distortion. Therefore, developing a no-reference quality assessment algorithm for real-world distorted images in complex mine environments is crucial.

[0003] However, the task of labeling image datasets in the complex environment of mines faces numerous challenges, resulting in relatively high costs. Therefore, there is an urgent need to research low-label quality assessment algorithms to reduce the model's dependence on the number of samples. Due to the small number of samples in quality assessment datasets, existing mainstream quality assessment algorithms mostly adapt pre-trained models to the quality assessment task. Pre-trained models are mostly designed for image understanding tasks. On the one hand, although pre-trained models have strong semantic understanding capabilities, their semantic understanding capabilities are significantly reduced during fine-tuning due to catastrophic forgetting, which in turn causes the model to overfit to the quality assessment task. This reduces the original content understanding capabilities of the pre-trained models, affecting the model's ability to handle content diversity, resulting in reduced prediction accuracy and generalization capabilities. On the other hand, image quality assessment requires understanding image content, which is crucial for quality assessment tasks. However, due to the limited generalization capabilities of pre-trained models, the model's ability to understand the content of quality assessment scene samples is often insufficient, making it difficult to fully utilize the specific content of the image to evaluate image quality. Finally, how to effectively handle content diversity is a fundamental issue in quality assessment tasks. This problem becomes more prominent as the number of training samples decreases, leading to a sharp decline in the model's ability to handle content diversity and evaluation accuracy. Therefore, for the image dataset annotation work in the complex environment of mines, manual labeling is not only time-consuming and labor-intensive, but also has problems such as ignoring insignificant areas and brief annotations. It is often difficult to obtain image content labels for quality evaluation scenarios. Summary of the Invention

[0004] In response to the problems existing in the above-mentioned prior art, the present invention provides a low-label quality assessment method and system based on content understanding in complex mine environments. This method can not only enhance the model's ability to understand image content and handle content diversity when fine-tuning the pre-trained model to adapt to quality assessment tasks, but also improve the model's quality assessment accuracy based on content understanding. It is particularly suitable for image dataset labeling work in complex mine environments.

[0005] To achieve the above objectives, this low-label quality assessment method based on understanding the complex environment of mines specifically includes the following steps:

[0006] Step 1: Obtain a detailed description of the image content based on the visual language model: Use an unsupervised approach to describe the image content in detail using the visual language model.

[0007] Step 2: Obtain "actual-antonymous" text description pairs based on the language model and content description: Using the language model, the detailed image content description text is processed to obtain three "actual-antonymous" text description pairs. The actual description text is extracted based on the detailed content description, and the antonym description is a description with the opposite meaning generated based on the actual description.

[0008] Step 3: Preserve the pre-trained model's content understanding capability based on global-local image-text similarity distillation: Based on the Long-CLIP model, the similarity between the image and the full text description, as well as the similarity between the image and the partial description, is obtained. Relationship distillation technology is used to maintain the invariance of this relationship during training, thereby preserving the pre-trained model's understanding of the image content.

[0009] Step 4: Understanding the content of complex mine environment images based on "actual-antonymous" text description pairs: Using a multi-task approach, with the guidance of "actual-antonymous" text description pairs, the model is trained to approach "actual descriptions" and move away from "antonymous descriptions";

[0010] Step 5, visual quality assessment of few-label images based on content understanding of complex mine environments: While being guided by global-local relationship distillation and "actual-antonymous" text descriptions, a multi-task learning approach is used to fine-tune the pre-trained model to adapt to the quality evaluation task.

[0011] Furthermore, Step 1 is as follows:

[0012] Step 1-1, select the visual language model: select ShareGPT4V-Captioner as the visual language model;

[0013] Step 1-2: Design prompt words to generate image content description: After inputting the image to be described, input the prompt word "Analyze the image in a comprehensive and detailed manner" into the visual language model. The visual language model will give a detailed text description, expressed as:

[0014] text=Cap(I,p)

[0015] Where: Cap represents the image description generation model ShareGPT4V-Captioner; I represents the image; p represents the input prompt word; text is the generated text.

[0016] Furthermore, Step 2 is as follows:

[0017] Step 2-1, select the language model: select Llama3-8B-Instruct as the language model;

[0018] Step 2-2, design prompt words to generate text description pairs: Input the image to be described, and input the prompt words to the language model: "Generate three pairs of opposite text descriptions based on the following text:" + "<content description text>". The language model will generate three pairs of "actual-antonymous" text description pairs, expressed as:

[0019]

[0020] Where: LLM represents the language model Llama3-8B-Instruct; p2 represents the input prompt word; text is the content text generated in Step 1; They represent three different pairs of “actual-antonymous” text descriptions obtained by the model.

[0021] Furthermore, Step 3 is as follows:

[0022] Step 3-1, segmenting the detailed text description: The detailed image content text description is segmented into n different sub-descriptions, where n = 10. First, perform the initial segmentation based on the period. If the number of initial segments is greater than n, merge them to make the number of paragraphs equal to n. If the number of initial segments is less than n, combine different sub-paragraphs to supplement them until the number of sub-paragraphs equals n. This is expressed as:

[0023] s1,s2,…,s n =split(text)

[0024] Where: split represents the split operation; text represents the detailed text description; s1, s2, ..., s n Describe the sub-paragraphs of the local text obtained by segmentation;

[0025] Step 3-2, calculate the global-local image-text similarity: input the image to be described, as well as the global text description text and the local description text s1, s2, ..., s n , the similarity between image and text is calculated with the help of LongCLIP model, which is expressed as:

[0026] s g ,s1,s2,...,s n =Softmax[LongCLIP(I,text,s1,s2,...,s n )]

[0027] Where: s g is the global text-image similarity, which represents the image I and the local description text s1, s2, ..., s nSimilarity; Softmax is the activation function;

[0028] Step 3-3, calculate the similarity relationship distillation loss function: First, obtain the similarity distribution relationship of the original Long-CLIP model, and in the process of training the model, distill this relationship to maintain the original content understanding ability, which is expressed as:

[0029]

[0030] Where: Represents the similarity distribution relationship between image I and global-local text; Similarity distribution relationship between an antonym image I and global-local text; LongCLIP pre Represents the pre-trained LongCLIP model; LongCLIP IQA represents the quality evaluation model after fine-tuning; KL represents the Kullback-Leibler divergence; l d represents the distillation loss function.

[0031] Furthermore, Step 4 is as follows:

[0032] Step 4-1, calculate the similarity between the image and text description pair: input the image to be described and the "actual-antonym" text description pair The similarity between images and text is calculated using the LongCLIP model, which is expressed as:

[0033]

[0034] Where: Indicates the actual description text; Indicates antonym description text; Represents image I and actual description text similarity; Represents image I and antonym text similarity;

[0035] Step 4-2, calculate the text prediction loss function: with the help of the text prediction loss function "actual-antonymous" text description pair guidance, the training model is close to the "actual description" and away from the "antonymous description", where the actual description text Similarity score label with the image Antonym description text The similarity score label with the image should be Then the text prediction loss function is expressed as:

[0036]

[0037] Where: l p Represents the “actual-antonym” text description prediction loss function.

[0038] Furthermore, Step 5 is as follows:

[0039] Step 5-1, image quality assessment based on prompt learning: Input the image to be described and input the quality assessment prompt word pair p1, p2 to the model: "p1 = ******high quality", "p2 = ******low quality", where "*" represents a learnable parameter, which is updated through prompt learning. The similarity between the image and the quality assessment prompt word is calculated with the help of the LongCLIP model, which is expressed as:

[0040] q1,q2=Softmax[LongCLIP(I,p1,p2)]

[0041] Where: q1 represents the similarity between image I and text p1; q2 represents the similarity between image I and text p2;

[0042] Let "high quality" and "low quality" be 1 and 0 points respectively, then q1 represents the score of the image;

[0043] Step 5-2, calculation of the overall loss function: normalize the image quality score to 0-1 and record it as y, then the overall loss function is expressed as:

[0044] l total =MSE(q1,y)+λ1l d +λ2l p

[0045] Where: l total is the overall loss function; MSE is the mean square error loss function; λ1 is the weight parameter of the distillation loss function; λ2 is the weight parameter of the “actual-antonym” text description prediction loss function.

[0046] A low-label quality assessment system based on content understanding of complex mine environments includes an image content description module, a text segmentation and distillation module, a text description pair generation module, a content understanding enhancement module, and an image quality evaluation module.

[0047] Considering that when the mainstream algorithm fine-tunes the pre-trained model to adapt to the quality evaluation task, the model's ability to understand the content will be reduced and it will overfit to the quality evaluation task. This problem is particularly prominent in the few-label scenario. Therefore, compared with the existing technology, this few-label quality assessment method based on content understanding in complex mine environments uses a large model to enhance the algorithm's understanding of the content to improve the quality evaluation accuracy. First, with the help of a large visual language model, a detailed content description aligned with the quality score is generated, and a large language model is used to obtain the "actual-antonym" text pairs of the content description. Then, while adapting the quality evaluation, the content understanding ability of the pre-trained model is retained by distilling the relationship between local and global text-image similarities, and the model is trained to be close to the "actual description" and away from the "antonym description", thereby enhancing the model's understanding of the content of the target image and improving the algorithm accuracy. This low-label quality assessment method based on content understanding in complex mine environments uses global-local relationship distillation and "actual-antonym" text description guidance, while adopting multi-task learning to fine-tune the pre-trained model to adapt to the quality assessment task. It can not only enhance the model's ability to understand image content and handle content diversity, but also improve the model's quality assessment accuracy based on content understanding. It is particularly suitable for image dataset annotation work in complex mine environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow chart of the present invention;

[0049] Figure 2 It is an algorithm block diagram of the present invention;

[0050] Figure 3 1 is a diagram illustrating ablation experiment results on different data sets provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] This low-label quality assessment method based on content understanding in complex mine environments first uses a large visual language model to generate detailed content descriptions aligned with quality scores, and uses the large language model to obtain "actual-antonymous" text pairs of the content descriptions. Then, while adapting the quality evaluation, it retains the content understanding ability of the pre-trained model by distilling the relationship between local and global text-image similarities, and trains the model to be close to the "actual description" and away from the "antonymous description", thereby enhancing the model's content understanding of the target image and improving the algorithm's accuracy.

[0052] like Figure 1 As shown in the figure, the low-label quality assessment method based on the understanding of the complex environment of the mine specifically includes the following steps:

[0053] Step 1: Obtain a detailed description of the image content based on the visual language model: Using an unsupervised approach, the visual language model is used to describe the image content in detail, thus avoiding the problems of manual annotation, such as high cost, low efficiency, inadequate description, and neglect of insignificant image areas. The details are as follows:

[0054] Step 1-1: Select a large visual language model. This paper selects ShareGPT4V-Captioner as the large visual language model. ShareGPT4V has 100,000 high-quality image text labels collected from the advanced GPT4-Vision. By building a labeling model on this subset, it is expanded to 1.2 million highly descriptive image description text labels, surpassing the diversity and information content of other existing datasets. It covers world knowledge, object attributes, spatial relationships, and aesthetic evaluations, and can generate detailed image content description text that can focus on areas that are easily overlooked by the human eye.

[0055] Step 1-2: Design prompt words to generate image content description. After inputting the image to be described, input the prompt word "Analyze the image in a comprehensive and detailed manner" to the model, and the model will give a detailed text description, which can be expressed as:

[0056] text=Cap(I,p)

[0057] Where: Cap represents the image description generation model ShareGPT4V-Captioner; I represents the image; p represents the input prompt word; text is the generated text.

[0058] Step 2: Obtain "actual-antonymous" text description pairs based on the language model and content description: Using the language model, the detailed image content description text is processed to obtain three "actual-antonymous" text description pairs. The actual description text is extracted based on the detailed content description, and the antonym description is a description with the opposite meaning generated based on the actual description. The details are as follows:

[0059] Step 2-1: Select a large language model. This paper selects Llama3-8B-Instruct as the large language model. This model has 8 billion parameters and has set a new performance record among large language models of the same scale, becoming one of the best performing models in the industry at this parameter level. Llama 3 significantly reduces the false rejection rate, enhances the consistency of the model, and enriches its response diversity by optimizing the post-training process. In addition, Llama 3 has also made significant progress in reasoning, code generation, and following instructions, making the model more flexible and easy to use.

[0060] Step 2-2: Design prompt words to generate text description pairs. Input the image to be described and give the model the prompt words: "Generate three pairs of opposite text descriptions based on the following text:" + "<content description text>". The model will then generate three pairs of "actual-antonymous" text descriptions, which can be expressed as:

[0061]

[0062] Where: LLM represents the language model Llama3-8B-Instruct; p2 represents the input prompt word; text is the content text generated in Step 1; They represent three different pairs of “actual-antonymous” text descriptions obtained by the model.

[0063] Step 3: Preserve the pre-trained model's ability to understand content based on global-local image-text similarity distillation: Based on the Long-CLIP model, we obtain the similarity between the image and the full text description, as well as the similarity between the image and the partial description, i.e., global-local image-text similarity. We use relationship distillation technology to maintain the invariance of this relationship during training, thereby preserving the pre-trained model's understanding of the image content. The details are as follows:

[0064] Step 3-1, segment the detailed text description. The detailed text description of the image content is segmented into n different sub-descriptions. In this invention, n=10. First, perform the initial segmentation based on the period. If the number of initial segments is greater than n, merge them until the number of sub-paragraphs equals n. If the number of initial segments is less than n, combine different sub-paragraphs to supplement them until the number of sub-paragraphs equals n. It can be expressed as:

[0065] s1,s2,…,s n =split(text)

[0066] Where: split represents the split operation; text represents the detailed text description; s1, s2, ..., s n The local text description sub-paragraph obtained by segmentation.

[0067] Step 3-2, calculate the global-local image-text similarity. Input the image to be described, the global text description text and the local description text s1, s2, ..., s n, we use the LongCLIP model to calculate the similarity between images and texts. The LongCLIP model is modified based on the Contrastive Language-Image Pre-Training (CLIP) model and can process longer texts. It can be expressed as:

[0068] s g ,s1,s2,...,s n =Softmax[LongCLIP(I,text,s1,s2,...,s n )]

[0069] Where: s g is the global text-image similarity, which represents the image I and the local description text s1, s2, ..., s n Softmax is the activation function.

[0070] Step 3-3, calculate the similarity relationship distillation loss function. When training the LongCLIP model to adapt to the quality evaluation task, the text-image similarity score output by the model will change with the change of parameters, while the similarity distribution relationship between the image and the global-local text can remain unchanged, that is, The distribution of can remain unchanged, so the present invention first obtains the similarity distribution relationship of the original Long-CLIP model, and in the process of training the model, by distilling this relationship, maintains the original content understanding ability, which can be expressed as:

[0071]

[0072] Where: Represents the similarity distribution relationship between image I and global-local text; Similarity distribution relationship between an antonym image I and global-local text; LongCLIP pre Represents the pre-trained LongCLIP model; LongCLIP IQA represents the quality evaluation model after fine-tuning; KL represents the Kullback-Leibler divergence; l d represents the distillation loss function.

[0073] Step 4: Understanding the content of complex mine environment images based on the guidance of "actual-antonymous" text description pairs: Using a multi-task approach and with the guidance of "actual-antonymous" text description pairs, the training model is close to the "actual description" and away from the "antonymous description".

[0074] The details are as follows:

[0075] Step 4-1, calculate the similarity between the image and text description. Input the image to be described and the “actual-antonym” text description pair The similarity between images and texts is calculated using the LongCLIP model, which can be expressed as:

[0076]

[0077] Where: Indicates the actual description text; Indicates antonym description text; Represents image I and actual description text similarity; Represents image I and antonym text similarity.

[0078] Step 4-2, calculate the text prediction loss function. In order to further enhance the model's ability to understand the image content of the quality evaluation data, the training model is close to the "actual description" and away from the "antonym description" with the guidance of the text prediction loss function "actual-antonym" text description pair. Similarity score label with the image Antonym description text The similarity score label with the image should be Then the text prediction loss function is expressed as:

[0079]

[0080] Where: l p Represents the “actual-antonym” text description prediction loss function.

[0081] Step 5: Visual quality assessment of few-label images based on understanding the complex environment of the mine: While using global-local relationship distillation and "actual-antonymous" text description guidance, a multi-task learning approach is used to fine-tune the pre-trained model to adapt to the quality assessment task. The details are as follows:

[0082] Step 5-1: Image quality evaluation based on prompt learning. Input the image to be described and the quality evaluation prompt word pair p1, p2 to the model: "p1 = ******high quality", "p2 = ******low quality", where "*" represents a learnable parameter. It is updated through prompt learning and the similarity between the image and the quality evaluation prompt word is calculated using the LongCLIP model, expressed as:

[0083] q1,q2=Softmax[LongCLIP(I,p1,p2)]

[0084] Where: q1 represents the similarity between image I and text p1; q2 represents the similarity between image I and text p2;

[0085] Let "high quality" and "low quality" be 1 and 0 points respectively, and q1 represents the score of the image.

[0086] Step 5-2, calculate the overall loss function. Normalize the image quality score to 0-1 and record it as y. The overall loss function can be expressed as:

[0087] l total =MSE(q1,y)+λ1l d +λ2l p

[0088] Where: l total is the overall loss function; MSE is the mean square error loss function; λ1 is the weight parameter of the distillation loss function; λ2 is the weight parameter of the “actual-antonym” text description prediction loss function.

[0089] The system framework diagram of this low-label quality assessment method based on the understanding of the complex environment of the mine is as follows: Figure 2 As shown in the figure, the system includes an image content description module, a text segmentation and distillation module, a text description pair generation module, a content understanding enhancement module and an image quality evaluation module. The algorithm flow is as follows: first, prompt words are designed for the few-label dataset, and then a detailed description of the image content is obtained with the help of the visual language large model ShareGPT4V-Captioner; then, the text segmentation of the detailed content is performed to distill the local and global image-text similarity relationships, and at the same time, the "actual-antonym" text description pairs are generated with the help of the language large model to guide the model to understand the image content; then, while enhancing the model's understanding of the content and alleviating the overfitting problem of few labels, with the assistance of relationship distillation and content understanding, the image quality can be evaluated under the guidance of a small number of labeled samples.

[0090] The technical effects of the present invention are described in detail below in combination with performance tests and experimental analysis. Existing no-reference quality assessment algorithms typically use 80% of the images to train the model, while the present invention aims to reduce the model's training data. To demonstrate the performance of this model under low-label conditions, the present invention uses 5% and 10% of the image training samples from the real distortion dataset KonIQ-10k. The results are compared with existing algorithms: using free energy principle for blind image quality assessment (NFERM), blind image quality assessment based on high order statistics aggregation (HOSA), deep neural networks for no-reference and full-reference image quality assessment (WaDIQaM-NR), deep meta-learning for no-reference image quality assessment (MetaIQA), and active learning-based sample selection for label-efficient blind image quality assessment (AL-IQA). The results are shown in the following table:

[0091] Number of training samples 5% KonIQ-10k 10% KonIQ-10k NFERM 0.615 0.651 HOSA 0.730 0.751 WaDIQaM-NR 0.678 0.723 MetaIQA 0.796 0.821 AL-IQA 0.859 0.888 Algorithm of the present invention 0.866 0.891

[0092] As can be seen from the above table, compared with the above-mentioned best existing model algorithms, the algorithm of the present invention can achieve the best few-label prediction ability.

[0093] To further verify the effectiveness of the low-label quality assessment method based on the understanding of the complex environment of the mine, the algorithm of the present invention was subjected to an ablation experiment on KonIQ-10k. The prompt words and CLIP model directly trained by the present invention were used as baseline models and compared with the algorithm of the present invention. All models were trained using 1% of the KonIQ-10k images. The comparison results are shown in Figure 2. Figure 3 As shown. Figure 3 It can be seen that if the relationship distillation and text pair prediction strategies used in the algorithm of the present invention are not adopted, the prediction ability of the baseline model will be significantly reduced.

[0094] This low-label quality assessment method based on content understanding in complex mine environments uses a large model to enhance the algorithm's understanding of content to improve the accuracy of quality evaluation. By enhancing the model's understanding of image content, it can improve the accuracy of the low-label, no-reference quality evaluation algorithm and reduce the number of labels for constructing new scene quality evaluation algorithms, which is conducive to the efficient and low-cost construction and deployment of quality evaluation algorithms. While guiding global-local relationship distillation and "actual-antonym" text descriptions, multi-task learning is used to fine-tune the pre-trained model to adapt to the quality evaluation task, which can not only enhance the model's ability to understand image content and handle content diversity, but also improve the model's quality evaluation accuracy based on content understanding. It is particularly suitable for image dataset annotation work in complex mine environments.

Claims

1. A low-label quality assessment method based on understanding the complex environment of mines, characterized by: The specific steps include: Step 1: Obtain a detailed description of the image content based on the visual language model: Use an unsupervised approach to describe the image content in detail using the visual language model. Step 2: Obtain "actual-antonymous" text description pairs based on the language model and content description: Using the language model, the detailed image content description text is processed to obtain three "actual-antonymous" text description pairs. The actual description text is extracted based on the detailed content description, and the antonym description is a description with the opposite meaning generated based on the actual description. Step 3: Preserve the pre-trained model's content understanding capability based on global-local image-text similarity distillation: Based on the Long-CLIP model, the similarity between the image and the full text description, as well as the similarity between the image and the partial description, is obtained. Relationship distillation technology is used to maintain the invariance of this relationship during training, thereby preserving the pre-trained model's understanding of the image content. Step 4: Understanding the content of complex mine environment images based on "actual-antonymous" text description pairs: Using a multi-task approach and guided by "actual-antonymous" text description pairs, the model is trained to approach "actual descriptions" and avoid "antonymous descriptions." Step 5: Visual quality assessment of few-label images based on content understanding of the complex mine environment: While guided by global-local relationship distillation and "actual-antonym" text descriptions, a multi-task learning approach is used to fine-tune the pre-trained model to adapt to the quality evaluation task.

2. The low-label quality assessment method based on understanding the complex environment of a mine according to claim 1 is characterized in that: Step 1 is as follows: Step 1-1, select the visual language model: select ShareGPT4V-Captioner as the visual language model; Step 1-2, design prompt words to generate image content description: After inputting the image to be described, input the prompt word "Analyze the image in a comprehensive and detailed manner" into the visual language model. The visual language model will give a detailed text description, expressed as: text=Cap(I,p) Where: Cap represents the image description generation model ShareGPT4V-Captioner; I represents the image; p represents the input prompt word; text is the generated text.

3. The low-label quality assessment method based on understanding the complex environment of a mine according to claim 2 is characterized in that: Step 2 is as follows: Step 2-1, select the language model: select Llama3-8B-Instruct as the language model; Step 2-2, design prompt words to generate text description pairs: Input the image to be described, and input the prompt word "Generate three pairs of opposite text descriptions based on the following text:" + "<content description text>" into the language model. The language model will generate three pairs of "actual-antonymous" text description pairs, expressed as: Where: LLM represents the language model Llama3-8B-Instruct; p2 represents the input prompt word; text is the content text generated in Step 1; They represent three different pairs of "actual-antonymous" text descriptions obtained by the model.

4. The low-label quality assessment method based on understanding the complex environment of a mine according to claim 3 is characterized in that: Step 3 is as follows: Step 3-1, segmenting the detailed text description: The detailed image content text description is segmented into n different sub-descriptions, where n = 10. First, perform the initial segmentation based on the period. If the number of initial segments is greater than n, merge them to make the number of paragraphs equal to n. If the number of initial segments is less than n, combine different sub-paragraphs to supplement them until the number of sub-paragraphs equals n. This is expressed as: s1,s2,…,s n =split(text) Where: split represents the split operation; text represents the detailed text description; s1, s2, ..., s n Describe the sub-paragraphs of the local text obtained by segmentation; Step 3-2, calculate the global-local image-text similarity: input the image to be described, as well as the global text description text and the local description text s1, s2, ..., s n , the similarity between image and text is calculated with the help of LongCLIP model, which is expressed as: s g ,s1,s2,...,s n =Softmax[LongCLIP(I,text,s1,s2,...,s n )] Where: s g is the global text-image similarity, which represents the image I and the local description text s1, s2, ..., s n Similarity; Softmax is the activation function; Step 3-3, calculate the similarity relationship distillation loss function: First, obtain the similarity distribution relationship of the original Long-CLIP model, and in the process of training the model, distill this relationship to maintain the original content understanding ability, which is expressed as: Where: Represents the similarity distribution relationship between image I and global-local text; Similarity distribution relationship between an antonym image I and global-local text; LongCLIP pre Represents the pre-trained LongCLIP model; LongCLIP IQA represents the quality evaluation model after fine-tuning; KL represents Kullback-Leibler divergence; l d represents the distillation loss function.

5. The low-label quality assessment method based on understanding the complex environment of a mine according to claim 4 is characterized in that: Step 4 is as follows: Step 4-1, calculate the similarity between the image and text description pair: input the image to be described and the "actual-antonym" text description pair The similarity between images and text is calculated using the LongCLIP model, which is expressed as: Where: Indicates the actual description text; Indicates antonym description text; Represents image I and actual description text similarity; Represents image I and antonym text similarity; Step 4-2, calculate the text prediction loss function: with the help of the text prediction loss function "actual-antonymous" text description pair guidance, the training model is close to the "actual description" and away from the "antonymous description", where the actual description text Similarity score label with the image Antonym description text The similarity score label with the image should be Then the text prediction loss function is expressed as: Where: l p Represents the "actual-antonym" text description prediction loss function.

6. The low-label quality assessment method based on understanding the complex environment of a mine according to claim 5 is characterized in that: Step 5 is as follows: Step 5-1, image quality evaluation based on prompt learning: Input the image to be described, and input the quality evaluation prompt word pair p1, p2 to the model: "p i =******high quality", "p2 =******low quality", where "*" represents a learnable parameter, which is updated through prompt learning. The similarity between the image and the quality evaluation prompt word is calculated with the help of the LongCLIP model, which is expressed as: q1,q2=Softmax[LongCLIP(I,p1,p2)] Where: q1 represents the similarity between image I and text p1; q2 represents the similarity between image I and text p2; Let "high quality" and "low quality" be 1 and 0 points respectively, then q1 represents the score of the image; Step 5-2, calculation of the overall loss function: normalize the image quality score to 0-1 and record it as y, then the overall loss function is expressed as: l total =MSE(q1,y)+λ1l d +λ2l p Where: l total is the overall loss function; MSE is the mean square error loss function; λ1 is the weight parameter of the distillation loss function; λ2 is the weight parameter of the "actual-antonym" text description prediction loss function.

7. A low-label quality assessment system based on mine complex environment content understanding based on the low-label quality assessment method based on mine complex environment content understanding according to claim 1, characterized in that: It includes image content description module, text segmentation and distillation module, text description pair generation module, content understanding enhancement module and image quality evaluation module.

Citation Information

Patent Citations

  • Adversarial pre-training of machine learning models

    CN115485696A

  • Real scene severe weather image restoration method based on visual language model

    CN118537264A

  • Image aesthetics quality evaluation method based on prompt learning

    CN118865387A

  • Video quality assessment model training method, quality assessment method, device and medium

    CN119763018A