Mine disease identification method based on vision-language model and prompt optimization
Through the visual-language pre-training model and prompt word optimization method, the problems of insufficient data and poor adaptability of complex scenes in mine disease detection are solved, and high-precision and low-dependence automated disease recognition are achieved, which is suitable for zero-sample and small-sample environments.
Patent Information
- Application Number
- CN202510319157.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There are problems in the detection of mine road diseases, such as scarcity of data samples, insufficient detection accuracy and poor adaptability of complex scenarios. Especially in the zero-sample and small-sample environment, existing deep learning models and manual detection methods are difficult to effectively apply.
Combining the visual-language pre-training model and prompt word optimization method, high-precision classification and detection of mine diseases are achieved by constructing domain-specific multimodal embedding, multi-view angle enhancement, dynamic threshold matching and pseudo-label adaptive learning.
Under zero sample and few sample conditions, high-precision detection of mine diseases is achieved, the dependence on large amounts of labeled data is reduced, and the adaptability and robustness of the model in complex scenarios is improved.
Smart Images

Figure CN120354196A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the application of multimodal machine learning in industrial inspection, and specifically to a mine disease identification method based on a vision-language model and prompt optimization. Background Art
[0002] Mine roads are crucial infrastructure in mining operations. Due to the complex mine environment and frequent vehicle passage, mine roads often suffer from various forms of damage and diseases, such as cracks, settlements, potholes, and mud pumping. These diseases not only affect the flatness and safety of mine roads but also exacerbate equipment wear, increase operating costs, and may even cause the interruption of mine operations in severe cases, resulting in huge economic losses. Therefore, timely and accurate detection of mine road diseases is of great significance for the daily maintenance and safety guarantee of mines.
[0003] Traditional mine road disease detection methods mainly rely on manual observation and manual recording. Usually, on-site inspectors need to check the road surface section by section to judge the type and severity of diseases. Although this method is intuitive, it has the following several significant defects:
[0004] 1. High labor cost: Manual detection requires a large number of inspectors. Especially in large mines, the labor and time costs are very high, and the inspection efficiency is low.
[0005] 2. Limited accuracy: Manual judgment is subjective, and the detection results may vary due to the experience and fatigue of inspectors, resulting in judgment errors and affecting the accuracy of detection.
[0006] 3. Lack of timeliness: Manual inspection cannot achieve real-time monitoring of mine roads, and diseases may deteriorate rapidly in a short time, leading to potential safety hazards.
[0007] In order to improve the detection efficiency, in recent years, deep learning technology has been gradually introduced into the field of mine disease detection, and image classification and object detection models are used to achieve partial automated disease identification.
[0008] However, this method has some insurmountable problems:
[0009] 1. Data dependence: The training of deep learning models requires a large number of labeled image data sets. However, it is extremely difficult to obtain mine disease data, and the labeling workload is large, resulting in limited data samples.
[0010] 2. Insufficient generalization ability: The types, forms, and distributions of mine diseases vary significantly due to differences in mine environments. The generalization ability of conventional deep learning models is difficult to cope with complex and variable disease characteristics.
[0011] 3. Difficult model adjustment: Due to the significant differences in mine environments, the distribution of disease data in each mine may vary greatly. The model usually needs to be fine-tuned for specific scenarios, resulting in a complex application process and difficulty in wide promotion.
[0012] To address the problems of insufficient data and poor model generalization, in recent years, multi-modal learning methods based on vision-language pre-trained models have emerged. Such models (e.g., CLIP) combine the features of images and texts. By learning the common representation space of images and texts, the model can achieve good recognition results even in zero-shot or few-shot scenarios. The multi-modal learning ability of the CLIP model enables it to automatically classify and detect diseases by combining image and text prompts without a large amount of labeled data.
[0013] However, there are still the following difficulties in directly applying vision-language pre-trained models to mine disease detection:
[0014] 1. Insufficient construction of domain prompts: The prompts of existing vision-language pre-trained models are mainly general domain descriptions, which are difficult to accurately express the specific features of mine diseases, resulting in the model's difficulty in effectively utilizing the visual features of disease images and the semantic information of prompts.
[0015] 2. Low adaptability to complex scenarios: The image data of mine diseases is complex and variable. Existing models lack data augmentation methods specifically for mine scenarios, leading to insufficient recognition ability of the model for diseases with multiple perspectives, scales, and orientations.
[0016] 3. Imperfect few-shot learning strategy: In zero-shot and few-shot scenarios, the model needs to dynamically learn and iteratively optimize disease features. However, existing pre-trained models lack adaptive learning strategies and are difficult to continuously optimize the recognition performance through limited samples.
[0017] To address the above problems, the present invention proposes a mine road disease recognition method and system that combines a vision-language pre-trained model and prompt optimization, using domain-specific prompt construction, multi-modal embedding optimization, adaptive threshold adjustment, and pseudo-label adaptive learning mechanisms to effectively improve the accuracy and robustness of mine road disease recognition. Summary of the Invention
[0018] The objective of the present invention is to solve the problems of scarce data samples, insufficient detection accuracy, and poor adaptability to complex scenarios in mine road disease detection. Especially in zero-shot and few-shot environments, existing traditional deep learning models and manual detection methods are difficult to be effectively applied. For this reason, the present invention proposes a mine road disease recognition method and system based on a vision-language pre-trained model and prompt optimization to achieve high-precision classification and detection of mine diseases.
[0019] To solve the above technical problems, the present invention is implemented through the following technical solutions:
[0020] In the first aspect, a mine disease recognition method based on a vision-language model and prompt optimization, the method includes the following steps:
[0021] S1: Multimodal embedding generation: Input the mine disease image and the text prompt words constructed based on domain-specific knowledge into the vision-language pre-trained model respectively. The model extracts the visual feature vector V i and the text feature vector T j and maps the two into the same embedding vector space; the core representation of the embedding generation is:
[0022] E i = f(V i , T j ) = W v V i + W t T j ;
[0023] where W v and W t are the weight matrices of visual and text embeddings respectively;
[0024] S2: Prompt word construction and optimization: Based on the specific manifestation forms of mine diseases, extract the description words related to diseases from the dictionaries and professional term libraries in the mining and construction engineering fields, construct a text prompt word set covering various aspects of disease types, disease causes, disease morphologies, etc., extract the high-frequency words as the core description word set through semantic analysis, and further perform semantic analysis on the text prompt words to ensure that the prompt words have a high semantic relevance in the multimodal embedding space, defined as:
[0025] T core = {t k | freq(t k ) > δ}
[0026] where freqt k represents the frequency of the word t k , and δ is a preset frequency threshold for screening the core words of disease descriptions;
[0027] S3: Disease recognition and classification: Based on the cosine similarity S ij of the multimodal embedding vectors, calculate the similarity score between the input disease image and the prompt words, classify the disease type by selecting the prompt word with the highest score. If the similarity does not reach the preset confidence threshold, generate a pseudo-label for the detection result and add it to the training set, and perform iterative learning, so as to gradually improve the classification performance of the model in the few-shot environment;
[0028]
[0029] Complete the disease type classification by selecting the prompt word with the highest similarity. If the similarity S ij < τ, generate pseudo-labels and improve the few-shot detection performance through iterative optimization.
[0030] In this application, the automatic detection and classification of mine road diseases are specifically realized through the following technical solutions:
[0031] 1. Multimodal embedding generation method
[0032] Input the mine disease image and the domain-specific text prompt word into a vision-language pre-training model (such as CLIP) to generate a multimodal feature representation in a unified embedding space. Specifically, the model extracts the image feature vector V i and the text feature vector T j , and combines the weight matrix to fuse the two to generate the embedding vector E i , which is expressed as:
[0033] E i = W v V i + W t T j
[0034] where W v and W t are the weight matrices for visual and text embeddings. This embedding generation method ensures that the model can fully utilize the multimodal information of images and texts in a few-shot environment, improving the accuracy of disease detection. Brief Description of the Drawings
[0035] 2. Prompt word construction and optimization strategy
[0036] The present invention constructs a domain-specific prompt word set suitable for mine diseases, covering information such as the type, cause, and morphology of diseases. By extracting specific words from dictionaries in the mining and construction fields and performing semantic analysis, a prompt word set covering multiple disease types is generated. In addition, the high-frequency words are refined through the word frequency analysis method to form a core prompt word set, and the screening formula is:
[0037] T core = {t k | freq(t k ) > δ}
[0038] where req(t k ) represents the occurrence frequency of the word t k , and δ is the screening threshold. This optimization strategy ensures the high semantic relevance of the prompt words in the embedding space, enhancing the recognition effect of the model in a few-shot situation.
[0039] 3. Multi - perspective Enhancement - based Visual Prompt Generation Method
[0040] In view of the diversity of mine disease image data, the present invention proposes a data enhancement strategy. By operations such as perspective rotation, scale change, and mirror flipping, multi - perspective visual prompts are generated to cover disease characteristics at different angles, different sizes, and different orientations. The specific operations are as follows:
[0041] Perspective rotation: Rotate the disease image at fixed angles (such as 45 degrees, 90 degrees, 180 degrees) to generate multi - perspective images;
[0042] Scale change: Perform various scaling operations on the disease image to simulate the performance of the disease at different distances;
[0043] Mirror flipping: Obtain multi - azimuth disease images through horizontal and vertical mirror flipping. Through the above - mentioned enhancement methods, the model can handle complex and changeable mine disease images.
[0044] 4. Dynamic Threshold Matching Strategy
[0045] The present invention dynamically adjusts the similarity matching threshold based on the detection confidence of the disease type. On the basis of the initial threshold setting, if the detection confidence does not reach the preset standard, the threshold standard is automatically increased to reduce misjudgment and optimize disease classification. The specific adjustment formula is as follows:
[0046] τ new =τ init +Δτ
[0047] where τ new is the dynamically adjusted threshold, and Δτ is the incremental adjustment value. This strategy ensures that the model gradually improves the accuracy and stability of classification in multiple - round detection tasks by progressively optimizing the detection threshold.
[0048] 5. Pseudo - label Generation and Adaptive Learning Method
[0049] The present invention realizes adaptive learning for the few - sample environment through a pseudo - label generation and iterative optimization mechanism. For detection results with insufficient confidence, pseudo - labels are generated and added to the pseudo - sample set. The model continuously updates the embedding distribution through multiple - round iterative training to enhance the adaptive ability to complex disease types. The conditions for pseudo - label generation are:
[0050] S ij <τ conf
[0051] where S ij is the similarity between the image and the prompt. If it is lower than the confidence threshold τ conf , then a pseudo - label is generated and added to the training set for dynamic update.
[0052] In a specific implementation of the first aspect, a multi-modal embedding is generated for the images and text prompts of mine diseases through a vision-language pre-trained model. The system includes the following steps:
[0053] A1: Generation of disease image embeddings: Input the disease image data one by one into the vision encoding module of the vision-language pre-trained model, and extract the visual feature vector V through the image encoding layer i and convert it into a high-dimensional vector representation to capture the visual details of the disease characteristics;
[0054] V i = σ(W enc *I i + b)
[0055] where W enc is the encoding weight matrix, I i is the input image, * is the self-attention mechanism, and σ is the activation function; the generated visual embedding features ensure that the image details can be accurately characterized;
[0056] A2: Generation of text prompt embeddings: By inputting the constructed disease text prompts into the language encoding module of the same model, generate text feature vectors to capture the disease semantic information and maintain a consistent representation with the image embeddings; the text encoding formula is:
[0057] T j = σ(W text ·T raw + b text )
[0058] where W text is the text encoding matrix, T raw is the original text vector, ensuring that the text features can be consistently embedded in the high-dimensional space;
[0059] A3: Embedding alignment and optimization: Use the principal component analysis (PCA) method to perform dimensionality reduction on the embedding vectors of the images and text prompts, and project the embedding features of different disease types into the same high-dimensional feature space to maximize the feature separation between different disease categories, thereby improving the accuracy and stability of disease classification.
[0060] In a specific implementation of the first aspect, by constructing and optimizing a domain-specific multi-level text prompt set, the performance of the vision-language pre-trained model in zero-shot and few-shot learning is improved. The method includes the following steps:
[0061] B1: Domain vocabulary extraction and construction: Extract relevant terms covering common diseases from the professional glossaries in the fields of mine geology and construction engineering, and generate hierarchical complete sentence descriptions T according to information such as disease types, disease causes, and disease morphologiesfull , used to cover complex diseases in the mine environment;
[0062]
[0063] where t k is the disease vocabulary, and α k is the vocabulary weight coefficient, used to optimize the integrity of semantic information;
[0064] B2: Semantic analysis and core word extraction: For the generated complete sentence prompt words, high-frequency words are screened through word frequency analysis, and the core terms of disease descriptions are extracted to form a core word list, ensuring that the model can still accurately capture the main semantic features of the prompt words under low-sample data conditions;
[0065] B3: Prompt word formatting and standardization: Standardize the sentence format of the complete sentence prompt words and the core word list to ensure the grammatical consistency and vocabulary standardization of the prompt words, enabling the model to more efficiently understand the semantic connotation of the text prompt words.
[0066] In a specific implementation manner of the first aspect, multi-perspective visual prompt words are generated by performing various data augmentation processes on mine disease images to enhance the robustness of mine disease recognition, including the following steps:
[0067] C1: Perspective rotation and scale change: The original disease images are respectively rotated at multiple angles, including angle adjustments such as 45 degrees, 90 degrees, and 180 degrees, and at the same time, the images are scaled at different scaling ratios to generate multi-perspective and multi-scale visual prompt words, enabling the model to handle the complex manifestations of diseases at different angles and scales in the mine;
[0068] C2: Multi-scale transformation operation: The disease images are magnified and reduced from micro to macro to simulate the visual characteristics of diseases at different shooting distances, ensuring that the visual prompt words can cover the full-size feature information of mine diseases;
[0069] C3: Mirror flipping and symmetry enhancement: Symmetry enhancement processing is performed on the images through horizontal flipping and vertical flipping to obtain disease visual prompt words in multiple directions, ensuring that the model has a robust recognition ability for multi-directional perspectives of diseases in the mine.
[0070] In a specific implementation manner of the first aspect, high-precision classification of mine diseases is achieved by dynamically adjusting the similarity matching threshold for disease detection, including the following steps:
[0071] D1: Threshold Initialization and Confidence Setting: Set initial similarity matching thresholds for each disease type characteristic, and determine the confidence range based on the preliminary detection results of disease samples to ensure the basic classification accuracy of the model for major disease types;
[0072] D2: Dynamic Threshold Adjustment and Automatic Optimization: During the classification process, if the detection confidence is lower than the preset standard, the system will automatically increase the matching standard of the similarity threshold to effectively reduce the misjudgment rate of disease recognition;
[0073] D3: Progressive Threshold Matching Optimization: In multiple rounds of detection tasks, the system gradually tightens the similarity threshold for each type of disease, and progressively optimizes the threshold standard to handle the detection of complex disease types, ensuring the high accuracy and stability of the classification results.
[0074] In a specific implementation manner of the first aspect, by generating pseudo-labels and iteratively optimizing the disease recognition model to improve the recognition ability under few-shot and zero-shot conditions, the following steps are included:
[0075] E1: Pseudo-label Generation and Verification: Label the disease detection results with low confidence to generate pseudo-labels, and conduct consistency verification on the generated pseudo-label samples to ensure that the labeling of pseudo-labels conforms to the actual disease characteristics; The pseudo-sample set is defined as:
[0076]
[0077] Add the pseudo-sample set to the original training dataset for dynamic update to optimize the model's adaptability to the low-sample environment;
[0078] E2: Iterative Training and Dynamic Update: Add the pseudo-label samples to the original set of prompt words, and dynamically update the feature distribution of the multi-modal embedding to gradually strengthen the model's adaptive learning ability for different disease types;
[0079] E3: Multi-round Pseudo-label Optimization and Model Expansion: Through the iterative process of multi-round pseudo-label generation, verification, and update, continuously expand the coverage of prompt words and the disease feature distribution, enabling the model to have stronger robustness and extensive recognition ability in the few-shot environment.
[0080] In the second aspect, a mine disease detection system based on a vision-language model and prompt optimization, which is composed of the following modules:
[0081] Disease Data Acquisition Module: Used to automatically collect and preprocess disease images in the mine environment, and generate structured image data for model processing;
[0082] Multi-modal Embedding Module: Input the disease image and the constructed text prompt words into the vision-language pre-trained model to generate multi-modal embedding vectors that fuse image and text semantics;
[0083] Similarity calculation and classification module: Using the cosine similarity calculation model of multi-modal embedding vectors, it completes the similarity matching between the disease image and the text prompt, and realizes the classification of disease types;
[0084] Adaptive threshold adjustment module: According to the confidence level of disease detection, it automatically adjusts the similarity matching threshold to ensure the accuracy of the classification result;
[0085] Pseudo-label adaptive learning module: In the few-shot environment, it generates pseudo-label samples and introduces an iterative learning mechanism to improve the model's adaptive detection ability for mine diseases.
[0087] The beneficial effects of the present invention are as follows:
[0088] 1. High precision and high adaptability
[0089] Through the combination of vision-language pre-training model and prompt optimization, the present invention realizes high-precision detection of mine diseases in zero-shot and few-shot environments, and is especially suitable for complex scenarios that are difficult to cover by traditional deep learning.
[0090] 2. Low data dependence
[0091] By constructing domain-specific prompts and pseudo-label adaptive learning, the present invention reduces the dependence on a large amount of labeled data and realizes automated disease detection based on limited samples.
[0092] 3. Adaptive learning strategy with strong adaptability
[0093] The dynamic threshold matching and pseudo-label iterative optimization mechanism enable the model of the present invention to have strong adaptability in different mine environments and variable disease characteristics, and can effectively improve the robustness of detection. Description of the drawings
[0095] Figure 1 It is a schematic diagram of the overall system architecture of the present invention. Specific implementation manners
[0096] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0097] Such as Figure 1 The mine disease recognition method based on the vision-language model and prompt optimization shown.
[0098] In the data preprocessing stage, the collected mine disease images are standardized so that the size and resolution of the images meet the input requirements of the vision-language pre-training model (such as CLIP). The preprocessing steps include size adjustment, grayscale processing, and denoising to reduce unnecessary information interference in the images. Finally, the standardized image data I is obtained. i .
[0099] 2. Prompt Construction and Optimization
[0100] To meet the specific detection requirements of mine diseases, a domain-specific prompt set T = {t1, t2, …, t n} is constructed. The prompts include descriptions of common mine diseases such as "crack", "settlement", and "pothole". During the construction process, the prompt set is optimized through semantic analysis and word frequency screening to ensure that the core prompts have high semantic relevance. The screening condition is:
[0101] T core = {t k | freq(t k ) > δ}
[0102] where freq(t k ) represents the frequency of the vocabulary t k in the disease description, and δ is the set screening threshold. The optimized prompt set will be used as the text input of the model.
[0103] 3. Multimodal Embedding Generation
[0104] During the multimodal embedding generation process, the preprocessed disease image I i and the optimized text prompt T j are respectively input into the visual encoding and language encoding modules of the vision-language pre-training model to generate the feature vectors of the image and text:
[0105] V i = VisualEncoder(I i ), T j = TextEncoder(T j )
[0106] By combining the image feature V i and the text feature T j , a unified multimodal embedding E i is generated for disease detection in the multimodal space:
[0107] E i = W v V i + W t T j
[0108] Among them, W and W t are the weight matrices of the model, ensuring the effective fusion of image and text features.
[0109] 4. Data Augmentation and Multi-view Prompt Generation
[0110] To improve the robustness of the model in complex scenarios, various data augmentation processes are performed on the disease image data, including perspective rotation, scale change, and mirror flipping, to generate multi-view and multi-scale visual prompts:
[0111] Perspective rotation: Rotate the disease image by a preset angle (such as 45 degrees, 90 degrees, 180 degrees) to generate multi-angle images:
[0112] I θ = R θ (I)
[0113] Scale change: Generate image data of different sizes through scaling operations:
[0114] I s = S s I
[0115] Mirror flipping: Generate multi-directional disease images through horizontal and vertical flips. These enhanced visual prompt sets can further enrich the input data of the model, ensuring its high generalization ability in variable mine scenarios.
[0116] 5. Dynamic Threshold Matching
[0117] To improve the accuracy of disease classification, a dynamic threshold matching strategy is adopted. Set a basic similarity matching threshold τ init at the initial classification stage. If the detection confidence is lower than the expected standard, the similarity threshold is automatically increased. The formula for dynamic adjustment is:
[0118] τ new = τ init + Δτ
[0119] where Δτ is the adjustment increment. During multiple rounds of detection, as the detection task becomes more complex, a progressive optimization strategy is used to gradually tighten the threshold:
[0120] τ n+1 = τ n × 1 + α
[0121] where α is the optimization coefficient. This strategy gradually improves the classification accuracy in multiple rounds of detection tasks, ensuring the detection accuracy of complex diseases.
[0122] 6. Pseudo-label Generation and Adaptive Learning
[0123] Under few-shot conditions, the recognition effect of the model is improved through pseudo-label adaptive learning. Pseudo-labels are generated for detection results with insufficient confidence and added to the pseudo-sample set.
[0124]
[0125] Among them, S ij represents the similarity between the disease image and the prompt word, and pseudo-labels are generated when it is lower than the confidence threshold τ conf The pseudo-label samples and the original training set are used for iterative learning together. The pseudo-label update formula is:
[0126]
[0127] Through multiple rounds of iterative updates, the model gradually improves its adaptability and robustness to few-shot disease features.
[0128] This implementation process, through steps such as data preprocessing, prompt optimization, multi-modal embedding generation, data augmentation, dynamic threshold matching, and pseudo-label adaptive learning, makes full use of the multi-modal capabilities of the vision-language model, effectively improves the accuracy and robustness of mine disease detection, and achieves efficient recognition under zero-shot and few-shot conditions.
[0129] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A mine disease recognition method based on a vision-language model and prompt optimization, characterized in that The method includes the following steps: S1: Multimodal Embedding Generation: Input the mine disease image and the text prompt words constructed based on domain-specific knowledge into the vision-language pre-trained model respectively. The model extracts the visual feature vector V i and the text feature vector T j and maps the two into the same embedding vector space. The core representation of embedding generation is: E i = f(V i , T j ) = W v V i + W t T j ; Among them, W v and W t are the weight matrices for visual and textual embeddings respectively; S2: Prompt construction and optimization: Based on the specific manifestation forms of mine diseases, extract the description words related to diseases from the dictionaries and professional term libraries in the fields of mining and construction engineering, construct a text prompt word set covering various aspects of disease types, disease causes, disease forms, etc., extract high-frequency words as the core description word set through semantic analysis, and further perform semantic analysis on the text prompt words to ensure that the prompt words have a high degree of semantic relevance in the multi-modal embedding space, defined as: T core = {t k | freq(t k ) > δ} where freq(t k ) represents the frequency of the word t k , and δ is a preset frequency threshold for screening the core words of disease descriptions; S3: Disease Identification and Classification: Cosine Similarity S Based on Multimodal Embedding Vectors ij , calculate the similarity score between the input disease image and the prompt words, classify the disease type by selecting the prompt word with the highest score. If the similarity does not reach the preset confidence threshold, generate pseudo-labels for the detection results and add them to the training set, and perform iterative learning, so as to gradually improve the classification performance of the model in the few-shot environment; Complete the disease type classification by selecting the prompt with the highest similarity. If the similarity S ij < τ, generate pseudo-labels and improve the few-shot detection performance through iterative optimization.
2. The mine disease identification method based on visual - language model and prompt optimization according to claim 1, wherein, Generate multi-modal embeddings for the images and text prompts of mine diseases through a vision-language pre-trained model. The system includes the following steps: A1: Disease Image Embedding Generation: Input disease image data into the visual encoding module of the vision-language pre-trained model one by one, and extract the visual feature vector V through the image encoding layer i , convert it into a high-dimensional vector representation to capture the visual details of disease features; V i = σ(W enc * I i + b) Among which W enc is the encoding weight matrix, I i is the input image, * self-attention mechanism, and σ is the activation function; the generated visual embedding features ensure that the image details can be accurately characterized; A2: Generation of text prompt embeddings: By inputting the constructed disease text prompts into the language encoding module of the same model, generate text feature vectors to capture the semantic information of the diseases and maintain a consistent representation with the image embeddings; the text encoding formula is: T j = σ(W text ·T raw + b text ) Among which W text is the text encoding matrix, and T raw is the original text vector, ensuring that text features can be consistently embedded in the high-dimensional space; A3: Embedding alignment and optimization: Use the principal component analysis (PCA) method to perform dimensionality reduction on the embedding vectors of the images and text prompts, project the embedding features of different disease types into the same high-dimensional feature space, so as to maximize the feature separation degree between different disease categories, thereby improving the accuracy and stability of disease classification.
3. The mine disease identification method based on visual - language model and prompt optimization according to claim 1, wherein By constructing and optimizing a domain-specific multi-level text prompt word set to improve the performance of the vision-language pre-trained model in zero-shot and few-shot learning. The method includes the following steps: B1: Domain Vocabulary Extraction and Construction: Extract relevant terms covering common diseases from the professional glossaries in the fields of mine geology and construction engineering, and generate a complete hierarchical sentence description T according to information such as disease types, disease causes, and disease forms, full for covering complex diseases in the mine environment; where t k is a disease vocabulary, and α k is the vocabulary weight coefficient, which is used to optimize the integrity of semantic information; B2: Semantic analysis and core word extraction: For the generated complete sentence prompts, screen out high-frequency words through word frequency analysis and extract the core terms for disease description to form a core word list, ensuring that the model can still accurately capture the main semantic features of the prompts under low-sample data conditions; B3: Prompt formatting and standardization: Perform standardization processing on the sentence formats of the complete sentence prompts and the core word list to ensure the grammatical consistency and vocabulary standardization of the prompts, enabling the model to more efficiently understand the semantic connotations of the text prompts.
4. The mine disease identification method based on visual - language model and prompt optimization according to claim 1, characterized in that Generate multi-view visual prompts by performing various data augmentation processes on mine disease images to enhance the robustness of mine disease recognition, including the following steps: C1: Perspective rotation and scale change: Perform multi-angle rotation operations on the original disease images, including angle adjustments such as 45 degrees, 90 degrees, 180 degrees, etc., and at the same time scale the images at different scaling ratios to generate multi-view and multi-scale visual prompts, enabling the model to handle the complex manifestations of diseases at different angles and scales in the mine; C2: Multi-scale transformation operation: Perform multi-scale magnification and reduction of the disease images from micro to macro to simulate the visual features of the diseases at different shooting distances, ensuring that the visual prompts can cover the full-size feature information of the mine diseases; C3: Mirror flipping and symmetry enhancement: Perform symmetry enhancement processing on the images through horizontal flipping and vertical flipping to obtain disease visual prompts in multiple directions, ensuring that the model has a robust recognition ability for multi-directional perspectives of diseases in the mine.
5. The mine disease identification method based on visual-language model and prompt optimization according to claim 1, characterized in that, Achieve high-precision classification of mine diseases by dynamically adjusting the similarity matching threshold for disease detection, including the following steps: D1: Threshold Initialization and Confidence Setting: Set initial similarity matching thresholds for each disease type characteristic, and determine the confidence range based on the preliminary detection results of disease samples to ensure the basic classification accuracy of the model for major disease types; D2: Dynamic Threshold Adjustment and Automatic Optimization: During the classification process, if the detection confidence is lower than the preset standard, the system will automatically increase the matching standard of the similarity threshold to effectively reduce the misjudgment rate of disease recognition; D3: Progressive Threshold Matching Optimization: In multiple rounds of detection tasks, the system gradually tightens the similarity thresholds for each type of disease, and progressively optimizes the threshold criteria to handle the detection of complex disease types, ensuring the high precision and stability of the classification results.
6. The mine disease identification method based on visual - language model and prompt optimization according to claim 1, characterized in that, By generating pseudo-labels and iteratively optimizing the disease recognition model to improve the recognition ability under few-shot and zero-shot conditions, including the following steps: E1: Pseudo-label Generation and Verification: Label the disease detection results with low confidence to generate pseudo-labels, and verify the consistency of the generated pseudo-label samples to ensure that the labeling of pseudo-labels conforms to the actual disease characteristics; The pseudo-sample set is defined as: Add the pseudo-sample set to the original training dataset for dynamic update to optimize the model's adaptability to the low-sample environment; E2: Iterative Training and Dynamic Update: Add the pseudo-label samples to the original set of prompt words, and dynamically update the feature distribution of the multi-modal embedding to gradually strengthen the model's adaptive learning ability for different disease types; E3: Multi-round Pseudo-label Optimization and Model Expansion: Through the iterative process of multi-round pseudo-label generation, verification, and update, continuously expand the coverage of prompt words and the disease feature distribution, enabling the model to have stronger robustness and extensive recognition ability in the few-shot environment.
7. A mine disease detection system based on a vision-language model and prompt optimization, characterized in that, This system consists of the following modules: Disease Data Acquisition Module: Used to automatically collect and preprocess disease images in the mine environment, generating structured image data for model processing; Multi-modal Embedding Module: Input disease images and constructed text prompt words into a vision-language pre-trained model to generate multi-modal embedding vectors that fuse image and text semantics; Similarity Calculation and Classification Module: Use the cosine similarity calculation model of multi-modal embedding vectors to complete the similarity matching between disease images and text prompt words, and achieve the classification of disease types; Adaptive Threshold Adjustment Module: Automatically adjust the similarity matching threshold according to the confidence of disease detection to ensure the accuracy of classification results; Pseudo-label Adaptive Learning Module: In the few-shot environment, generate pseudo-label samples and introduce an iterative learning mechanism to improve the model's adaptive detection ability for mine diseases.
Citation Information
Cited By
Congenital heart disease perioperative period risk early warning method and system based on fundus color photo
CN120895240A