Coal mine underground target detection method and video analysis equipment based on visual large model
By adopting a visual big model-based method in underground target detection of coal mines, combining image enhancement, hyperparameter optimization, model fine-tuning and knowledge distillation technology, the problem that existing models can only realize single-task scenario detection, and the ability of multi-task scenario detection in underground coal mines and the lightweight deployment of models is achieved.
Patent Information
- Application Number
- CN202410539005.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-04-30
AI Technical Summary
The existing underground target detection model of coal mines can only realize the detection of a single task scenario, and the application scenarios are limited, making it difficult to meet the needs of multi-task scenarios.
Using a visual big model-based method, the Grounding DINO model combines the Bayesian optimization method, fine-tuning and knowledge distillation technology of tree structure to achieve multi-task scene detection. Specific steps include image enhancement, hyperparameter optimization, model fine-tuning, knowledge distillation and object detection.
The ability of multi-task scenario detection in the coal mine is realized, the generalization ability and detection accuracy of the model are improved, and the model is lightweight through knowledge distillation, which is suitable for on-site deployment in the coal mine.
Smart Images

Figure CN118411512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a coal mine underground target detection method and video analysis equipment based on a large visual model. Background Art
[0002] The underground detection algorithm of coal mines mainly involves the target detection technology in the field of computer vision. The existing target detection technologies include: early traditional target detection algorithms and target detection algorithms based on deep learning. Traditional target detection algorithms such as VJ detectors and HOG detectors mainly realize simple target detection based on manually designed features; target detection based on deep learning mainly uses convolutional neural networks (CNN) to realize automatic feature extraction through model training, which can be divided into two-stage detectors and one-stage detectors.
[0003] The two-stage detector divides the target detection task into two stages: candidate region generation and target classification and positioning. For example, models such as RCNN (Region with Convolutional Neural Network), Fast RCNN (Fast Regional Convolutional Neural Network), and Faster RCNN (Faster Regional Convolutional Neural Network); the single-stage detector regards target detection as a regression problem. Compared with the two-stage detector, it has faster detection speed and stronger generalization ability, such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) models.
[0004] Traditional target detection algorithms have a limited scope of application, low detection accuracy, and can only perform simple target detection, so they have been eliminated and abandoned.
[0005] At present, the mainstream deep learning target detection mainly adopts Faster RCNN and YOLO models, and its target detection accuracy has been significantly improved, and it has been widely used in coal mine underground detection. However, in actual applications, coal mine underground detection involves multiple task scenarios such as underground personnel detection, safety helmet detection, illegal smoking, etc. The existing deep learning target detection model can only realize the "one-to-one mode" of training one model to realize one task scenario, and the application scenarios are limited. Summary of the invention
[0006] The present invention provides a coal mine underground target detection method and video analysis equipment based on a large visual model, the main purpose of which is to achieve multi-task scene detection.
[0007] In a first aspect, an embodiment of the present invention provides a method for detecting underground coal mine targets based on a large visual model, comprising:
[0008] Acquire a video image from underground in a coal mine, perform image enhancement on the video image to obtain an enhanced image, and use the enhanced image and annotated text as a training set;
[0009] The tree-structured Bayesian optimization method is used to optimize hyperparameters and determine the values of hyperparameters in the Grounding DINO model.
[0010] Based on the training set, fine-tuning the Grounding DINO model to obtain a fine-tuned Grounding DINO model;
[0011] Perform knowledge distillation on the fine-tuned Grounding DINO model to obtain the Grounding DINO model after knowledge distillation;
[0012] The image to be detected in the coal mine is input into the Grounding DINO model after knowledge distillation to obtain the target detection result.
[0013] In a second aspect, an embodiment of the present invention provides a coal mine underground target detection system based on a visual large model, comprising:
[0014] An enhancement module is used to acquire a video image of a coal mine, perform image enhancement on the video image to obtain an enhanced image, and use the enhanced image and the annotated text as a training set;
[0015] The hyperparameter determination module is used to optimize the hyperparameters based on the tree-structured Bayesian optimization method and determine the values of the hyperparameters in the Grounding DINO model.
[0016] A fine-tuning module, used for fine-tuning the Grounding DINO model based on the training set to obtain a fine-tuned Grounding DINO model;
[0017] The knowledge distillation module is used to perform knowledge distillation on the fine-tuned Grounding DINO model to obtain the Grounding DINO model after knowledge distillation;
[0018] The detection module is used to input the image to be detected in the coal mine into the GroundingDINO model after knowledge distillation to obtain the target detection result.
[0019] In a third aspect, an embodiment of the present invention provides a video analysis device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for detecting underground targets in coal mines based on a large visual model when executing the computer program.
[0020] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, wherein the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for detecting underground coal mine targets based on a large visual model are implemented.
[0021] The invention proposes a coal mine underground target detection method and video analysis equipment based on a visual big model. Compared with the single target detection scene of the existing coal mine underground target detection algorithm, the Grounding DINO visual big model that integrates text and image understanding capabilities is introduced. It only needs to adjust the text prompts of specific target detection to achieve multi-task scene detection, and has a strong model generalization ability. In addition, based on the pre-training ability of the visual big model with its large-scale data set, the vertical field ability of the coal mine underground scene is further enhanced by fine-tuning a small number of samples; guided filtering is used to extract illumination components, two-dimensional gamma function illumination correction, adaptive histogram equalization operator and other image enhancement technologies to effectively improve the uneven illumination and similar background problems of coal mine underground images and improve the detection effect; finally, the lightweight of the visual big model is achieved through knowledge distillation based on the Bayesian optimization method, the detection accuracy and detection speed are balanced, and it is convenient for field deployment underground in coal mines. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of a method for detecting underground coal mine targets based on a large visual model provided by an embodiment of the present invention;
[0023] Figure 2 A flowchart of enhancing a video image in an embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of fine-tuning the Grounding DINO model in an embodiment of the present invention;
[0025] Figure 4 is a flow chart of fine-tuning the Grounding DINO model in an embodiment of the present invention;
[0026] Figure 5 A schematic diagram of the structure of a coal mine underground target detection system based on a large visual model provided by an embodiment of the present invention;
[0027] Figure 6 A schematic diagram of the structure of a video analysis device provided by an embodiment of the present invention.
[0028] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0029] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0030] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0031] In the embodiments of the present application, at least one refers to one or more; multiple refers to two or more. In the description of the present application, words such as "first", "second", "third", etc. are only used to distinguish the purpose of description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0032] References to "one embodiment" or "some embodiments" etc. described in this specification mean that one or more embodiments of the present application include a particular feature, structure or characteristic described in conjunction with the embodiment. Thus, in this specification, the terms "include", "comprises", "has" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0033] Figure 1 A flow chart of a method for detecting underground coal mine targets based on a large visual model is provided in an embodiment of the present invention, such as Figure 1 As shown, the method includes:
[0034] S110, acquiring a video image of an underground coal mine, performing image enhancement on the video image to obtain an enhanced image, and using the enhanced image and the annotated text as a training set;
[0035] In order to be better applicable to underground coal mine detection in the embodiment of the present invention, a real-time video image of the underground coal mine is first obtained. The video image adopts the COCO (Common Objects in COntext) format, including target category information, image information and annotation information. The annotation information can be determined according to actual conditions, such as detecting coal miners, detecting safety helmets, detecting smoking, etc., and the embodiment of the present invention does not make specific limitations on this.
[0036] Among them, the COCO format is a dataset that can be used for image recognition.
[0037] Due to the special environment underground in coal mines, such as high light sources, object occlusion, similar background, etc., the video images may have problems of poor quality and low recognition rate. Therefore, in an embodiment of the present invention, image enhancement processing is performed on the video images to improve the image quality in the coal mine environment, improve the visibility and detail clarity of the image, enhance the image quality and the diversity of training data, and help improve the accuracy of underground coal mine detection.
[0038] As an implementation method, Figure 2 FIG. 1 is a flow chart of enhancing a video image in an embodiment of the present invention. Figure 2 As shown, the step of performing image enhancement on the video image to obtain an enhanced image includes:
[0039] Acquire the brightness of the video image and extract the illumination component through guided filtering;
[0040] Processing the illumination component using a two-dimensional gamma function to obtain a corrected brightness component;
[0041] Resynthesize the image based on the hue and saturation of the video image in combination with the corrected brightness component;
[0042] The synthesized image is processed by CLAHE to obtain an enhanced image.
[0043] The initial color mode of the video image is RGB (Red Green Blue). Convert RGB to HSV (Hue Saturation Value) space to obtain the brightness V, hue H and saturation S of the video image. Guided filtering is performed on the brightness to extract the illumination component. It should be noted that guided filtering has extremely strong stability for the target edge of the video image. For the video surveillance image of the coal mine, due to the illumination of the light source, a halo may appear around the target, resulting in inaccurate extraction of the target illumination component. Accurate extraction of the illumination component plays an important role in image enhancement.
[0044] (1) Guided filtering is performed on the underground coal mine video image P through the guide map I to extract the illumination component, thereby effectively maintaining the edge, achieving edge smoothing, and reducing gradient deformation.
[0045] Function construction: Assume that the guided filter function satisfies a linear relationship between the input and output video images, that is:
[0046]
[0047] Among them, q is the value of the video image pixel; I is the value of the guide image pixel; i and k are pixel indexes; a and b are coefficients of the linear function.
[0048] Optimal solution: The optimal solution is obtained by the least squares method. The linear coefficient of each filter window can be expressed by the following formula:
[0049]
[0050] Where: μ k is the average value of the guide image I in the filtering window ω; k is the average value of the image to be filtered p in ω; is the variance of I in ω; |ω| is the number of pixels in ω; Ε is a parameter established to prevent a from being too large in regularization.
[0051] (2) A two-dimensional gamma function is used to perform illumination correction on the illumination component to obtain a corrected brightness component.
[0052] The traditional one-dimensional gamma function uses a fixed value for global adjustment, while the present invention proposes a two-dimensional gamma function that can adaptively adjust the brightness value of each pixel according to its characteristics, thereby achieving dynamic correction of the uneven lighting area of the coal mine lamp.
[0053] The expression of the two-dimensional gamma function is:
[0054]
[0055] Wherein: F(x, y) is the corrected brightness component; m is the mean value of the illumination component; v(x, y) is the brightness component of the video image; V(x, y) is the corrected brightness component.
[0056] This function can dynamically adjust the brightness value of the pixel according to the specific lighting conditions, reduce the brightness of the overexposed area, and increase the brightness of the dark area, effectively solving the problem of uneven lighting.
[0057] The hue and saturation of the video image and the corrected brightness component are resynthesized. Since the image tone in the mine is simple and the image contrast is low, it is difficult to distinguish the details of the underground objects, and the video image may also have some noise and other influencing factors, the contrast accounts for a large proportion of the image imaging effect. Therefore, the embodiment of the present invention uses the CLAHE algorithm for processing.
[0058] The CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm includes two steps: threshold limiting and pixel value average distribution. Setting a reasonable threshold for the mine image histogram can suppress the noise amplification problem; pixel value average distribution redistributes the pixel values exceeding the threshold to the entire grayscale range (0-255), expands the dynamic range, and enhances the contrast. Compared with traditional histogram equalization, this algorithm can effectively overcome the problem of noise amplification.
[0059] After the above three enhancement operations, an enhanced video image can be generated, which solves the problems of uneven lighting and similar background in the mine environment, and prepares for subsequent target detection.
[0060] S120, performing hyperparameter optimization based on a tree-structured Bayesian optimization method to determine the values of hyperparameters in the Grounding DINO model;
[0061] Among them, the Grounding DINO model, as a large visual model, combines the image transformer-based detector (DINO) and the object detector based on the language positioning pre-trained model, which can realize object detection based on a given image and text description. By using the text and image alignment understanding ability of the Grounding DINO large visual model, multi-task scene detection can be achieved by simply adjusting the text prompts for specific object detection. For example:
[0062] For the task of detecting personnel in underground coal mines, the sample images are labeled with the areas where the coal mine personnel are located. The Grounding DINO model is trained using the labeled samples, so that the Grounding DINO model has the ability to detect personnel in underground coal mines.
[0063] For the helmet detection task, the sample image is annotated with the area where the coal mine workers’ helmets are located. The annotated samples are used to train the Grounding DINO model, so that the Grounding DINO model has the ability to detect whether the coal mine workers are wearing helmets. When it is detected that they are not wearing helmets, the corresponding management personnel can be notified and a safety warning can be issued.
[0064] For the illegal smoking detection task, the sample images are annotated to determine whether coal mine personnel have smoking behavior. The Grounding DINO model is trained using the annotated samples so that the Grounding DINO model can detect whether coal mine personnel have smoking behavior. When smoking behavior is detected, the corresponding management personnel can be notified and a safety warning can be issued.
[0065] The enhanced video image is input into the image encoder of the Grounding DINO model to extract the feature expression of the image and generate the original image features. At the same time, the annotated text corresponding to the enhanced video image is input into the text encoder to obtain the original text features. The two are respectively obtained through the feature enhancer to obtain the enhanced image features and text features.
[0066] Here, we mainly adjust the image encoder’s image block size (patch size), moving window size (window size), latent vector dimension (hidden size), hidden layer dropout ratio (dropout rate) and model training learning rate (learning rate) to perform hyper-parameter optimization, so as to enhance the feature expression of image coding and improve target detection capabilities.
[0067] The hyperparameters include image block size, moving window size, latent vector dimension, dropout ratio of hidden layer, and learning rate. The specific meanings of each hyperparameter are shown in Table 1:
[0068] Table 1
[0069]
[0070]
[0071] As an implementation mode, the tree-structured Bayesian optimization method performs hyperparameter optimization to determine the value of the hyperparameter in the Grounding DINO model, including:
[0072] Defining the objective function and the value range of the hyperparameters;
[0073] Selecting a current hyperparameter combination within the value range, setting a first Gaussian mixture model for the current hyperparameter combination, and setting a second Gaussian mixture model for other hyperparameter combinations;
[0074] Modeling the conditional probability distribution and marginal probability distribution of the current hyperparameter combination respectively, and calculating the posterior probability distribution by using the Bayesian formula;
[0075] Select the expected improvement as the acquisition function, and determine the optimization objective equation according to the first Gaussian mixture model, the second Gaussian mixture model, and the posterior probability distribution;
[0076] Based on the posterior probability distribution, obtain the optimization generation number of the current hyperparameter combination. If the optimization generation number does not meet the preset requirements, use the hyperparameter combination corresponding to the maximum ratio of the first Gaussian mixture model to the second Gaussian mixture model as the current hyperparameter combination again, and repeat the above steps until the optimization generation number meets the preset requirements, and obtain the values of the hyperparameters.
[0077] In the embodiments of the present invention, in order to make full use of limited training resources, a tree-structured Bayesian optimization method (Tree-structured Parzen Estimator, TPE) is used for hyperparameter optimization. TPE uses Gaussian mixture models to learn model hyperparameters, reduces the parameter search space, and improves the search efficiency.
[0078] For the hyperparameter sequence x, TPE maintains the first Gaussian mixture model l(x) for the hyperparameters related to the best objective value, maintains the second Gaussian mixture model g(x) for the remaining hyperparameters, and selects the hyperparameters corresponding to the maximum of l(x) / g(x) as the next set of search values.
[0079] TPE models the conditional probability distribution p(x|y) and the marginal probability distribution p(y) respectively, and calculates the posterior probability distribution p(y|x) through Bayes' formula. Where:
[0080]
[0081] TPE selects the expected improvement (EI) as the acquisition function (the acquisition function is used to determine where to collect the next sample point), and the final optimization objective equation is:
[0082]
[0083] Where: y* represents the threshold, let γ = p(y < y*), which is used to divide l(x) and g(x), and the range is between (0, 1).
[0084] Bayesian optimization process:
[0085] (1) Define the objective function and the value range of the hyperparameters;
[0086] (2) Sample a part of the parameter combinations in the value range as the current hyperparameter combination, and evaluate the performance of the current hyperparameter combination on the objective function;
[0087] (3) Based on historical observation data, a Gaussian mixture model is established to approximate the posterior distribution of the target function;
[0088] (4) Using the posterior distribution to obtain the new sample with the largest expected improvement, add it to the observed data;
[0089] (5) Repeat steps (3) and (4) until the preset optimization number or performance requirements are met.
[0090] Through the above process, the optimal hyperparameter combination of the encoder can be efficiently determined, so as to further improve the encoding quality of the visual encoder and generate more refined feature expressions.
[0091] S130, fine-tuning the Grounding DINO model based on the training set to obtain a fine-tuned Grounding DINO model;
[0092] The embodiment of the present invention uses the GroundingDINO visual model, which is a general visual understanding model that has achieved leading performance on multiple benchmark datasets. However, if it is directly applied to coal mine scenarios, the generalization ability may be poor. Therefore, it is necessary to fine-tune it on the mine dataset to adapt the large model to the target field.
[0093] The basic idea of model fine-tuning is to use the knowledge learned in the pre-trained model to solve new tasks, which can save training time and computing resources and improve the generalization ability and performance of the model.
[0094] The specific method is: obtain a coal mine underground scene dataset, and fine-tune the GroundingDINO model under the selected model hyperparameter conditions; through fine-tuning with a few samples, the model's target detection capability in this industry scenario can be further improved.
[0095] Fine-tuning of the large visual model only updates the parameters of the higher or last layer in the Grounding DINO model network structure to avoid overfitting or catastrophic forgetting.
[0096] As an implementation manner, the Grounding DINO model is fine-tuned based on the training set to obtain a fine-tuned Grounding DINO model, including:
[0097] Load the pre-trained weight parameters of the Grounding DINO model;
[0098] Read the training data in the training set, input it into the Grounding DINO model, and perform forward operation on the Grounding DINO model;
[0099] Calculate the loss value based on the forward operation result and the true value;
[0100] If the set number of evaluation steps is reached, the model is evaluated and the target detection mAP evaluation index is calculated. Otherwise, back propagation is performed to update the network parameters of the Grounding DINO model.
[0101] Repeat the above steps until the set training index or the maximum number of rounds is reached, and the fine-tuned Grounding DINO model is obtained, and the fine-tuning is finished.
[0102] Figure 3 is a schematic diagram of fine-tuning the Grounding DINO model in an embodiment of the present invention, Figure 4 FIG. 1 is a flow chart of fine-tuning the Grounding DINO model in an embodiment of the present invention, such as Figure 3 and Figure 4 As shown in Figure 2, the fine-tuning target model is represented by fitting a pre-trained model to a regression task, and its optimization objective is expressed as:
[0103]
[0104] Where: W Guided and W Hint Represent the Grounding DINO model and pre-trained model parameters respectively.
[0105] (2) Fine-tuning process, the process of fine-tuning with a small number of samples based on the coal mine underground scene dataset is as follows:
[0106] 1. Build training and validation data sets;
[0107] 2. Build the model and load the pre-trained weight parameters;
[0108] 3. Traverse and read batch training data;
[0109] 4. Perform model forward operation based on input data;
[0110] 5. Calculate the loss between the forward operation result and the true value;
[0111] 6. If the set number of evaluation steps is reached, the model is evaluated and the target detection mAP evaluation index is calculated; otherwise, back propagation is performed to update the network parameters of the model;
[0112] Repeat steps 3 to 6 until the set training index or the maximum number of epochs is reached, save the model file, and end fine-tuning.
[0113] Through this fine-tuning strategy, a high-performance target detection model suitable for coal mine scenarios can be obtained. It is worth noting that fine-tuning only requires a small amount of domain data to converge, which greatly reduces the amount of data annotation and improves the efficiency of model development.
[0114] S140, performing knowledge distillation on the fine-tuned Grounding DINO model to obtain a knowledge distilled Grounding DINO model;
[0115] In response to the real-time requirements of coal mines, the knowledge distillation method is used to reduce the number of model parameters and calculations while ensuring model accuracy, thereby improving the model's running speed.
[0116] As an implementation mode, performing knowledge distillation on the fine-tuned Grounding DINO model to obtain the Grounding DINO model after knowledge distillation includes:
[0117] The fine-tuned Grounding DINO model is subjected to Patch-group distillation and Anchor-point distillation respectively to obtain the Grounding DINO model after knowledge distillation.
[0118] Among them, knowledge distillation is a process of "teaching" the student model through the teacher model, so that the small model can have the knowledge / capabilities of the large model. That is, the output result q of the teacher model is used as the target of the student model, and the student model is trained so that the result p of the student model is close to q. The loss function is written as:
[0119] L = αCE(y,p) + CE(q,p);
[0120] Where: CE is cross entropy; y is the onehot encoding of the true label.
[0121] (1) TaT method:
[0122] In order to make the student model fully mimic the feature space components of the teacher model, Target-Ware Transformer (TaT) is used to reconfigure the semantics of student features at specific locations at the pixel level.
[0123] L TaT =||f s′ -f t || 2 ;
[0124] Where: f tis the Teacher feature, f s′ is the reconfigured Student feature.
[0125] (2) Layered distillation method:
[0126] In order to further reduce the computational complexity of TaT, a hierarchical distillation strategy is proposed, which is divided into two steps:
[0127] 1. Patch-group distillation: split the entire feature map into smaller patches, extract local information from the teacher model to the student model;
[0128] 2. Anchor-point distillation, merge local patches into a vector and extract global information.
[0129] The output of the teacher model is used as the training target of the student model, and the student is trained for knowledge distillation. The optimization goal is:
[0130]
[0131] Min(L_dis+λ_1Lpatch+λ_2Lanchor);
[0132] Where L CE is the distillation loss, is the patch-group loss, is the anchor-point loss.
[0133] Through end-to-end joint optimization, the student model can maximize the generalization ability of the teacher while maintaining efficient computing.
[0134] S150, inputting the image to be detected in the coal mine into the Grounding DINO model after knowledge distillation to obtain a target detection result.
[0135] Finally, a lightweight model based on knowledge distillation is used to perform field target detection on the input coal mine underground video images.
[0136] The invention proposes a method for underground coal mine target detection based on a visual big model. Compared with the single target detection scenario of the existing underground coal mine target detection algorithm, the GroundingDINO visual big model that integrates text and image understanding capabilities is introduced. It only needs to adjust the text prompts of specific target detection to realize multi-task scene detection, and has a strong model generalization ability. In addition, based on the pre-training ability of the large-scale data set of the visual big model, the vertical field ability of the underground coal mine scene is further enhanced by fine-tuning a small number of samples. Image enhancement technologies such as guided filtering to extract illumination components, two-dimensional gamma function illumination correction, and adaptive histogram equalization operator are adopted to effectively improve the problems of uneven illumination and similar background of underground coal mine images, thereby improving the detection effect. Finally, the lightweight of the visual big model is realized through knowledge distillation based on the Bayesian optimization method, and the detection accuracy and detection speed are balanced, which is convenient for field deployment underground in coal mines.
[0137] Figure 5 A schematic diagram of a structure of a coal mine underground target detection system based on a visual large model provided by an embodiment of the present invention, such as Figure 5 As shown, the system includes:
[0138] The enhancement module 510 is used to obtain a video image of a coal mine, perform image enhancement on the video image to obtain an enhanced image, and use the enhanced image and the annotated text as a training set;
[0139] A hyperparameter determination module 520 is used to perform hyperparameter optimization based on a tree-structured Bayesian optimization method to determine the values of hyperparameters in the Grounding DINO model;
[0140] A fine-tuning module 530, configured to fine-tune the Grounding DINO model based on the training set to obtain a fine-tuned Grounding DINO model;
[0141] A knowledge distillation module 540 is used to perform knowledge distillation on the fine-tuned Grounding DINO model to obtain a knowledge-distilled Grounding DINO model;
[0142] The detection module 550 is used to input the image to be detected in the coal mine into the GroundingDINO model after knowledge distillation to obtain a target detection result.
[0143] This embodiment is a system embodiment corresponding to the above method embodiment, and its specific implementation process is the same as that of the above method embodiment. For details, please refer to the above method embodiment, and this system embodiment does not make specific limitations on this.
[0144] Each module in the above-mentioned coal mine underground target detection system based on visual large model can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the video analysis device in the form of hardware, or can be stored in the memory in the video analysis device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0145] Figure 6 A schematic diagram of the structure of a video analysis device provided in an embodiment of the present invention. The video analysis device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The video analysis device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the video analysis device is used to provide computing and control capabilities. The memory of the video analysis device includes a computer storage medium and an internal memory. The computer storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the computer storage medium. The database of the video analysis device is used to store data generated or obtained during the execution of a method for detecting underground coal mine targets based on a large visual model, such as the values of hyperparameters in the Grounding DINO model, training sets, etc. The network interface of the video analysis device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for detecting underground coal mine targets based on a large visual model is implemented.
[0146] In one embodiment, a video analysis device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of a method for detecting underground coal mine targets based on a large visual model in the above embodiment are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in this embodiment of a system for detecting underground coal mine targets based on a large visual model are implemented.
[0147] In one embodiment, a computer storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a method for detecting underground coal mine targets based on a large visual model in the above embodiment are implemented. Alternatively, when the computer program is executed by a processor, the functions of each module / unit in the above embodiment of a system for detecting underground coal mine targets based on a large visual model are implemented.
[0148] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0149] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0150] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for underground target detection in coal mines based on a large visual model, characterized in that: include: Acquire a video image from underground in a coal mine, perform image enhancement on the video image to obtain an enhanced image, and use the enhanced image and annotated text as a training set; The tree-structured Bayesian optimization method is used to optimize hyperparameters and determine the values of hyperparameters in the Grounding DINO model. Based on the training set, fine-tuning the Grounding DINO model to obtain a fine-tuned Grounding DINO model; Perform knowledge distillation on the fine-tuned Grounding DINO model to obtain the Grounding DINO model after knowledge distillation; Inputting the image to be detected in the coal mine into the Grounding DINO model after knowledge distillation to obtain the target detection result; The step of performing image enhancement on the video image to obtain an enhanced image includes: Acquire the brightness of the video image and extract the illumination component through guided filtering; Processing the illumination component using a two-dimensional gamma function to obtain a corrected brightness component; Resynthesize the image based on the hue and saturation of the video image in combination with the corrected brightness component; Perform CLAHE processing on the synthesized image to obtain an enhanced image; The tree-structured Bayesian optimization method is used to optimize hyperparameters and determine the values of hyperparameters in the Grounding DINO model, including: Defining the objective function and the value range of the hyperparameters; Selecting a current hyperparameter combination within the value range, setting a first Gaussian mixture model for the current hyperparameter combination, and setting a second Gaussian mixture model for other hyperparameter combinations; Modeling the conditional probability distribution and marginal probability distribution of the current hyperparameter combination respectively, and calculating the posterior probability distribution by using the Bayesian formula; Selecting expected improvement as an acquisition function, and determining an optimization objective equation according to the first Gaussian mixture model, the second Gaussian mixture model, and the posterior probability distribution; Based on the posterior probability distribution, the optimization algebra of the current hyperparameter combination is obtained. If the optimization algebra does not meet the preset requirements, the hyperparameter combination corresponding to the maximum ratio of the first Gaussian mixture model to the second Gaussian mixture model is used as the current hyperparameter combination again, and the above steps are repeated until the optimization algebra meets the preset requirements, and the value of the hyperparameter is obtained; The Grounding DINO model is fine-tuned based on the training set to obtain a fine-tuned Grounding DINO model, including: Load the pre-trained weight parameters of the Grounding DINO model; Read the training data in the training set, input it into the Grounding DINO model, and perform forward operation on the Grounding DINO model; Calculate the loss value based on the forward operation result and the true value; If the set number of evaluation steps is reached, the model is evaluated and the target detection mAP evaluation index is calculated. Otherwise, back propagation is performed to update the network parameters of the Grounding DINO model. Repeat the above steps until the set training index or the maximum number of rounds is reached, and the fine-tuned GroundingDINO model is obtained, and the fine-tuning is completed.
2. The method for detecting underground coal mine targets based on a large visual model according to claim 1, characterized in that: The fine-tuned Grounding DINO model is subjected to knowledge distillation to obtain a knowledge-distilled Grounding DINO model, including: The fine-tuned Grounding DINO model is subjected to Patch-group distillation and Anchor-point distillation respectively to obtain the Grounding DINO model after knowledge distillation.
3. The method for detecting underground coal mine targets based on a large visual model according to claim 1 or 2, characterized in that: The hyper parameters include image block size, moving window size, latent vector dimension, dropout ratio of hidden layer, and learning rate.
4. A video analysis device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for underground target detection in a coal mine based on a visual large model as described in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Machine learning model hyper-parameter tuning method based on multi-model Bayesian optimization
CN116862013A
Remote sensing image target detection and segmentation method and device, equipment and storage medium
CN117831042A